Mr.PlanB Logo

    Newsletter

    Subscribe our newsletter

    Get new infrastructure guides, comparison reports, and migration notes in your inbox.

    Infrastructure notes, guides, and new tools. Unsubscribe anytime.

    Back to Blog
    Kubernetes
    MariaDB
    Database
    Operators
    Stateful Workloads

    MariaDB Operator 25.10 and Stateful Workloads on Kubernetes

    October 31, 2025
    8 min read

    Running databases in Kubernetes has always felt a bit like fitting a square peg in a round hole. Kubernetes was designed for stateless applications, and for the longest time databases, those pesky stateful, disk-hungry, fail-sensitive creatures, have been treated like second-class citizens in the cloud-native ecosystem. With the release of MariaDB Operator 25.10, that is starting to shift.

    This update is a significant leap forward, especially for teams that want to run stateful workloads like MariaDB inside Kubernetes without duct-taping together half-baked solutions. The release makes asynchronous replication a fully supported feature, adds automated replica recovery, and bakes in several smart operational improvements that make running production-grade MariaDB clusters inside Kubernetes both possible and sane.

    Asynchronous replication goes GA, and it's actually solid

    The headline feature is the general availability (GA) of asynchronous replication. That may not sound thrilling unless you've watched a database go sideways at 3 AM. For most users, asynchronous replication means something pretty straightforward: a primary database server does all the writes, and one or more replicas follow along, pulling changes over as fast as they can.

    It's hardly bleeding-edge tech, since MySQL and MariaDB have supported it for ages. What changed is that MariaDB Operator now understands it deeply. You define a simple Kubernetes manifest, flip the replication switch, and you've got a primary-replica setup.

    apiVersion: k8s.mariadb.com/v1alpha1
    kind: MariaDB
    metadata:
      name: mariadb-repl
    spec:
      storage:
        size: 1Gi
        storageClassName: rook-ceph
      replicas: 3
      replication:
        enabled: true
    
    

    That's it. Behind the scenes, the operator sets up users, manages credentials, syncs the binary logs, and monitors replication lag. You don't need to babysit it; you define the desired state and let the operator handle the dirty work.

    Failover you can actually trust

    Automated primary failover is another big win. If your primary pod dies, and eventually it will, the operator automatically picks the most up-to-date replica and promotes it. It doesn't cross its fingers and hope the new primary has all the writes. The operator checks replication lag and relay log application status to make sure the candidate is clean.

    Here's a sample of what that looks like during a failover:

    NAME           READY   STATUS                                  PRIMARY          UPDATES
    mariadb-repl   False   Switching primary to 'mariadb-repl-1'   mariadb-repl-0   ReplicasFirstPrimaryLast
    ...
    NAME           READY   STATUS    PRIMARY          UPDATES
    mariadb-repl   True    Running   mariadb-repl-1   ReplicasFirstPrimaryLast
    
    

    That transition takes seconds. It's the kind of zero-touch recovery most teams wished for when managing databases manually, and now it's built in.

    You can even control failover behavior with settings like autoFailoverDelay to tune how aggressively the system promotes a new primary. That matters a lot for high-availability setups where uptime is measured in dollars per second.

    Replica recovery that doesn't suck

    Then there's the elephant in the cluster, replica corruption. Anyone who's dealt with asynchronous replication knows the pain of error code 1236, the dreaded "replica can't catch up because the primary purged the binary logs" situation. It's a silent killer that leaves your cluster in a weird limbo.

    MariaDB Operator 25.10 solves this with automated replica recovery using a construct called PhysicalBackup. If a replica can't recover normally, the operator triggers a recovery flow that takes a volume-level snapshot from a healthy replica and restores it into the broken one, all without manual intervention. Best of all, it actually works:

    kubectl get mariadb
    NAME           READY   STATUS                PRIMARY
    mariadb-repl   False   Recovering replicas   mariadb-repl-1
    ...
    kubectl get mariadb
    NAME           READY   STATUS    PRIMARY
    mariadb-repl   True    Running   mariadb-repl-1
    
    

    Recovery time depends on your storage driver and data size, but it's typically fast enough that you don't need to scramble. For teams running production-grade workloads, this is a godsend, turning replica recovery from a 30-minute firefight into a non-event.

    Smarter, safer scaling and backups

    The release goes beyond failover and recovery. It also offers flexible strategies for scaling out, including support for different backup methods. You can use fast, local VolumeSnapshots for rapid scaling, or switch to mariadb-backup for longer-term durability.

    That gives teams more control over how they balance performance and reliability. For example, you can keep one PhysicalBackup spec for nightly S3 backups and another for instant snapshot-based recovery. The operator supports both, and choosing the right one is as easy as plugging in a different template.

    The community's fingerprints are all over this

    Much of what makes 25.10 so good came from real-world feedback. Users in the open-source community reported issues with early replication support, submitted manual recovery runbooks, and pushed the maintainers to refine the operational experience.

    The maintainers, especially mmontes11, who appears to be leading a lot of the development, deserve props for listening and iterating. You can feel the difference between a feature built in a bubble and one forged through actual production use.

    As one user noted in the release discussion, many features exist today because people kept breaking their clusters and wanted better recovery paths. That kind of evolution is rare in projects trying to be everything to everyone.

    Not perfect, but getting close

    There's still room to grow. Right now the operator only supports replication within a single Kubernetes cluster, so replication across clusters or regions is off the table. That's a limitation for teams building multi-region failover systems, but given how fast things are moving, cross-cluster support feels like a matter of "when" rather than "if."

    The usual caveats about performance apply too. If you're running Kubernetes on-prem, local storage is a must, because networked volumes can become a bottleneck, especially with write-heavy workloads. As the maintainer put it, "Don't make any assumptions: run sysbench."

    Even with those constraints, MariaDB Operator 25.10 brings a level of confidence that stateful workloads inside Kubernetes have often lacked. It has grown out of the bolt-on experiment stage into something production-ready, thoughtfully built, and backed by a community that clearly cares.

    TL;DR

    MariaDB Operator 25.10 makes asynchronous replication work the way you'd want it to, automatically, intelligently, and resiliently. The release includes:

    • General availability of async replication
    • Automated failover to the most up-to-date replica
    • Snapshot-based replica recovery on error code 1236
    • Flexible backup strategies for different use cases

    That makes it a milestone release for anyone looking to move stateful workloads into Kubernetes with minimal drama. If you've been waiting for a sign that running a database in k8s isn't reckless, this is it.