Mr.PlanB Logo

    Newsletter

    Subscribe our newsletter

    Get new infrastructure guides, comparison reports, and migration notes in your inbox.

    Infrastructure notes, guides, and new tools. Unsubscribe anytime.

    Back to Blog
    proxmox
    ceph
    storage
    high-availability
    starwind
    nas

    Ceph, StarWind or Something Else for HA Storage in Proxmox?

    January 18, 2026
    9 min read

    Every Proxmox admin hits a very specific moment. You've got a small cluster humming along, live migration works, HA restarts mostly behave, and your dashboards look clean. Then you ask the forbidden question: what if I want my storage to just stay up when a node dies?

    You don't mean "eventually recover" or "degraded but fine if I don't touch it." You want boring, enterprise-style availability, with the same IP, the same mount and the same Samba or NFS share, so that one node disappears and everything keeps chugging like nothing happened. That's where things get weird.

    On paper, Proxmox gives you lots of options. In practice, most people end up stuck in an uncomfortable middle ground between "this technically works" and "why is this so much harder than my Synology." That tension is where Ceph, StarWind VSAN, clustered filesystems, and DIY NAS ideas start colliding.

    The Synology-shaped hole in Proxmox

    If you've ever run Synology High Availability, you know how spoiled it makes you. You get two boxes, one active and one passive, behind a floating IP. CIFS and NFS don't care which box is alive, Plex doesn't panic, and containers don't need therapy afterward.

    Trying to replicate that experience inside a Proxmox cluster feels like reverse engineering a magic trick. You can absolutely get shared storage, redundancy, and even decent performance. The hard part is getting all of them at once without turning your setup into a science experiment.

    Most people come into this thinking, "I already have HA storage, so I'll just layer network services on top of it." That's where reality starts pushing back.

    Ceph: the answer everyone gives, even when it hurts

    Ceph comes up immediately in these conversations, for good reason. It's native and integrated, and Proxmox practically nudges you toward it every time you click "Datacenter."

    Block storage with RBD is solid, and HA VM disks work great. CephFS for shared access is where things get interesting. CephFS can give you a single shared filesystem that multiple nodes mount at once, which sounds like exactly what you want for media, surveillance footage, and bulk storage workloads. In theory, you point your containers at CephFS and call it a day.

    In practice, CephFS inherits all of Ceph's baggage. Small clusters struggle, and two-node pools are a durability nightmare. Running size=2 replicas might look fine until something corrupts and Ceph has no idea which copy is "correct." HDD-backed pools feel slow unless you throw more spindles and faster networking at the problem, and recovery traffic can stomp all over your normal workloads unless you designed the network correctly from day one. When Ceph fails, it fails loudly and with opinions.

    That's why experienced admins keep saying the same thing: Ceph works best when every node participates, the failure domain is hosts, and the network is fast enough to absorb rebuilds without flinching. With anything less you're accepting risk, whether you admit it or not.

    StarWind VSAN: great block storage, awkward file services

    StarWind VSAN sits in an interesting spot. As a replicated iSCSI solution, it does exactly what it promises: high availability block devices, predictable behavior, and fewer moving parts than Ceph. For VM disks, it's clean and reliable. Bulk file access is where things get complicated.

    The moment you say "I want Samba or NFS with a floating IP," you've stepped outside StarWind's comfort zone. Now you need something on top of that replicated block device that understands clustering, fencing, and failover.

    Mounting the same LVM-backed volume group from multiple nodes is a hard no, because that's how you corrupt data fast. So you start looking at clustered filesystems like GFS2 or OCFS2, and suddenly your simple storage question turns into a thesis defense.

    Clustered filesystems work, but they aren't forgiving. Locking behavior and latency both matter, containers don't always behave nicely with them, and troubleshooting becomes an exercise in reading man pages written by people who assume you already know the answers. StarWind gives you strong primitives, but it doesn't give you a Synology-style experience out of the box.

    The "just pass it to containers" trap

    A common instinct is to avoid VMs entirely and pass storage straight into containers. It sounds elegant, with less overhead, fewer layers and more control. That works until you remember that containers don't magically solve shared-write problems.

    If multiple nodes might run the same container, or restart it during failover, they all need consistent access to the same data. That means either a clustered filesystem or a single active node at any given time.

    Once you introduce "single active," you're basically rebuilding HA from scratch. You need leader election and IP failover, and you have to make sure the storage is mounted in exactly one place at a time. At that point, you're halfway to running a NAS VM anyway.

    Why "active-active NAS" is rarer than you think

    People often ask why more open solutions don't just do what NetApp or enterprise arrays do. The answer is simple: it's hard, and it's expensive. True active-active NAS requires tight coordination, rock-solid locking, and decades of edge-case handling. NetApp charges what it charges because that polish didn't happen overnight.

    Open-source projects tend to pick a lane. Ceph focuses on distributed storage primitives, TrueNAS on being a NAS, and Proxmox on virtualization. None of them fully replaces the others without tradeoffs.

    You can run TrueNAS on top of HA storage, float IPs using keepalived or Corosync, and glue it all together. Nobody pretends that's simple, though.

    Performance expectations vs reality

    A lot of these debates come down to expectations. Wanting 200 MB/s read and write speeds isn't outrageous. What trips people up is how that performance behaves during a failure.

    Two-node replication means every write goes to both nodes. Lose one node and you're instantly degraded. Rebuilds hammer the remaining disk, latency spikes, and services stutter.

    It's still usable, as long as you're honest about what "high availability" looks like at home. Sometimes it means "it comes back quickly" rather than "no one notices."

    Backups matter more than perfect HA

    One of the most grounded takes in these discussions is also the least exciting: focus on backups.

    Running Proxmox Backup Server, replicating to a NAS, and keeping local short-term backups alongside off-node long-term ones saves more homelabs than any exotic HA design. HA keeps things running, while backups save your data. Confusing the two is how people end up angry at Ceph for doing exactly what it was designed to do.

    So what's the least bad answer?

    The uncomfortable part is that there isn't a single answer.

    If you want tight Proxmox integration and can scale properly, Ceph remains the cleanest solution, especially if you accept its rules and design around them.

    If you already trust StarWind VSAN, treat it as block storage and layer services carefully, knowing you're building something custom.

    If you want Synology-level polish, there's no shame in letting a NAS be a NAS and focusing Proxmox on compute.

    The middle ground exists, but it isn't smooth, and hardly anyone warns you about that part.