Ceph, HA, and the Minimum Viable Cluster for SMBs
If you're a small or medium business eyeing high availability (HA) with Proxmox and Ceph, the obvious question hits early: what's the smallest cluster setup that actually makes sense? Can you get by with just two nodes and a spare Raspberry Pi pretending to be a quorum device, or is that a crash waiting to happen?
That question kicked off a fiery discussion among infrastructure enthusiasts, and the answers had a lot more nuance than you might expect. It turns out "minimum" means "smallest you can sleep soundly with," which is a stricter bar than "smallest that technically works."
The HA illusion of the 2-node cluster
On paper, you can set up a Proxmox HA cluster with just two nodes and a QDevice for quorum, and the config is technically viable. A lot of users do it, especially in labs or very small deployments, and some even run production VMs this way. There's a reason many in the community get nervous when they see "2-node HA cluster" and "production" in the same sentence, though.
As one user bluntly put it: "3 is the bare minimum, but I'd never run a production workload on just 3."
A 2-node cluster with a QDevice is inherently brittle. The extra vote avoids split-brain, but you're still running on a knife's edge. When one node goes down, the survivor is already at full capacity and may struggle to handle the load. And the QDevice is often something janky, like a NAS, a VM, or a Pi. It does no heavy lifting, but if it fails, you've got quorum issues.
In real-world terms, you want your HA setup to get through a failure without drama, which asks more than simply surviving it.
What Ceph actually needs to breathe
If you're planning to run Ceph for storage in a Proxmox cluster, things get heavier fast. Ceph is resilient, scalable, and performs well under pressure, but it doesn't like to be cramped.
The consensus is that you need at least four real Ceph nodes to run it safely, and most people recommend five before they start feeling secure. That number protects replication, performance, and durability when things get messy.
One experienced admin explained it like this: "Four real Ceph nodes and a quorum vote Proxmox VM, and a very strong network backbone, no joke, like 40Gbit/100Gbit."
That last bit isn't optional. A Ceph cluster on a 10Gbit network is doable, but it caps performance. At the very least you need a tightly designed 10G infrastructure, or clever use of DACs and directly connected NICs to cut switch costs and latency, and even that gets dicey as your node count rises.
Why 3 nodes is the community sweet spot (with caveats)
Ask ten people on a Proxmox forum what the minimum cluster size should be, and the most common answer is "Three real nodes."
That's where quorum gets stable without relying on external gadgets. It also gives you a spare box for updates, testing, or failover without hitting panic mode every time you reboot something, and it lets you run Ceph at a minimum level, although you're still toeing the line on redundancy.
One user noted they use a fourth node without a quorum vote purely for compiling binaries and testing patches, just to avoid risking the cluster. Real-world deployments get that careful even at small scale.
And while QDevices can help in 2-node scenarios, several users argued that managing two full Proxmox nodes plus a QDevice (on a backup server, Pi, or embedded switch container) is more complex than just adding a third node. As another contributor put it: "Two plus QDevice is harder to administer than just having three nodes."
Storage matters: Ceph vs alternatives for SMBs
Ceph isn't the only game in town, especially for lightweight VM workloads. Some admins recommend ZFS replication across nodes for HA, with no need for shared block storage. Others go with iSCSI SANs or newer tech like LINSTOR/DRBD or SeaweedFS, though these come with their own quirks and integration headaches.
Ceph's appeal is that it's deeply integrated with Proxmox and scales nicely, but you pay for that in both complexity and hardware demand.
One user described skipping Ceph altogether in favor of ZFS replication with just two nodes: "We run a 2-node HA cluster with ZFS replication at work using the two_node corosync option and it works fine."
It's not fancy, but for SMBs that don't need hyper-resilient object storage and don't have hundreds of VMs chugging away, it's enough.
Is it production or just a lab in disguise?
One question kept bubbling up in the thread: what actually counts as production? One user answered it best: "Production is defined by the workloads and how you treat them rather than by size."
If you're running a mission-critical service, whether it's a 911 call center or your company's core ERP, on a 2-node cluster with some shaky shared storage, that's production. That setup might also be one hiccup away from a long, painful day.
If you're calling it production, treat it like production. That means redundancy, testing, monitoring, backups, and maybe spending on that third or fourth node instead of getting clever with QDevices and edge cases.
Final thoughts: build small, but smart
If you're an SMB trying to do Proxmox HA and Ceph on a budget, there are two roads.
You can start lean but solid, with three Proxmox nodes that have enough juice to handle VM failover, paired with ZFS replication and good backups, adding Ceph only if you need distributed storage.
Or you can go all in with four or five nodes, a proper Ceph setup, high-speed networking, and a clean HA setup with real quorum. It's pricey, but future-proof.
Forcing high availability into a 2-node setup is often as much about avoiding upfront complexity as saving money. But if that "simpler" setup makes your life harder the minute something fails, is it really simpler?
The smallest cluster that's "worth it" is the smallest one that keeps you sleeping at night, which can be bigger than the smallest one that works. In most real-world cases, that starts at three.