Why My 3-Node Ceph Cluster Is Hitting a Wall at 25GbE
I spun up what felt like a monster on paper: a three-node Proxmox cluster with four NVMe OSDs per node, all enterprise-grade, plus separate 25Gb NICs and switches for Ceph public and private traffic. It was clean, overbuilt, and fast.
Do You Really Need a 10Gb Network for Proxmox Ceph?
Or at least it was supposed to be, until the benchmark numbers came back: 2756 MB/sec of bandwidth, 689 IOPS, and 23ms average latency. Suddenly that shiny 25Gb network didn’t feel so shiny anymore.
This was a learning lab and a playground, not a production system. Still, when you spend real money on NVMe drives and 25Gb networking, you expect fireworks. Instead I got a ceiling, and the comments started rolling in.
The setup should’ve screamed
Here’s the hardware:
- 3x MS-02 Ultra 285HX nodes
- 64GB DDR5 5600 per node
- 4x Micron 7450 Pro NVMe drives per node
- PCIe Gen4
- Dual 25Gb networking (separate public and cluster networks)
On paper, this thing should fly.
The write test was run with:
Coderados bench -p ultra-pool 20 write -t 16 --object_size=4MB --no-cleanup
The result was consistent at around 2.7 to 2.8 GB/sec. It wasn’t spiky or unstable, just… capped.
One reply summed it up bluntly: You’re hitting the network limit. Another user did the math out loud: almost 3 GBytes/sec is about 24 Gbits, which is basically 25GbE maxed out.
It stings when someone else says it, because deep down you already know.
“Ceph shines at scale,” and that’s the catch
One of the first responses cut through the noise:
It’s hard to tell. But I’d say the small size of the cluster. Ceph shines at scale. 3 nodes is the very bare minimum.
People don’t always want to hear that. Three nodes is the minimum viable cluster, the point where Ceph barely starts being Ceph.
Distributed storage doesn’t flex its muscles until you give it room to breathe. Once you add more OSDs and more nodes and spread the load wider, you start to see real parallelism.
With only three nodes, you’re boxed in. Replication traffic bounces between the same machines, network saturation shows up fast, and there’s nowhere for the data to hide. The benchmark numbers reflect that layout more than any weakness in Ceph.
The IOPS fear
What really scared one commenter was the IOPS, more than the bandwidth: 689 on average.
Someone chimed in saying they get roughly 127 IOPS out of their SAS SSDs, so comparatively this looks better. Let’s be real, though: these are Micron 7450 Pro NVMe drives, and locally they can do absurd numbers.
So why does distributed storage look so… ordinary?
Because a write here goes far beyond a single NVMe talking to a CPU over PCIe. It involves:
- Client write
- Network hop
- Primary OSD write
- Replica OSD write(s)
- Acknowledgment chain
- Journaling
- Commit
Every 4MB object goes through that ceremony of consensus and safety. You don’t get NVMe marketing numbers in a replicated cluster. What you get is durability, fault tolerance, and survival, and there’s a price for that.
Is 25GbE the villain?
Now for the network. 25GbE sounds fast, and it is, until you start pushing multiple 4MB objects across replication streams.
At ~2.8 GB/sec, you’re basically saturating a 25Gb link once overhead is included. You’re already at the limit.
One commenter offered a practical suggestion: LAG or ECMP.
Individual streams won’t exceed 25Gb. But you can run 50Gb per second no problem.
That’s the nuance. A single TCP stream won’t magically exceed 25Gb, but multiple flows across multiple ports change the picture.
ECMP can spread traffic across multiple links. LAG can work too, but it gets weird past two ports on some gear. Suddenly you’re deep in networking architecture on top of benchmarking storage. This is how it happens: you start with a cluster and end up redesigning your switching fabric.
The 4K block size rabbit hole
Then things got more interesting. Someone asked whether I was running a switch or a full mesh network, and then came the suggestion to reformat the NVMe drives to a 4K block size. That’s when you know you’re officially in the weeds.
The Micron 7450 drives might be running 512e instead of native 4Kn. That mismatch can cost performance, especially in distributed storage systems that care deeply about alignment.
Reformatting NVMe drives isn’t casual, though. You don’t click a button. You:
- Remove an OSD
- Wait for rebalance
- Reformat drive
- Re-add as new OSD
- Wait for rebalance
- Repeat
You do it slowly and carefully, one disk at a time, which is closer to surgery than tuning.
Emotionally, when you’re new to Ceph, even asking whether you need to redo the whole cluster feels overwhelming. You don’t want to blow up the lab you just built, but that’s how you learn.
Benchmarks vs reality
Another voice of reason popped up:
The best benchmark is probably your real workload or a fio benchmark which resembles your workload.
That’s the awkward part about storage benchmarks. rados bench is synthetic, clean, and controlled. It doesn’t look like VM traffic, databases, or backups.
You can chase benchmark numbers forever, or you can ask whether this handles what you actually plan to run. If your workload isn’t saturating 25Gb, maybe you’re not bottlenecked in real life, and this “limit” only exists in lab stress tests. That perspective matters.
Is distributed storage really that bad?
At one point in the discussion, someone basically asked whether distributed storage is really that bad.
It isn’t bad. It is expensive, architecturally as well as financially. You’re trading raw local NVMe speed for:
- Redundancy
- Self-healing
- Data distribution
- Node failure tolerance
Ceph is designed to survive a node dying at 2am without you noticing. For a cluster like this, that matters more than any single-drive benchmark number.
The beginner phase is the best phase
One comment stuck with me:
New to all this, but it’s been fun learning so far.
That’s the right energy, because this phase, where you question everything, is where deep understanding forms. Why does Ceph scale with nodes? Why does replication amplify network load? Why does block size alignment matter? Why does ECMP change throughput characteristics?
You don’t really get distributed systems until they disappoint you and you start pulling threads.
So what’s actually happening here?
Zooming out, the numbers look like this:
- ~2.7 GB/sec write bandwidth
- ~24 Gbit/sec effective throughput
- 25Gb network
- 3-node replicated cluster
- 4MB objects
- 16 threads
This is almost textbook network saturation. The drives aren’t the bottleneck, PCIe Gen4 isn’t the bottleneck, and the CPUs (285HX) aren’t gasping. The fabric is the limit, and with only three nodes there’s limited horizontal scaling.
Adding more nodes increases performance because replication spreads wider and aggregate bandwidth grows. Adding more links breaks past the single-link ceiling. Tuning block size and alignment might squeeze out extra efficiency.
Nothing here looks broken; this is what the physics of a saturated link looks like.
What “overkill” labs teach you
There’s something humbling about realizing your “overbuilt” lab is already maxed out. You thought 25Gb was huge, and it turns out to be finite. You thought enterprise NVMe meant absurd cluster numbers, but it doesn’t override replication math. You thought three nodes was serious scale, and it’s the starting line. That’s distributed systems pulling the curtain back.
The good news
The cluster is stable, the bandwidth is consistent, the latency isn’t chaotic, and the system behaves predictably. Predictability is gold in storage.
I found a limit, and limits are where real architecture decisions begin. Do you:
- Add nodes?
- Add network links?
- Switch to 100Gb?
- Tune OSD settings?
- Change replication?
- Accept the ceiling?
Those are design decisions now.
Three nodes, twelve NVMe drives, and 25GbE got me nearly 3 GB/sec writes. That’s the moment you realize distributed storage responds to your topology, whatever you expected from it. And honestly, that’s what makes it addictive.