Mr.PlanB Logo

    Newsletter

    Subscribe our newsletter

    Get new infrastructure guides, comparison reports, and migration notes in your inbox.

    Infrastructure notes, guides, and new tools. Unsubscribe anytime.

    Back to Blog
    Storage
    Ceph
    NVMe

    Scaling a 40GB/s Ceph Cluster From Five Nodes to Six

    May 21, 2026
    7 min read

    Storage engineers get excited about things nobody outside the room notices, like latency dropping or a dashboard turning green across every node. You spend weeks planning, tuning and adjusting workloads, watching the graphs climb toward performance that stays just out of reach, and then one day the numbers come in higher than you expected.

    I Built a 200K IOPS Monster—Then Proxmox Turned It Into a 6K Joke

    A five-node storage cluster was already moving at absurd speed: roughly 40 GB/s average reads, peak writes touching 11 GB/s, and more than 2 million IOPS with over 30 clients pushing directly against it. It ran on AMD EPYC systems with 200 Gb networking and Ceph on RBD, tested with direct I/O. These were real numbers from real hardware getting hammered until its limits started to show, not figures from a vendor slide.

    The next question was what happens when you add more. Nothing was broken and capacity wasn't running out. People who build systems like this simply want to know where the ceiling is.

    Chasing limits is part of the job

    Hardware got faster over the years, networking jumped forward, and NVMe shifted expectations so far that older storage architectures started to look ancient. Speed alone no longer settles much, though. Modern storage systems are judged by how they scale. A cluster that performs beautifully with five nodes means little if adding a sixth sets off rebalancing storms, downtime, latency spikes or unexpected bottlenecks, and anyone who has lived through a bad storage expansion remembers it.

    That history makes people skeptical. One camp treats every scaling claim with suspicion until it holds up under pressure, because vendors love charts and smooth expansion stories while reality sometimes delivers late-night incidents and emergency troubleshooting. Another camp considers scaling a solved problem when the architecture is designed correctly, since distributed storage exists precisely so you can add resources, redistribute workloads and keep going.

    People who have built enough clusters sit somewhere in between. They know success depends less on the technology itself than on implementation details nobody talks about until something breaks: networking choices, failure domains, workload patterns, placement groups, rebalancing strategy and hardware consistency. Those details decide whether scaling feels effortless or catastrophic, which is what made this test interesting. The cluster was expanded and then pushed hard.

    The satisfaction of watching a system survive abuse

    Stressing infrastructure on purpose is oddly satisfying. You throw clients at it, push throughput, hammer writes and watch the latency charts for the first warning sign.

    This cluster was already seriously capable before the expansion. Forty gigabytes per second of average reads is well beyond small-company infrastructure, and two million IOPS changes expectations entirely. At those levels mistakes surface quickly, whether they are weak hardware choices, bad balancing decisions or design shortcuts.

    Instead, the expansion reportedly went almost suspiciously smoothly. A sixth NVMe node joined, Ceph redistributed and rebalanced the data, and the system kept running with zero downtime and very little drama. For people conditioned to expect pain whenever infrastructure grows, that feels strange.

    One anonymous commenter reacted with humor more than admiration, joking that traditional spinning disks would be "crying" trying to keep pace with numbers like these. The joke works because expectations really have changed. Storage performance targets that sounded impossible a few years ago now feel normal in some environments, and hardware moved faster than many people expected.

    Why storage expansion usually hurts more than anyone admits

    People love talking about performance numbers and say much less about operational pain. Scaling storage isn't glamorous. Clusters rebalance, network traffic surges, recovery operations compete with production workloads, performance shifts for a while, and administrators check their dashboards more often than usual. Sometimes an expansion creates problems that only show up weeks later.

    That is why smooth scaling stories matter. When growth stops hurting, teams plan more aggressively, their architectures evolve differently, and expansion stops being a dangerous event that needs elaborate preparation rituals.

    Skepticism is still healthy. Some engineers argue benchmarks only matter when the workload matches reality, since synthetic test environments can hide ugly surprises and large transactional systems behave differently from direct I/O runs. Others answer that testing limits on purpose exposes weaknesses before production finds them. Both positions have a point, and the strongest infrastructure teams weigh benchmark results and operational history together before they trust a system.

    Ceph's biggest promise was never raw speed

    Performance grabs the headlines, but scalability is what keeps distributed storage alive. The interesting detail in this case was less the throughput than how the expansion reportedly happened: a few dashboard actions, a rebalance, and continued operation without downtime.

    That simplicity matters because complexity compounds. Clusters grow, client counts rise and storage demands multiply, and organizations rarely shrink their infrastructure. Technology that feels manageable at small scale sometimes collapses operationally at a larger one. Distributed systems promise a way out of that trap, even if they don't promise perfection: you can grow without rebuilding everything from scratch, without a forklift migration, without a months-long redesign because growth broke an assumption, and without a downtime window to explain.

    When scaling goes right, it is almost invisible. People outside technical operations rarely notice, while everyone on the infrastructure team does.

    The infrastructure arms race isn't slowing down

    Performance expectations have changed dramatically. Five years ago, many environments measured storage success differently. Today AI systems need enormous throughput, virtualization density keeps climbing, analytics pipelines move staggering amounts of data, and modern software stacks assume storage won't be the bottleneck.

    That assumption puts pressure on the people building infrastructure. They are expected to deliver bandwidth, IOPS, consistency, failure tolerance, recovery speed and room to expand all at once, and there is no finish line. Hit 40 GB/s average reads and someone asks about 50. Reach two million IOPS and people immediately wonder what comes next.

    That mindset drives innovation, and it also wears people out. Infrastructure teams are asked to push harder while staying reliable, scale larger while staying manageable, and move faster without causing outages. Stories like this one resonate because they show preparation paying off. Good architecture absorbs growth, and bad architecture fights it.

    Performance means nothing without confidence

    Confidence is hard to measure. Dashboards don't show it and benchmarks only partly capture it. You can see it when teams stop fearing expansion, when scaling stops feeling risky and when systems behave predictably under pressure.

    The storage world already has plenty of stories about painful growth: expansions that went badly, hardware mismatches nobody anticipated, unexpected bottlenecks, migration nightmares, long rebalancing windows and a lot of lost sleep. Those horror stories spread quickly because nearly everyone collects one eventually, so the smoother experiences stand out. One successful scaling event changes how a team approaches the next decision. Growth feels less dangerous, experimentation gets easier and ambitions grow, which matters more than any benchmark screenshot.

    Good infrastructure disappears into the background, and people only notice it when something fails. The best compliment a storage system gets can sound boring: "It just kept running."

    The real test still hasn't happened yet

    Adding the sixth node answered one question and raised another: how much more performance is left? Infrastructure people rarely stop after an expansion. They benchmark again, retune, measure and push harder, because reaching one limit usually reveals the next.

    Now that the cluster has survived the growth, the open questions are whether throughput rose significantly, whether write patterns improved, whether scaling efficiency held, where latency lands and whether client counts can climb even higher. There is always another graph to watch and another ceiling to find. Systems improve, expectations rise and hardware evolves, so the limits keep moving.

    Somewhere in a room full of dashboards, someone is watching the numbers climb past yesterday's and smiling, because finding out that a fast system can go faster is one of the best feelings in this line of work.