Mr.PlanB Logo

    Newsletter

    Subscribe our newsletter

    Get new infrastructure guides, comparison reports, and migration notes in your inbox.

    Infrastructure notes, guides, and new tools. Unsubscribe anytime.

    Back to Blog
    Ceph
    Storage
    API
    Disaster Recovery
    Proxmox
    Debugging

    How One Bad API Call Took Down an Entire Ceph Cluster

    November 9, 2025
    12 min read

    It started with a single command, a simple curl request that was supposed to pull a few harmless stats. Instead, it brought an entire Ceph cluster to its knees.

    A home lab admin, tinkering with Ceph's RESTful API, sent what should've been a routine call:

    curl -k -X POST "https://USERNAME:API_KEY@HOSTNAME:PORT/request?wait=1" -d '{"prefix": "df", "detail": 1}'
    

    That line doesn't look dangerous. It's the kind of one-liner anyone running Ceph has probably used or tested in some form: a quick peek at pool stats, something you could usually pull with ceph df detail. This time, though, that tiny detail parameter, "detail": 1, triggered something far worse than a parsing error. Within seconds, the monitor services (ceph-mon) crashed across the entire cluster.

    Every VM tied to the system froze. The Proxmox dashboard went red, the storage daemons screamed in silence, and the cluster at the heart of a self-hosted infrastructure was dead, all because of one malformed API request.

    The day the monitors died

    At first, it looked like a minor hiccup. The admin, who goes by packetsar, noticed Ceph's monitors throwing errors. Instead of recovering, they crashed and kept crashing. Logs filled with a hauntingly specific message:

    ceph-mon[278661]: terminate called after throwing an instance of 'ceph::common::bad_cmd_get'
    ceph-mon[278661]: what():  bad or missing field 'detail'
    

    Anyone who's wrangled Ceph knows that when the monitors (MONs) fail, you have a big problem. MONs keep track of the cluster map: they know which OSDs hold what data, where pools live, and who's in quorum. When they're offline, the cluster effectively loses its memory.

    And this went well beyond a single node crash. It cascaded across all monitors. Each one picked up the poisoned request from a shared database and followed it right off a cliff, and every attempt to restart just repeated the failure.

    "Reboots, service restarts, nothing worked," wrote packetsar. "Ceph and Proxmox cluster are hard down and VMs have stopped at this point."

    The self-inflicted poison pill

    The weirdest part is that the command shouldn't have been that harmful. Ceph's API is supposed to validate input and gracefully handle malformed JSON or unknown fields. Instead, the "detail": 1 parameter got interpreted in a way that caused a fatal exception inside the MON process.

    Essentially, Ceph ingested the bad request, choked on it, and then saved it, replaying the same invalid command each time the monitors tried to restart. It's like a crash loop cycle: the cluster remembered the bad command and kept feeding it back to itself.

    "Looks like that request got put in a shared ceph-mon database and caused all the monitor services to crash," the admin wrote.

    So one bad command turned into persistent corruption. The poisoned state lived inside the monitor store, meaning even clean restarts couldn't shake it.

    One comment summed it up perfectly:

    "Yet another example of bad input validation? Don't forget to file this as a bug towards Ceph."

    The lesson is simple and harsh: when you build distributed systems, input validation isn't optional, because one unchecked field can bring everything down.

    "You can poison the whole cluster with a single REST call"

    That's how packetsar described it later in the thread. It sounds hyperbolic until you remember Ceph's architecture.

    Ceph is built around consensus. The monitors replicate data to maintain quorum, and when one gets bad state data, it propagates it to others, assuming it's valid. So if a malformed command makes its way into the monitor database, it gets copied faithfully across nodes. The system trusts its peers, which is its strength and, in this case, its weakness.

    When someone asked whether the same command worked fine via CLI, the admin confirmed that ceph df detail ran perfectly from the command line. It was the REST API layer that failed to validate properly. That distinction matters because it puts the problem in Ceph's API wrapper rather than its core logic, and that wrapper is a layer meant to make automation safer and simpler. Here it became a single point of failure.

    Rebuilding from ashes

    Once the scope of the failure sank in, packetsar had to figure out how to bring a dead Ceph cluster back when the monitors wouldn't even start.

    There's no magic "undo" button for Ceph's monitor database, so packetsar had to go old-school and rebuild it manually from the OSDs (Object Storage Daemons). It's a method outlined deep in Ceph documentation, usually reserved for extreme corruption scenarios.

    Here's the recovery process, simplified:

    1. Shut down all OSDs across every host to stop further writes.
    2. Pick one host to start with and rebuild a fresh monitor database using data from its local OSDs.
    3. Rsync that new store.db to other hosts, one by one, rebuilding and merging as you go.
    4. Once the database is complete, replace the production store.db on all monitors with the new one.
    5. Bring up the MONs again and let them reach quorum.
    6. Finally, rebuild the manager (mgr) daemons one at a time and reconfigure settings like the REST API module.

    That's not something you do lightly. It's slow, nerve-wracking work, especially when your storage cluster is the foundation of your entire virtualization setup. But it worked.

    "I was able to get VMs back online about an hour after I started slowly working through the rebuild process," he wrote.

    That's remarkable resilience, and it shows how deeply open-source systems can be understood and repaired when you have full control and patience.

    Lessons from the meltdown

    Nobody should take "stop tinkering with your cluster" from this, since tinkering is how you learn. But distributed systems, especially ones like Ceph, demand a particular respect.

    A few takeaways stand out from this debacle.

    1. APIs aren't always safer

    We tend to think APIs abstract away risk and that using structured calls instead of raw commands makes automation safer. This incident shows the opposite can be true if the API isn't validating inputs properly. Bad input through an API shouldn't take down the service it's meant to expose. That's development 101.

    2. Persistence cuts both ways

    Ceph's design makes it resilient, keeping data consistent across nodes even through crashes. That persistence also makes it unforgiving. If it stores a bad state, that bad state becomes gospel until manually purged.

    3. Testing in production isn't testing

    It's tempting to test "just one small thing" in your live cluster, especially if you've done it a hundred times before. But a self-hosted lab isn't the same as a sandbox, and one malformed command can turn an evening experiment into a 4 a.m. recovery session.

    4. Documentation is survival

    Documentation saved this cluster, and luck had little to do with it. Few people know you can rebuild a monitor store from OSDs. The process is buried in Ceph's lower-level docs, the stuff only desperate admins end up reading, and it's also what brought the cluster back.

    The danger of silent failures

    This wasn't the first Ceph disaster story, and it won't be the last. Others chimed in with eerily similar experiences: clusters refusing to come back after power outages, monitors corrupted beyond repair, or phantom configurations that refused to die.

    One user wrote:

    "Managed to break Ceph by having an unexpected power outage… it never came back up and nothing I tried worked. Even starting from scratch kept bringing back old remnants. I sacked Ceph off and moved to ZFS with replication."

    ZFS isn't distributed storage; it trades scalability for simplicity and reliability, which is telling. When a system's recovery path is harder than a full migration, you should treat that as a warning sign.

    Ceph is powerful but unforgiving. It gives you incredible flexibility, with object, block, and file storage in one platform, but it assumes you'll handle that power responsibly.

    "Public shaming might be the best medicine"

    One commenter half-joked that if the bug was reproducible, "the Ceph devs should be ashamed." And honestly, they have a point.

    This is more than an obscure corner case. An authenticated REST call should never be able to crash core services. Whether you're running a home lab or a data center, that kind of fragility undermines trust in the stack.

    Open-source software thrives on transparency. Bugs happen, but ones like this show the need for better testing around API input validation, particularly for admin-level commands.

    Aftermath and reflection

    The story has a happy ending: the cluster was revived, the VMs came back online, and the admin learned more about Ceph's internals in a few hours than most do in months.

    It also left a mark. "It seems crazy that you can poison the whole cluster with a single REST call," he said. And he's right.

    Distributed systems are supposed to be resilient, but resilience doesn't mean invincibility. Sometimes it means recoverability. This case was far from perfect, but it shows clearly that you can claw your way back when everything goes wrong.

    The bigger picture

    There's a broader theme here about the fragility of automation in modern infrastructure. We build layers upon layers, APIs on top of daemons and daemons on top of databases, all to make management simpler. Each layer also introduces new complexity and new ways to fail.

    In cloud-native environments, that risk multiplies. Every service talks to another over APIs, and every command, deployment script, and automation job carries the potential to trigger a cascade.

    When something like Ceph, one of the most respected open-source storage systems in the world, can be brought down by a single malformed JSON field, it's a reminder that complexity always comes at a cost.

    Epilogue: one line to rule them all

    Somewhere in a terminal history sits that fateful command. It's small and unassuming, a line of text that fits in a tweet.

    Behind it is a story about trust, fragility, and resilience in distributed systems: how even the most sophisticated software can fall apart from a tiny mistake, and how, with the right knowledge, persistence, and documentation, it can be brought back to life again. Sometimes the difference between total data loss and recovery is just knowing which file to rsync next.