Mr.PlanB Logo

    Newsletter

    Subscribe our newsletter

    Get new infrastructure guides, comparison reports, and migration notes in your inbox.

    Infrastructure notes, guides, and new tools. Unsubscribe anytime.

    Back to Blog
    Kubernetes
    DevOps
    Infrastructure

    Real Kubernetes Production Failures From Control Plane to IPs

    February 18, 2026
    6 min read

    Someone asked a simple question: what actually goes wrong with Kubernetes in production? They wanted real failures, the kind that never make it into best-practice guides or whiteboard architecture.

    The answers were unpolished accounts from people who had been burned. Read together, they show a pattern: most Kubernetes failures trace back to very human decisions. Here is what actually broke.

    1. When the control plane implodes

    One engineer accidentally added about 60 machines to the API server pool instead of the node pool. etcd "got REALLY angry and collapsed under its own weight."

    That is operational blast radius more than a bug. The strange part is that workloads kept running. They didn't reschedule, heal or scale, but they kept serving traffic in their last known state.

    Most people don't understand this until they see it live. Kubernetes tolerates this kind of split in a weird way: the data plane keeps running even while the control plane is on fire, at least for a while. Recovery meant manually restoring etcd from its own data directory and rejoining the members, which is nobody's idea of a fun Tuesday.

    2. IP address exhaustion

    "Subnet size for k8s created too small." That one sentence hides weeks of pain. Clusters run out of IPs, pods fail to schedule, nodes scale but can't attach ENIs, firewall rules need rewriting and networking teams get dragged into meetings.

    Another comment bluntly said:

    Not using IPv6 is the first mistake.

    It may have been sarcastic, but the pain behind it is real. Someone else said disabling warm ENI allocation in EKS freed thousands of IPs, the kind of lesson you only learn after watching nodes fail to scale during an incident. IP math is dull until production goes read-only.

    3. DockerHub rate limits: the self-inflicted DDoS

    "DockerHub rate limits are a major chicken and egg." The sequence goes like this:

    • You scale nodes.
    • Nodes pull images.
    • DockerHub throttles you.
    • Pods fail.
    • The autoscaler adds more nodes.
    • They also fail.

    You end up DDoSing your own supply chain. One engineer described DDoSing their internal container registry during a node pool rollout. Kubernetes amplifies mistakes, because rollouts are multiplicative events. If you haven't set up a private registry, a pull-through cache, an ECR/ACR/GCR mirror, Harbor or an image swapper, you will eventually learn this the hard way.

    4. etcd is small, until it isn't

    etcd rarely shows up in architecture diagrams as the villain, yet it is the heart of the cluster. When it is slow, the control plane feels slow, scheduling is delayed and API calls hang. When it is overloaded, you are restoring from snapshots and praying. Most teams don't monitor etcd closely until after their first outage, and that theme runs through the whole thread.

    5. Capacity planning isn't optional

    Someone casually admitted:

    Bad capacity planning from our side.

    That's common. Clusters are built for today's scale, and six months later there are more namespaces, services and pods, more IP usage and more control plane load. As Kubernetes approaches its limits it degrades gradually, and gradual degradation is harder to diagnose than an explosion.

    6. The registry lesson: cache everything

    One response mentioned adding image caching to ECR after getting burned, and another said installing Harbor helped massively. The pattern is clear: at scale, external dependencies become internal single points of failure. Your cluster can be healthy while your registry isn't, and Kubernetes doesn't care whose fault it is. It just reports "ImagePullBackOff."

    7. IPv4 assumptions

    There's a reason someone said:

    If you're deploying a cluster and you think “a /16 won’t be enough” then yes IPv6

    Most teams assume a /24 is enough, a /22 is generous and a /16 is massive. Then pod-per-node density increases, secondary ENIs allocate, warm pools reserve IPs and sidecars double the pod count. Networking is where cloud-native optimism meets physical limits.

    8. The cause is usually not Kubernetes

    Read all of this together and very few of the horror stories involve Kubernetes bugs. They involve misconfiguration, over-scaled control planes, undersized subnets, registry dependencies and capacity blind spots. Kubernetes mostly did what it was told, and the humans told it the wrong thing.

    9. Observability gaps

    The original question asked about observability gaps, and most of the failures described had little to do with applications. They were networking constraints, control plane collapse, infrastructure bottlenecks and registry throttling. None of that shows up in your APM dashboard. It shows up in etcd metrics, cloud subnet utilization, ENI allocations, image pull latency and API server saturation, so if you only watch pod CPU and memory, you won't see it coming.

    10. The lesson nobody likes

    Kubernetes is powerful, and that power multiplies blast radius. At scale, a small mistake becomes 60 misconfigured API servers, thousands of exhausted IPs or a registry meltdown. The system itself is deterministic, but its outcomes at scale are hard to predict.

    What actually goes wrong?

    Distilled, the thread comes down to six causes:

    1. Control plane overload
    2. IP/subnet exhaustion
    3. Registry bottlenecks
    4. Capacity miscalculations
    5. Networking assumptions
    6. Overconfident rollouts

    YAML indentation and container crashes barely came up. Most of it was infrastructure math.

    Why production clusters fail

    Kubernetes in production rarely fails because of an exotic zero-day exploit or a scheduler bug. It fails because:

    • Someone assumed the cluster wouldn't grow that fast.
    • Someone underestimated IP math.
    • Someone scaled node pools without caching images.
    • Someone added machines to the wrong pool.
    • Someone didn't model worst-case rollouts.

    The cluster amplified those decisions. Maybe that's why one commenter called these stories "horror short stories for K8s folks." Once you have watched one of these incidents unfold live, a simple kubectl apply never looks quite the same again.