Mr.PlanB Logo

    Newsletter

    Subscribe our newsletter

    Get new infrastructure guides, comparison reports, and migration notes in your inbox.

    Infrastructure notes, guides, and new tools. Unsubscribe anytime.

    Back to Blog
    Kubernetes
    DevOps
    Infrastructure

    Should You Use CPU Limits in Kubernetes Production?

    February 18, 2026
    3 min read

    The original thread asked two things: is the "stop using CPU limits" advice still valid, and do you use CPU limits in production?

    Yes, the mechanics still hold. Whether you should use limits depends on your workload and your cluster philosophy.

    What CPU limits actually do

    In Kubernetes, a CPU request drives scheduling, so that CPU is guaranteed at placement. A CPU limit is a hard cap enforced through CFS quota (the Linux Completely Fair Scheduler). When a container hits its limit it gets throttled.

    Why people say "stop using CPU limits"

    The Robusta article describes real behavior. With limits in place, pods get throttled while the node has idle CPU, latency spikes, and performance becomes unpredictable under burst.

    Kernel mechanics haven't changed, as one responder confirms:

    "Yes, still valid. The described mechanics have not changed."

    For bursty workloads, which covers most web services, limits hurt more than they help.

    The production reality

    The answer is neither "always" nor "never"; it comes down to intent.

    Case 1: burstable web services (most apps)

    APIs, frontends, event consumers, and typical SaaS workloads are mostly idle and spiky, and they're better off borrowing unused CPU. Many clusters run them with CPU requests set and no CPU limits.

    Memory is different. You almost always set memory limits, while CPU is elastic.

    Case 2: CPU-heavy batch jobs

    Data processing, video encoding, ML workloads, and CI/CD runners use CPU constantly, starve their neighbors, and gain nothing from bursting. Here limits earn their keep, and one commenter uses them for specific workloads or CI runners.

    What about QoS?

    Someone asked:

    any real reasons for qos in prod?

    QoS classes matter when a node is under pressure. Setting request == limit gives you Guaranteed, a request alone gives you Burstable, and neither gives you BestEffort.

    In most real-world clusters Burstable is perfectly fine, Guaranteed is useful for critical workloads, and BestEffort should be avoided in prod. QoS is still no reason to add limits blindly, since it matters more for memory eviction than for CPU throttling.

    The subtle gotcha: throttling is not fairness

    CPU limits don't "protect" the cluster. They only throttle the container itself and don't redistribute CPU fairly across workloads the way people imagine. At 30% cluster CPU (one commenter's case), limits do nothing except potentially hurt bursts.

    So what do mature teams actually do?

    Common production patterns in 2026 fall into three groups.

    Pattern A: modern SaaS

    This setup is extremely common now:

    • CPU requests set
    • No CPU limits
    • Memory requests + limits set
    • HPA on CPU or metrics

    Pattern B: mixed workloads

    Here limits are used selectively. Web workloads run without CPU limits, while CPU limits apply to:

    • Batch jobs
    • Data processing
    • CI runners

    Pattern C: multi-tenant clusters

    If multiple teams share a cluster and trust is low, CPU limits act as guardrails, often combined with ResourceQuota. That choice is driven by governance more than by performance.

    My direct answer

    Is the advice still valid? Yes. Do I use CPU limits? Not for bursty production services, but yes for heavy, long-running CPU jobs. I always set memory limits, and I always set CPU requests.

    Where teams go wrong is putting limits on everything without understanding throttling.