Mr.PlanB Logo

    Newsletter

    Subscribe our newsletter

    Get new infrastructure guides, comparison reports, and migration notes in your inbox.

    Infrastructure notes, guides, and new tools. Unsubscribe anytime.

    Back to Blog
    Kubernetes
    DevOps
    Infrastructure

    What Actually Load Balances Production Traffic in Kubernetes

    February 18, 2026
    7 min read

    Every network engineer hits a particular moment when diving into Kubernetes.

    You've spent years tuning hardware load balancers. You know the rhythm: client → load balancer → servers. TLS termination? Easy. Sticky sessions? Done. 200,000 concurrent TCP sessions? Bring it on.

    Then Kubernetes shows up and the vocabulary explodes into Services, Ingress, API Gateways, cloud load balancers, service meshes, NodePort, the LoadBalancer type, MetalLB, and Cilium. It feels like someone took a perfectly clean diagram and spilled abstractions all over it.

    So what's actually happening in production?

    Kubernetes doesn't replace your load balancer

    One of the most upvoted responses in the discussion cuts straight to it:

    “No, it just controls load balancers… Kubernetes does not come with a load balancer.”

    That sums it up. Kubernetes is not your F5. It isn't magically terminating TLS out of thin air, and it isn't secretly managing 200k TCP sessions inside a black box. It orchestrates and configures, and it tells something else to do the heavy lifting.

    That "something else" depends entirely on your environment. In cloud setups, a Service of type LoadBalancer triggers the cloud controller to provision an external load balancer automatically. In on-prem or bare metal environments, you might use MetalLB to advertise service IPs over BGP or ARP, or let Cilium handle BGP and L2/L3 distribution.

    Kubernetes itself is the control plane, and the traffic workhorse lives elsewhere. Once you see it that way, the rest of the vocabulary starts to fit.

    Ingress is not a load balancer (technically)

    This tripped up a few people in the thread, and it's worth slowing down for.

    An Ingress resource is a rule set, a configuration object. The real muscle sits in the Ingress controller running inside the cluster, which might be NGINX, Traefik, or HAProxy.

    One engineer explained it plainly: an Ingress resource for ingress-nginx gets translated into an nginx.conf file, and that config then handles routing. So Kubernetes defines the intent, the controller translates it into a real proxy configuration, and the proxy does the actual balancing.

    It's abstraction layered on abstraction, and it sounds messy at first. Once you realize it's just config automation, it clicks.

    So where does TLS actually terminate?

    This is where production setups start to diverge. There's no single "correct" answer, only a few common patterns.

    Pattern 1: Terminate at the edge (most common)

    An external hardware or cloud load balancer terminates TLS. The reasons:

    • You can run WAF checks at the edge
    • You reduce internal SSL handshakes
    • Certificate management is centralized

    Several engineers in the thread prefer this. Let the big iron (or cloud LB) handle TLS and keep the cluster focused on application routing.

    Some setups even re-encrypt traffic into the cluster: TLS terminates at the edge, a second TLS session runs to the Ingress, and mTLS takes over inside the cluster. You get layered security and a clean separation of responsibility.

    Pattern 2: Terminate at the Ingress controller

    Instead of offloading to the cloud LB, you configure cert management directly in Kubernetes. Annotations on an Ingress resource can trigger certificate provisioning via tools like cert-manager, and the Ingress controller handles TLS termination itself.

    This works well in:

    • Cloud-native setups
    • Smaller clusters
    • Teams that want full control inside Kubernetes

    It simplifies external infrastructure, but now your ingress layer needs to scale accordingly.

    Pattern 3: Offload to a service mesh

    This is where things get spicy. With Istio, TLS can be handled at the gateway or even via sidecar injection into workloads, and the mesh can automatically enforce mTLS between services.

    That's powerful, and it's also operationally heavy. You don't adopt this casually. You adopt it when you need service-to-service encryption, traffic shaping, observability, and policy control at scale.

    What about 200k TCP sessions?

    This part should reassure the network engineers in the room. If you're using a hardware load balancer or cloud LB in front, the performance story is exactly the same as it's always been. Kubernetes doesn't suddenly make TCP weaker. It uses the same Linux networking stack under the hood that many "traditional" load balancers rely on anyway.

    One commenter who had been building HPC clusters long before Kubernetes existed put it bluntly: it's the same networking stack, just automated differently.

    If you're running software-based ingress inside the cluster, scaling becomes horizontal, with more ingress pods, more nodes, and more endpoints. Autoscalers kick in and L2/L3 balancing distributes traffic. At that point it's ordinary capacity planning.

    Kubernetes doesn't count sessions itself. It delegates that responsibility to the infrastructure layer doing the balancing.

    Bare metal gets interesting

    Cloud makes things easy. Bare metal is where creativity shows up.

    Some engineers run MetalLB with BGP, advertising service IPs directly to their routers. One mentioned pairing it with a MikroTik RB5009 handling BGP routes.

    Others run a traditional reverse proxy outside the cluster, such as Traefik or HAProxy, and forward traffic to NodePort services inside Kubernetes.

    And yes, some still integrate actual hardware load balancers, letting Kubernetes dynamically manage configuration while infra teams keep control of the physical devices. That hybrid model is more common than people admit. Infra manages hardware, app teams manage routing rules, and everyone stays in their lane.

    The real production pattern

    Strip away the terminology and most real-world architectures look like this:

    Client → External Load Balancer → Ingress Controller → Service → Pods

    Everything else is implementation detail. Sometimes the external LB is cloud managed, sometimes it's F5 or Citrix, and sometimes it's MetalLB advertising VIPs via BGP.

    Kubernetes sits in the middle, coordinating how traffic should flow once it enters the cluster boundary. Load balancers are still there; what Kubernetes changes is that their configuration becomes industrialized.

    Why this feels so different

    In traditional environments, networking teams built and managed the entire traffic flow. In Kubernetes environments, application teams can define routing rules declaratively. That's a cultural shift as much as a technical one.

    You're no longer hand-configuring virtual servers on a hardware device. You're writing YAML, and that YAML triggers automation that configures something else. The indirection can feel uncomfortable if you're used to seeing every knob directly.

    It's also powerful, because now:

    • Developers can ship services without ticketing the network team
    • Infrastructure can enforce guardrails
    • Scaling becomes API-driven

    Most importantly, misconfigurations are reproducible and version-controlled.

    Does Kubernetes simplify or complicate load balancing?

    Both. It simplifies provisioning and complicates the mental model.

    Instead of one box labeled "Load Balancer," you now have:

    • Service objects
    • Ingress resources
    • Ingress controllers
    • Cloud controller managers
    • Optional service meshes
    • Optional in-cluster L2/L3 load balancers

    Under all of it, the same principles apply: Layer 4 vs Layer 7, TLS termination points, session handling, health checks, and routing logic. What differs is who manages what, and how automated it is.

    What production Kubernetes networking comes down to

    Kubernetes doesn't magically replace traditional load balancers. It standardizes how you talk to them.

    If you need massive throughput and high TCP concurrency, you still rely on proven infrastructure, whether that's cloud-native LBs, hardware appliances, or high-performance software proxies. Kubernetes just makes sure:

    • They're configured consistently
    • They're provisioned automatically
    • They follow declarative rules
    • Developers can't accidentally break global traffic patterns

    Networking itself stays the same, and Kubernetes orchestrates it. Once you stop expecting it to be the load balancer and start seeing it as the automation brain behind the scenes, everything makes sense. The old flow still exists; it just has a control plane now.