Prometheus Memory Usage: How Dashboards Drove Label Cardinality
Prometheus doesn't wake up one morning and decide to eat your RAM. When Prometheus memory usage jumps from "healthy" to "why is this pod requesting 64GB?", something usually earned it.
In the story behind "Prometheus: How We Slashed Memory Usage," the culprit was dashboards: high-cardinality metrics and labels powering Grafana panels that nobody had seriously audited in a long time. Traffic, scale and exotic TSDB bugs had little to do with it.
The fix involved no retention flags or obscure storage settings. The team investigated what was actually sitting inside the head block, and why it was there.
The moment you realize it's not "just growth"
Every Prometheus operator knows the feeling. Memory creeps up slowly, and you assume it's organic growth from more services, more pods and more metrics. When usage climbs faster than the workload grows, something is off.
That is where this story begins. Instead of throwing hardware at the problem, the team dug into cardinality and treated memory usage as a symptom to trace back to its cause. What they found is painfully familiar.
High cardinality: the silent multiplier
Prometheus memory depends heavily on the number of active time series. More unique label combinations mean more time series, and more time series mean more memory. The relationship is brutally linear.
The team focused on high-cardinality metrics and labels, especially the ones Grafana dashboards used. That detail matters, because dashboards can make bad metric design look justified. Someone adds a panel that slices by user_id or request_path. It looks useful and answers a debugging question, and nobody asks whether that label belongs in a production metric. Weeks pass, deployments multiply and pods churn, and the "temporary" label grows into tens of thousands of distinct series. Prometheus dutifully stores every one of them.
Grafana isn't the villain, but it enables bad habits
Grafana didn't cause the problem, but dashboards create incentives. When a panel needs a certain label to work, engineers are reluctant to remove that label even if it is inflating cardinality.
The article describes analyzing which metrics and labels the dashboards actually used, which is a subtle and useful change of question. Instead of listing which metrics exist or which labels are large, you ask which labels are justified by real usage. If a label isn't meaningfully powering a dashboard or an alert, it has no reason to sit in memory.
PromQL as a scalpel
The article promises helpful PromQL queries for finding high-cardinality metrics, and that is the step most teams skip. They feel the pain, assume they know the cause, and start deleting exporters or adjusting scrape intervals. PromQL can show you:
- Which metrics have the most series.
- Which labels have the highest distinct value counts.
- Which combinations are exploding.
When you query Prometheus about itself, you stop guessing and start measuring what each label costs. That is where discipline begins.
From "collect everything" to "collect intentionally"
A lot of Kubernetes-based setups start with broad defaults: kube-state-metrics, cAdvisor, node exporter, ingress controllers and service meshes, all scraping and exporting with rich labels. That works until it doesn't.
This team kept Prometheus's features and pruned metrics that were technically available but practically unnecessary, which takes a change of mindset. Observability culture often says to collect now and analyze later. Prometheus punishes that approach: it won't compress your indecision, and it keeps every series hot in memory.
Dashboards as technical debt
Dashboards accumulate like code. Someone creates one during an incident, another team clones it, and a third adds a new label for more granularity. Soon dozens of panels depend on subtle label combinations nobody remembers justifying, and removing a label feels risky. What if that dashboard breaks, or someone needs that breakdown during an outage? So the label stays and memory climbs.
By tracing cardinality back to dashboard usage, this team changed the question from "Can we afford to remove this label?" to "Is this label earning its cost?"
Slashing memory isn't about flags
When Prometheus gets heavy, it's tempting to look for runtime tweaks: adjust retention, tune compaction, increase the scrape interval or enable compression flags. Those can help, but they don't fix cardinality, which is structural. You can't GC your way out of bad metric design. This team reduced the series count by eliminating unnecessary high-cardinality labels, which is an architecture change more than a config change.
The hidden cost of "just one more label"
Engineers love labels because they make metrics flexible, slicing easy and dashboards powerful. Every new label dimension also multiplies the potential series. Take a metric with:
- 10 services
- 5 status codes
- 3 regions
That's 150 possible series. Add user_id with 10,000 possible values and you're at 1.5 million, which is how memory disappears in practice.
The worst part is that most dashboards don't need that breakdown. It exists because it might be useful someday, and Prometheus stores whatever you ask it to store.
What this story really teaches
The article frames the result as slashing memory usage, and the broader lesson is that observability systems need governance. Without it, metrics sprawl the same way logs do: dashboards grow unchecked, labels multiply and memory follows. Fixing it doesn't take exotic tooling. It takes visibility into series counts, a willingness to question dashboard assumptions, and the discipline to remove labels that aren't pulling their weight.
Prometheus scales if you respect it
These memory stories share a pattern. Prometheus usually fails because it is permissive: it will happily store millions of series and index every dynamic label, and it won't stop you. That is the trap.
This team looked inward and audited its own metric hygiene. They didn't blame the TSDB or immediately reach for a different time-series database. They asked what they were storing and why, and that question alone is worth asking of any Prometheus instance.
If your Prometheus is getting heavy
Ask yourself:
- Which metrics have the highest series count?
- Which labels explode in value cardinality?
- Are those labels actually required by dashboards or alerts?
- Are dashboards driving metric design instead of the other way around?
- Are you exporting labels that change every deploy?
If you don't know the answers, your memory graph probably does, and it's climbing. Slashing memory usage comes down to subtraction, and sometimes the fastest way to make Prometheus lighter is to admit you're measuring more than you need.