How One Team Cut Prometheus Memory From 60GB to 20GB
Most Prometheus operators hit the same wall sooner or later. The UI turns sluggish, scrapes start failing, PromQL queries time out, and dashboards hang just long enough to make you question your life choices. Then you check memory and find Prometheus casually chewing through 50 to 60GB of RAM. Memory at that level usually means your next big incident is already on its way.
One engineer described exactly that situation: a cluster where Prometheus had ballooned to about 60GB, with scrape reliability degrading, queries timing out and stability slipping. The fix they found was plain, ruthless cleanup.
The culprits: three predictable memory traps
The investigation turned up three classic Prometheus memory traps:
- Duplicate scraping
- Histogram overload
- Label explosion
If you have run Prometheus in Kubernetes for long, at least one of these has probably bitten you. Most teams get hit by all three.
1. Duplicate scraping: the accidental 2x multiplier
Prometheus was scraping ingress metrics from both the pods and a ServiceMonitor. That sounds harmless, but every duplicate scrape doubles:
- Time series count
- Memory usage
- WAL pressure
- Head block churn
Unless you deduplicate downstream, you have doubled cardinality for that entire metric set. You won't feel it right away, but it compounds across clusters, and one small config oversight can cost you tens of gigabytes.
The fix here was to disable the unnecessary pod-level scraping.
2. Histogram overload: death by buckets
Histograms are powerful and expensive. Metrics like *_duration_seconds_bucket can generate hundreds of thousands of time series when you have many buckets, many label combinations and many replicas. Multiply buckets by labels, pods and clusters, then by retention in the head block, and your memory graph starts to look like a hockey stick.
Exporters often enable histograms by default, and teams rarely go back to check whether they need every bucket or whether high-cardinality labels are attached to them. Prometheus ends up working as a histogram warehouse.
3. Label explosion: the cardinality monster
This one is the most dangerous. In this cluster, labels like these were producing 10k+ unique values:
replicasetpathcontainer_id
Each unique label combination is its own time series. If path includes raw URLs with embedded IDs, if container_id changes on every deploy, or if replicaset rotates constantly, every rollout creates thousands of new time series, and Prometheus keeps them all in memory. That is how Prometheus is designed to work.
The fix: disciplined cleanup
The remediation steps were straightforward:
- Drop unused metrics (after validating dashboards and alerts)
- Disable redundant scraping
- Remove high-cardinality labels that weren't actually used
- Write scripts to verify what was safe to drop
The last step matters, because blindly dropping labels can cause ingestion errors. A moderator added an important warning: labeldrop removes the label, not the series. If that label was needed to keep series unique, removing it can cause duplicate series collisions and ingestion failures. The series collapse into each other, and collapsing them without thinking leads to chaos. The author updated the post after the correction.
That exchange is a good example of the Prometheus behavior that trips people up. Tuning it takes precision.
The result: from 60GB to 20GB
After the cleanup, memory dropped from roughly 60GB to about 20GB and stability came back. Scrapes normalized, the UI responded again and PromQL stopped timing out, on the same cluster with the same workloads. The 40GB difference was waste that had been removed.
A pattern across teams
The comment section showed several teams dealing with similar issues. Someone joked "Laughs in VictoriaMetrics." Another replied "NetData ;)". The cycle is familiar:
- Prometheus grows.
- Memory balloons.
- People blame the database.
- Alternative TSDB vendors enter the chat.
Most Prometheus memory explosions come from the setup around Prometheus: duplicate scraping, unbounded label cardinality, overzealous histograms and the volume of metrics Kubernetes emits by default. One strong recommendation in the thread was to review the action: drop defaults in kube-prometheus-stack, because Kubernetes apiserver and cAdvisor metrics can overwhelm small setups. Kubernetes emits a lot of metrics, and ingesting all of them blindly is how you end up with a memory crisis.
Prometheus doesn't forgive laziness
Prometheus stores exactly what you tell it to store. It won't compress away bad labeling decisions, protect you from cardinality explosions or warn you when histograms multiply by replicas. It assumes you know what you are doing, which gives you a lot of power and plenty of room to hurt yourself.
When people say "Prometheus doesn't scale," they often mean "we didn't manage cardinality."
If you're sitting at 40GB+ RAM right now
Work through this checklist:
- Are you scraping the same endpoint twice?
- Do you actually need pod-level metrics for every service?
- How many histogram buckets are you exporting?
- Are you attaching user-level or path-level labels to high-volume metrics?
- Are container IDs or replica hashes in your labels?
- Have you audited unused metrics recently?
- Are you reviewing drop rules in kube-prometheus-stack?
If you haven't asked those questions yet, your memory graph is probably next.
Prometheus at 20GB can be healthy
The goal is to run Prometheus intentionally, which doesn't always mean running less of it. An instance using 20GB for a large Kubernetes cluster can be perfectly healthy. An instance using 60GB because it duplicates ingress metrics and tracks every container ID ever created is wasting most of that memory. In observability, this kind of discipline scales better than hardware.