Mr.PlanB Logo

    Newsletter

    Subscribe our newsletter

    Get new infrastructure guides, comparison reports, and migration notes in your inbox.

    Infrastructure notes, guides, and new tools. Unsubscribe anytime.

    Back to Blog
    Prometheus
    Counters
    Datadog
    Metrics

    Why Prometheus Counters Confuse Teams Moving From Datadog

    January 27, 2026
    8 min read

    Some frustration only shows up after a migration.

    You switch from something like Datadog to Prometheus, rebuild the dashboards and rewrite the alerts, and everything feels leaner, more "open," and more under your control. Then the counters start lying to you, or at least it feels that way.

    Low-frequency counters miss increments. Short-lived series disappear from totals. Success rates don't quite add up over 30 days. Alerts don't fire when they should, or worse, they fire late.

    One team described it bluntly: counters have been the biggest pain point since the switch. Things that "just worked" before now need careful thinking, and even when you think you've done it right, you're still uneasy. So is Prometheus unreliable, or are we asking it to behave like something it was never meant to be?

    The core complaint: slow counters plus dynamic labels equals anxiety

    The pain shows up in a few predictable places:

    • Alerting on a counter increase when the counter doesn't start at zero.
    • Calculating total increments over a time range, especially when short-lived series exist.
    • Viewing frequency of increments as a time series without weird artifacts.
    • Computing long-term success rates using sum(rate(success_total[30d])) / sum(rate(overall_total[30d])) and realizing short-lived series skew results.

    There's also a meta frustration: the raw data is there. If you eyeball the graph, you can often see what "should" be counted, yet rate() or increase() seems to understate things.

    That behavior is intentional. Prometheus prefers undercounting over overcounting when it detects edge cases like resets or missed scrapes. For SREs, that safety bias can feel backwards, because in many real-world setups a false negative alert is worse than a false positive. When your monitoring system chooses safety over sensitivity, it can feel like it's betraying you.

    The first hard truth: very slow counters are awkward

    One experienced voice cut straight to it: very slow-moving counters are a difficult issue with Prometheus.

    If you scrape every 30 seconds and something increments once every 20 minutes, you're working with sparse data. Now add dynamic labels, short-lived pods, restarts from deployments, and autoscaling events. You're no longer measuring a clean monotonic series. You're measuring fragments of counters across instances.

    Prometheus handles counter resets, but it cannot reconstruct missing history from pods that disappeared. Nobody is being incompetent there; it's physics.

    The cardinality trap

    One of the most consistent recommendations was to reduce cardinality for important SLO metrics.

    Too many teams add debugging-level labels directly to error counters:

    • user_id
    • request_id
    • feature_flag
    • shard
    • deployment hash

    It feels powerful, and it's also a recipe for sparsity. Sparse series make rate() less reliable because the time window might include series that only existed for a small part of that window. Your 30-day success rate query now includes dozens of micro-series that blinked in and out of existence.

    Metrics are supposed to answer "Is there a problem at X time?" They are not meant to replace your logs. When you overload counters with forensic-level labeling, you're stretching Prometheus beyond its design center.

    Short-lived workers: the silent saboteur

    Another pointed comment called short-lived metrics an anti-pattern.

    If you're using queue-dispatched ephemeral workers or FaaS-style compute, your counters may:

    1. Start.
    2. Increment once.
    3. Vanish.

    Your increase() calculation then has to reconcile series that only existed for a few scrapes, and that's where the weird undercounting shows up.

    One solution mentioned was accumulator exporters. Instead of each short-lived worker exposing its own counter, you push increments to a central accumulator (StatsD-style) and let Prometheus scrape that stable source. In modern stacks, you might use OpenTelemetry cumulative deltas feeding into a single aggregation collector.

    The theme is consistent: stabilize the counter before Prometheus sees it. Prometheus likes long-lived time series and does not love flickering ones.

    "Created timestamp injection" isn't magic

    For counters that don't start at zero, Prometheus now supports created timestamp zero injection via OpenMetrics. That helps with some startup ambiguity, but it doesn't eliminate every edge case.

    If a pod restarts and your scrape interval misses part of the lifecycle, you can still see partial data. Prometheus tries to handle resets intelligently, but it errs on the side of underestimating, by design. If you were expecting "perfect delta reconstruction," you're expecting something Prometheus doesn't promise.

    Recording rules: boring, powerful, underrated

    The Grafana SLO feature approach uses layered recording rules like:

    Codesum(sum_over_time((grafana_slo_success_rate_5m{})[28d:5m])) / sum(sum_over_time((grafana_slo_total_rate_5m{})[28d:5m]))
    

    It feels complicated, but there's a reason for the ceremony. Pre-aggregating into stable 5-minute deltas via recording rules makes long-range SLO math far more reliable.

    One key suggestion was to test your recording rules. Use promtool test rules, and add alerts if recording rules stop evaluating. If you don't test them, you're trusting invisible plumbing; if you do, they're surprisingly solid.

    The irony is that many teams trust raw PromQL more than recording rules, even though tested recording rules are often safer.

    The float problem (yes, it's real)

    Prometheus uses floats, which means counters lose +1 precision at very high magnitudes (around 2^53). Most web workloads will never hit that, but high-speed interfaces or massive aggregations might.

    One clever workaround mentioned was wrapping uint64 counters modulo 2^53 before export. It's niche, but it's a reminder that Prometheus made trade-offs early on and those decisions still ripple outward.

    Metrics vs. logs: the confidence crisis

    Underneath the technical complaints there's a bigger issue. When teams start saying "Maybe we should just use logs for this," they're talking about trust more than math.

    Metrics feel less credible when they miss edge cases. When a 30-day success rate query behaves differently depending on window size, confidence erodes.

    The uncomfortable part is that metrics were never meant to be perfect forensic accounting systems. They are coarse, aggregated signals. Logs are precise, and metrics are scalable. If you demand log-level exactness from counters, you'll always feel let down.

    The most interesting idea: "materialized metrics"

    One long-term proposal floated in the discussion was a new pipeline inside Prometheus that converts scrapes back into deltas, then re-materializes them into projected counters after label reduction. You would drop instance and other ephemeral dimensions, aggregate the deltas first, and then rebuild stable counters.

    It's ambitious, and it acknowledges a central pain: many users think in deltas, not cumulative counters. They want event-like semantics with metric-scale performance. If something like that ships, it could reshape how people think about counters entirely.

    So are counters "very unreliable"?

    No, but they are unforgiving.

    Prometheus counters are extremely reliable when:

    • Series are long-lived.
    • Cardinality is controlled.
    • Scrape intervals match event frequency.
    • Recording rules are used for long-range math.
    • You're measuring system health, not auditing transactions.

    They become uncomfortable when:

    • You mix debugging labels into SLO metrics.
    • You rely on short-lived workers.
    • You expect perfect reconstruction across restarts.
    • You stretch rate() across 30 days of fragmented series.

    The title calling them "very unreliable" was an admitted exaggeration, but the frustration is real. Prometheus isn't broken; Datadog-style mental models just don't transfer cleanly.

    Prometheus forces you to think about scrape intervals, series lifespan, label economics, and aggregation strategy. That cognitive load feels like a regression at first, until you realize it's just a different contract, and like most contracts in distributed systems, the fine print matters.