Why Tesla Picked ClickHouse Over Thanos for Prometheus at Scale
Every time a big company moves off a well-known stack, the internet reacts the same way: "Wait… why didn't they just scale it?"
That's exactly what happened when people saw Tesla talk about moving metrics workloads to ClickHouse. The claim floating around was that "Prometheus doesn't scale horizontally," and the obvious counter-question showed up immediately: why not just use Thanos or Cortex?
It's a fair question, and the answers are more interesting than the headline.
First, Prometheus does scale, with design choices
One myth needs to go right away. Prometheus absolutely scales, just not in the same way a distributed SQL database does.
Out of the box, a single Prometheus server is vertically scalable. You add RAM and CPU, tune retention, and control cardinality. At serious scale, you introduce horizontal patterns:
- Sharding by functional domain
- Federation
- Thanos
- Cortex
- Mimir
As one experienced operator bluntly put it:
Prometheus can handle 100 million cardinality. Thanos can handle billions.
That's real infrastructure, well past hobby scale. So the idea that "Prometheus doesn't scale horizontally" isn't quite accurate. It just doesn't scale the same way a distributed OLAP database does.
The Thanos argument: elastic and object-storage native
One of the strongest replies in the discussion came from someone running both Thanos and ClickHouse and choosing each for different signals. Their setup:
- Thanos for metrics
- ClickHouse for logs, traces, errors
That split is telling, because metrics workloads and log workloads are fundamentally different beasts.
They also point out that ClickHouse scaling horizontally is arguably harder in some ways because each shard has local persistent disk that needs care, and changing shard count is painful. That's a very real operational burden.
With Thanos, the picture looks different:
- Query components can scale horizontally.
- Store nodes scale with object storage.
- Deployment and StatefulSets make scaling straightforward.
- Data is sharded naturally via S3 object storage.
In other words, scaling Thanos is mostly orchestration and object storage, while scaling ClickHouse is shard math and disk planning. They're different trade-offs.
So why would Tesla choose ClickHouse?
This is where it gets interesting. The decision probably had less to do with "Prometheus can't scale" and more to do with what kind of scale they needed.
ClickHouse is an analytical database. It's optimized for:
- Massive columnar datasets
- Complex aggregations
- Ad-hoc queries
- Long retention
- Multi-signal analysis
Prometheus (even with Thanos) is optimized for:
- Time-series metrics
- Operational queries
- Real-time alerting
- High-ingest telemetry
- PromQL semantics
Those goals overlap, but they're not identical. If Tesla wanted any of the following, ClickHouse starts to look attractive:
- Cross-analysis of metrics with other telemetry
- Deep historical analytics
- Arbitrary SQL-style aggregations
- Very long retention at massive scale
The catch: they had to rebuild PromQL
One of the most interesting details in the discussion is that Tesla reportedly introduced their own transpiler (Comet) from PromQL to SQL, in cooperation with the ClickHouse team. That detail changes the tone.
If you move away from Prometheus storage but still want PromQL semantics, you have two options:
- Abandon PromQL.
- Recreate it.
They chose to recreate it, which suggests PromQL is hard to replace. Even if you change storage engines, the query model still matters. If ClickHouse were a drop-in Prometheus replacement, there'd be no need for a transpiler layer, so the fact that one was built tells you this wasn't trivial.
High cardinality isn't "solved" anywhere
Another thread in the discussion hit a core issue: what about high cardinality? Isn't that still a problem in Thanos or Cortex?
Correct. Switching backends doesn't magically solve high cardinality, because it's a data modeling problem. If your metric design includes the following, you will pay for it somewhere:
- Unbounded labels
- Per-user dimensions
- Dynamic identifiers
- Ephemeral workloads
Whether you run Prometheus, Thanos, or ClickHouse, the cost just moves around. Prometheus pays in memory, Thanos pays in object storage and index size, and ClickHouse pays in shard pressure and query planning. Cardinality doesn't disappear; it changes shape.
The operational culture angle
There's another layer here that doesn't get discussed enough. Thanos feels like "Prometheus, but bigger." ClickHouse feels like "we're running a distributed analytical database." Those are culturally different moves.
Thanos scaling looks like:
- Add store nodes.
- Add query nodes.
- Let object storage handle blocks.
ClickHouse scaling looks like:
- Plan shard counts.
- Balance partitions.
- Manage replication.
- Think about disk locality.
- Tune distributed SQL.
One commenter even said changing shard count in ClickHouse is painful, and that's a serious complaint. So Tesla's move was probably driven by capability rather than ease.
Metrics vs multi-signal data warehouse
Another possibility is that they wanted one system for more than metrics. ClickHouse is frequently used for:
- Logs
- Traces
- Events
- Business analytics
If you centralize telemetry into one analytical engine, you reduce system sprawl. You trade the operational simplicity of Prometheus + Thanos for the analytical flexibility of ClickHouse + SQL + custom tooling. Neither decision is simply "better"; they optimize for different targets.
What was their workload?
One commenter asked bluntly: what is high to you? That's the right question.
If Tesla was operating at the following, then a columnar analytics database starts to make sense:
- Hundreds of millions of active time series
- Massive long-term retention requirements
- Cross-domain analytics workloads
- Out-of-order ingestion windows
- Complex query workloads beyond PromQL's strengths
That doesn't invalidate Thanos. It just suggests their use case might have extended beyond "metrics monitoring."
The weird choice comment
One of the most candid replies called the move "a very weird choice indeed". That reaction says something, because for many Prometheus operators, Thanos is the natural scaling story: object storage, a stateless query layer, easy component scaling, and well-understood PromQL.
Moving to ClickHouse looks like abandoning that ecosystem. If you zoom out, though, it might have been convergence instead, with metrics becoming part of a broader analytics fabric.
What this actually teaches
The debate says more about architectural philosophy than about Tesla.
If you believe metrics are:
- Operational signals
- Time-series-first
- Alert-driven
- PromQL-native
Then Thanos (or Cortex/Mimir) feels like the cleanest scaling story.
If you believe metrics are:
- Just another data source
- To be queried alongside logs and events
- Part of a unified analytical lake
- Best explored via SQL and distributed compute
Then ClickHouse makes sense, but you pay for that flexibility in complexity.
Prometheus didn't "lose"
The most interesting detail remains the PromQL transpiler. Even after moving storage, they preserved PromQL semantics, which tells you Prometheus didn't fail conceptually. Its data model and query language were still valuable enough to emulate. The storage moved and the semantics stayed, which looks like evolution to me.
The takeaway
"Why not just scale Thanos?" is the right question. A better one is what problem they were actually trying to solve.
Scaling metrics horizontally is one problem. Unifying massive telemetry datasets under a powerful analytical engine is another. Prometheus plus Thanos can handle astonishing scale, with billions of series in some setups. ClickHouse can handle enormous analytical workloads but demands a different operational mindset.
The choice comes down to capability and to what kind of system you want to run. Once you understand that, the move looks intentional.