Prometheus cardinality is the number of unique time series your metrics produce once every label value combination is counted. It’s the most common reason a healthy Prometheus server runs out of memory. One unbounded label, like user_id on a request counter, can turn a few hundred thousand series into millions within days.

This guide shows how to find the metric and label responsible, and how to fix it with instrumentation changes, relabeling, scrape limits, and recording rules.

TL;DR

  • Cardinality is the count of unique label value combinations per metric name. Each combination is a separate time series held in memory.
  • Unbounded labels such as user IDs, request IDs, and raw URLs are the most common cause of explosions. Every new value creates a new series.
  • Prometheus has no built-in series limit, so memory is the real ceiling. Track prometheus_tsdb_head_series and alert on growth against a baseline.
  • Fix it as early as you can: remove the label at instrumentation, drop it with relabeling or scrape limits, or pre-aggregate with recording rules.
  • Middleware can strip high-cardinality labels inside your Kubernetes cluster with OpenTelemetry-native filters, before the data is sent or ingested.

Bring your Prometheus metrics into one view

Middleware scrapes annotated pods or accepts remote write, then puts those metrics on the same timeline as your logs and traces.

What cardinality actually means in Prometheus

Every Prometheus time series is identified by its metric name plus its full set of label key-value pairs. Change one label value and you don’t update an existing series. You create a new one with its own storage, its own chunk, and its own entry in every index that touches it. If you’re new to the data model, start with what Prometheus is and how it stores time series.

That’s the mechanic behind the official Prometheus guidance on labels: every unique combination of label pairs is a new time series. The docs specifically call out user IDs, email addresses, and other unbounded value sets as labels to avoid. Teams pairing Prometheus with Grafana inherit whatever cardinality decisions were made at instrumentation time.

Cardinality compounds multiplicatively, not additively:

Series per metric = values(label 1) × values(label 2) × … × values(label N)

A counter with a method label (4 values), a status_code label (6 values), and an endpoint label (50 values) produces 4 × 6 × 50 = 1,200 series from one metric definition.

Add a customer_id label with 10,000 distinct values, and that single counter now accounts for 12 million series. Nobody edited 12 million lines of config. One label did it.

How a cardinality explosion happens in the TSDB

Prometheus keeps its most recent data, the head block, in memory. Every active series in the head block needs a series descriptor, label index entries, and a chunk buffer. The exact cost per series depends on label count, label length, and churn. Measure it on your own server instead of relying on a rule of thumb:

process_resident_memory_bytes{job="prometheus"} / prometheus_tsdb_head_series

Prometheus has no built-in ceiling on how many series it will accept. It keeps allocating memory for new series until the process is OOM-killed, unless sample_limit in a scrape config rejects the scrape first.

That’s why cardinality problems rarely announce themselves with a clean error. They show up as slowly climbing memory, then sluggish queries, then a crash that looks unrelated to the cause.

Churn makes this worse than a static series count suggests. Every pod restart, autoscaling event, or short-lived batch job retires one set of label values (the old pod name, the old instance IP) and mints a new one. Kubernetes clusters do this constantly, so prometheus_tsdb_head_series can climb even when request volume is flat.

Our breakdown of how Prometheus works end to end covers how the scrape-to-storage path fits into collection, querying, and alerting.

Common cardinality sources and their blast radius

SourceTypical driverSeries impactPrimary fix
Unbounded app labelsuser_id, session_id, request_id on a counter or histogramGrows with unique users or requests, effectively unboundedRemove the label; move the detail to logs or traces
Raw URL pathsUnnormalized path label on HTTP metricsGrows with every dynamic segment (IDs in URLs)Normalize to route templates (/users/{id}) at instrumentation
Kubernetes metadatapod, instance, container_id on every scrapeMultiplies every metric by pod count, and again on every restartDrop or aggregate churn-prone labels with relabeling
Histogram bucketsDefault or custom le buckets on latency histogramsMultiplies base cardinality by bucket countFewer, purpose-fit buckets or native histograms
Federation and remote_writeRaw metrics forwarded upstream without aggregationDuplicates leaf cardinality at the central layerPre-aggregate with recording rules before forwarding

Where cardinality actually comes from

Unbounded label values

This is the classic cause, and it’s almost always accidental. A developer adds user_id or request_id to an existing counter because it seemed useful while debugging one incident. No alert fires, and the metric still looks fine on a dashboard.

Weeks later, after enough unique users have hit the endpoint, that one label has generated millions of series. Filtering it at query time doesn’t help, because the series already exist in the TSDB. The fix is at instrumentation: remove the label and push per-user detail to log monitoring or distributed tracing. Logs and traces are built for high-cardinality, per-event data in a way metrics aren’t. The same rule applies to high-cardinality attributes in OpenTelemetry metrics.

Takeaway: If a label’s value set grows with your user base or request volume, it doesn’t belong on a metric.

Kubernetes label multiplication

Kubernetes doesn’t need a bad instrumentation decision to hit high cardinality. In most Kubernetes scrape setups, every metric scraped from a pod carries pod, namespace, and instance labels. A metric exposed by a hundred pods becomes a hundred series before any application label is added. (See the Kubernetes metrics worth collecting for which ones earn their series.)

Deployments, autoscaling, and rolling restarts compound this by retiring old pod names and minting new ones. Head series climb over time even on a cluster with stable traffic. kube-state-metrics is a frequent contributor, since it exports per-object metrics for every pod, ReplicaSet, and job in the cluster, including ones that are long gone from your dashboards.

Takeaway: In Kubernetes, budget for a cardinality multiplier from pod churn alone. Restrict kube-state-metrics to the metrics and labels you actually use from day one, and keep per-pod scraping, such as a Prometheus PodMonitor, scoped to pods that need it.

Histogram bucket explosion

Histograms add one series per bucket boundary (the le label) on top of every other label combination. A latency histogram with 10 buckets on the earlier method/status_code/endpoint example doesn’t produce 1,200 series. It produces 12,000, plus the _sum and _count series.

Native histograms store bucket data in a single sparse series instead of one series per bucket. They’re a stable feature as of Prometheus v3.8. Scraping them still has to be enabled with the scrape_native_histograms setting, and your client library has to emit them, so they aren’t a drop-in fix for an existing histogram-heavy setup.

Takeaway: Audit histogram bucket counts as carefully as label counts.

Federation and remote-write amplification

Federating raw metrics from leaf Prometheus instances into a central one relocates the cardinality problem. It doesn’t solve it. The central instance inherits every leaf’s series plus its own overhead, and OOMs are common the first time someone runs a cross-cluster query.

The same applies to remote_write into Thanos, Mimir, or another long-term store. Forwarding raw series means the storage tier absorbs the full cardinality of every source instance.

Takeaway: Pre-aggregate with recording rules before federating or remote-writing. Send the rollup upstream, not the raw series.

No cardinality monitoring

The last cause is an operational gap, not a technical mistake. Teams that don’t track prometheus_tsdb_head_series over time have no baseline. A slow climb over weeks looks like nothing until it’s suddenly everything.

Takeaway: Alert on head series growth against a baseline, not just an absolute threshold, so growth gets caught while it’s still cheap to fix.

How to find high-cardinality metrics

Work from the top down: total series, then the worst metrics, then the worst label inside each.

1. Check total and churning series.

prometheus_tsdb_head_series
rate(prometheus_tsdb_head_series_created_total[1h])

A high creation rate with flat traffic means labels are churning.

2. Find the metrics with the most series.

topk(10, count by (__name__)({__name__=~".+"}))

This query touches every series, so run it off-peak on a large server.

3. Find the label driving a specific metric.

count(count by (path) (http_requests_total))

Swap path for each label on the metric. The one with the largest count is your multiplier.

4. Ask the TSDB directly.

Open the TSDB Status page in the Prometheus UI (Status → TSDB Status). It lists the top metric names, label names, and label-value pairs by series count, without running an expensive query. The same data is available from the /api/v1/status/tsdb endpoint if you want to pull it into a script.

For data already written to disk, the promtool tsdb analyze command reports the highest-cardinality labels and metric names for a given data directory.

How to reduce Prometheus cardinality

Fix it as early in the pipeline as you can. Instrumentation beats relabeling, and relabeling beats paying to store the series.

1. Fix instrumentation. Remove unbounded labels and normalize routes to templates. In Flask, for example, label by the route rule rather than the raw path:

REQUESTS.labels(
    method=request.method,
    route=request.url_rule.rule,  # "/users/<int:id>", not "/users/48213"
    status=str(response.status_code),
).inc()

2. Drop labels or metrics at scrape time. metric_relabel_configs runs after the scrape and before storage:

scrape_configs:
  - job_name: payments
    sample_limit: 50000   # reject the scrape if it exceeds this
    label_limit: 30       # reject the scrape if any series has more labels
    metric_relabel_configs:
      - action: labeldrop
        regex: request_id|session_id
      - source_labels: [__name__]
        regex: debug_.*
        action: drop

Dropping a label with labeldrop can make two series identical. Only drop labels when the remaining labels still uniquely identify each series.

If you run the Middleware Kubernetes agent, you can do the same job from the UI. The Attributes and Filter processors in OTel-Native Filters remove labels or drop whole metrics inside the cluster, without editing scrape configs.

3. Pre-aggregate with recording rules.

groups:
  - name: http_aggregates
    interval: 1m
    rules:
      - record: job:http_requests:rate5m
        expr: sum by (job, method, status_code) (rate(http_requests_total[5m]))

A recording rule makes dashboards and alerts cheaper, but the raw series still exist locally. To keep them out of long-term storage, forward only the rollups:

remote_write:
  - url: https://<your-remote-store>/api/v1/write
    write_relabel_configs:
      - source_labels: [__name__]
        regex: "job:.*"
        action: keep

4. Restrict kube-state-metrics. Use --metric-denylist (or --metric-allowlist) to stop exporting metric families nobody uses. Keep --metric-labels-allowlist limited to the Kubernetes labels you actually group by, for example --metric-labels-allowlist=pods=[app,team].

5. Alert on growth.

- alert: PrometheusHeadSeriesGrowing
  expr: prometheus_tsdb_head_series / (prometheus_tsdb_head_series offset 7d) > 1.25
  for: 1h
  labels:
    severity: warning

Drop noisy labels before they leave the cluster

Middleware Pipelines run OpenTelemetry-native filters inside your cluster, so churn-prone labels can be trimmed before they’re sent or stored.

How to decide which fix to apply

Start with three questions before touching any config:

  1. Is the label ever queried or aggregated on? If nobody runs sum by (that_label) or filters on it, it shouldn’t exist as a label.
  2. Does the cardinality come from instrumentation or infrastructure metadata? Instrumentation problems need a code change. Kubernetes metadata problems need relabeling.
  3. Do you need the raw series, or just an aggregate? If it’s an aggregate, a recording rule solves it without losing the query you actually run.

From there, match the fix to the situation:

  • A payments service that just shipped a customer_id label on its request counter. Pull the label in the next deploy. Replace it with a bounded customer_tier label (free/pro/enterprise) if segmentation is genuinely needed, and route per-customer detail to structured logs.
  • A cluster running kube-state-metrics with default settings on 300+ nodes. Deny the metric families nobody dashboards on, keep the label allowlist tight, and drop per-object labels you never group by with labeldrop or the Middleware Attributes processor.
  • A federation layer that OOMs the moment someone runs a cross-cluster query. Replace raw federation with recording rules computed at each leaf, so the central instance only ingests pre-aggregated series.
  • A latency histogram with 20 default buckets across a high-endpoint-count API. Cut the bucket list to the 6-8 boundaries that matter for your SLOs, or move to native histograms if your client library and server support them.
  • A team that keeps discovering cardinality problems after the on-call page. Add the growth-rate alert above, so a slow leak gets caught in week two instead of month three.

How Middleware helps you control Prometheus cardinality

Every fix above works in plain Prometheus. The catch is that each one lives somewhere different: instrumentation in code, relabeling in scrape configs, recording rules in rule files, and alerts in yet another file. Middleware puts the same controls in one OpenTelemetry-native platform, next to the logs and traces you need once an incident starts.

Cardinality problemPrometheus-native fixHow Middleware handles it
Churn-prone label such as a pod UID or container IDlabeldrop in metric_relabel_configsAttributes processor removes it inside the cluster
Label that is noisy only for some valuesRelabel rule with a regexCompact processor removes the attribute only when values match a pattern
Metric nobody queriesdrop action or kube-state-metrics denylistFilter processor drops it in-cluster, or toggle the metric off at ingestion
Slow series growthGrowth-rate alert on head seriesAnomaly detection with dynamic per-metric baselines and forecasting
Per-user detail needed for debuggingMove it to logs or tracesLogs and traces correlated with metrics on one timeline

Bring Prometheus metrics in without re-instrumenting

Your existing exporters and /metrics endpoints keep working. Middleware can collect them three ways:

  • Kubernetes annotation scraping. If you already run the Middleware Kubernetes agent, turn on annotation-based scraping for a cluster from the UI. Any pod with prometheus.io/scrape, prometheus.io/port, and prometheus.io/path annotations gets scraped. See Prometheus scraping on Kubernetes.
  • Host agent. Point the Middleware host agent at a Prometheus endpoint on a host by adding a job name and target URL.
  • Prometheus remote write. Keep your Prometheus server and forward its data to Middleware, as covered in our Prometheus PodMonitor guide.

Scraped metrics are available for custom dashboards and alerts as soon as they arrive, alongside Kubernetes monitoring data from the same cluster.

Strip high-cardinality labels before they leave the cluster

OTel-Native Filters are the agent-side layer of Middleware Pipelines. They run inside the OpenTelemetry Collector in your Kubernetes cluster and work on metrics, logs, and traces. You configure them from the UI under Settings → Pipelines, with no Collector YAML to maintain.

  • Attributes removes or renames attributes, for example dropping k8s.pod.uid from every metric.
  • Compact removes an attribute only when its value matches a noisy pattern, such as http.route for /health and /metrics.
  • Filter drops telemetry entirely based on conditions combined with AND/OR.
  • Redaction masks sensitive values such as emails, IPs, and UUIDs.

Processors run in order (Attributes, Compact, Filter, Redaction) before data is sent out of the cluster. That saves network and ingestion cost, because a label you strip here never reaches Middleware at all.

OTel-Native Filters are available for Kubernetes cluster sources only, with one filter set per signal in each pipeline. For metrics scraped on standalone hosts, fix the instrumentation or use metric_relabel_configs as shown earlier.

Turn off metrics you don’t use

Many cardinality problems come from metric families nobody looks at. Infrastructure monitoring lets you toggle individual metrics on or off at ingestion. An unused family stops adding series without a scrape config change or a redeploy.

Catch series growth before it becomes an outage

Static thresholds miss slow leaks. Middleware’s anomaly detection builds a dynamic baseline for each metric, and forecasting flags resources heading toward saturation. Alerts are evaluated every 15 seconds by default and can go to Slack, Microsoft Teams, PagerDuty, OpsGenie, or email. If you scrape your Prometheus server’s own /metrics endpoint, you can alert on prometheus_tsdb_head_series in Middleware too.

Keep per-user detail where it belongs

The main fix for unbounded labels is moving per-request detail out of metrics and into logs and traces. In Middleware, that detail stays connected. From a metric spike you can jump to the logs and traces for the same service and time window, so dropping user_id from a counter doesn’t cost you the ability to debug one customer’s request.

OpsAI, Middleware’s AI SRE agent, analyzes metrics, logs, traces, and Kubernetes telemetry together to surface the likely root cause when an alert fires.

Investigate without writing PromQL

Dashboards for hosts, Kubernetes, and containers are generated automatically on first deployment. For custom views, you can build a dashboard from a text prompt, which helps when not everyone on the on-call rotation writes PromQL.

Pricing that rewards cutting cardinality

Middleware uses usage-based, per-GB pricing across metrics, logs, and traces, so data you drop in the cluster is data you don’t pay to ingest. New accounts start with a 14-day free trial and can then move to a Free Forever plan with monthly usage limits. The same model scales from a single cluster to enterprise fleets, with a BYOC option for teams that need custom retention and a dedicated account team.

If you’re weighing a move off a Prometheus-only setup, see our roundup of Prometheus alternatives or how Prometheus, Grafana, and Middleware compare.

Know which metrics are driving your bill

Middleware pricing is usage-based, so series you cut at the source are series you don’t pay to ingest.

FAQs

What is a cardinality explosion in Prometheus?

A cardinality explosion is a rapid increase in the number of active time series. It’s usually caused by a label with unbounded values, such as user IDs or raw URLs. Each new value creates a new series, so memory use climbs until queries slow down or Prometheus is OOM-killed.

What is a good cardinality target for a single Prometheus metric?

The Prometheus instrumentation docs suggest keeping most metrics below a cardinality of 10, with the vast majority carrying no labels at all. Any metric that has a cardinality over 100, or could grow that large, is worth redesigning before it ships.

How much memory does each Prometheus series use?

It depends on label count, label length, and churn, so there’s no fixed number. Divide process_resident_memory_bytes by prometheus_tsdb_head_series on your own server to get your real per-series cost.

Does dropping a label at query time reduce cardinality?

No. Cardinality is set at ingestion, when Prometheus creates a new series. Aggregating a label away in a query only changes what a dashboard shows. It doesn’t remove the underlying series or the memory they use.

How do I find which metrics are causing high cardinality right now?

Run topk(10, count by (__name__)({__name__=~".+"})), or call the /api/v1/status/tsdb endpoint. The endpoint returns the top metric names and label names by series count directly from the TSDB.

Can sample_limit prevent a cardinality explosion?

It contains the blast radius. If a scrape exceeds sample_limit, Prometheus rejects the whole scrape, which stops one bad target from taking down the instance. The instrumentation problem still needs fixing, and the target’s metrics are missing until it is.

Can Middleware drop high-cardinality Prometheus labels?

Yes. On Kubernetes, OTel-Native Filters in Middleware Pipelines remove attributes, drop metrics by condition, or strip noisy values inside the cluster before data is sent. Individual metrics can also be toggled off at ingestion. For standalone hosts, drop labels at scrape time with metric_relabel_configs.

Is high cardinality always bad?

No. High cardinality is a design cost, not a defect. A label with a few thousand well-understood values, like endpoint on a large API, can be worth its storage if it’s queried. The problem is unbounded or unqueried cardinality.