The most detailed free FDE + DevOps library: 140+ lessons, 70+ labs and 80 long-form articles, in English and Turkish. Start learning →

DevOps Practitioner · Module 7: Applied observability and SLOs

Metrics, Labels and the Cardinality Budget

Labels turn one measurement into many answerable questions, and each label value multiplies the stored series. Cardinality is the bill: budget it like money or it bankrupts the monitoring.

10 min reading

Objectives

  • Explain what a metric with labels lets you ask and what it costs
  • Predict which label choices explode cardinality before collecting
  • Keep high-cardinality identifiers out of labels
  • Read a metric series and name what each label dimension contributes

Why this matters

A team adds user_id as a metric label to debug one incident. Series count grows from thousands to millions, the metrics store slows for everyone, retention is cut to pay the bill, and the historical data the next incident needs is already gone. One label, chosen in a hurry, taxed the whole organization. Label design is schema design: cheap to sketch, ruinous to fix after collection, so it gets decided deliberately up front.

Concepts

A metric names what is measured (http_requests_total); labels slice it by dimensions (method, status, route). Each unique label combination is a series the store keeps; cardinality is the count of those series. Low-cardinality labels (method with four values, status with five) multiply gently. High-cardinality values (user IDs, request IDs, timestamps, unbounded paths) multiply without bound and must never become labels. Route patterns (/users/:id) instead of raw paths (/users/48312) keep the dimension bounded while preserving the debugging value.

Choose counters, gauges and histograms by question. Counters (requests, errors) answer rates over time. Gauges (queue depth, temperature) answer current state. Histograms with explicit buckets answer latency distribution: what fraction of requests beat the budget. Averages of latency hide the slow tail users feel; quantiles computed from buckets show it, which is why the bucket boundaries are a design decision, not a default to accept blindly.

Budget cardinality per service: estimate series (metrics times label combinations), set a ceiling the store can hold, and alert on approaching it. When a new label is proposed, the review asks what question it answers that existing labels cannot; no question, no label. The L43 lab works a fixture series where a wrong threshold meets a real budget.

Worked example

A demo service exposes request metrics with method, status and route labels. The learner counts the series, then stages the fault: a raw-path label that explodes the count a hundredfold on fixture traffic. The fix buckets paths to patterns, the count returns to budget, and the before-after numbers are quoted. The lesson sticks because the bill is counted in series, not in adjectives.

Common wrong move

Labeling with request or user IDs to make debugging easy for one person. It makes debugging impossible for everyone when the store slows and retention shrinks. Correlate individual requests with logs and traces (lessons 2 and 3); keep metrics aggregated.

Quick check

An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.

Lesson feedback

No published feedback yet.

Log in and complete the lesson to leave feedback.

Exercise

On a fixture metrics endpoint, count the series, stage a raw-path label explosion, quote the before-after counts, and fix it with route patterns.

Pass criteria

The record shows the baseline series count, the exploded count with its cause named, and the fixed count back inside budget.

Sources

Log in to track progressFree account: stores only your lesson progress and quiz results.