DevOps Foundations · Module 7: Operations, observability and backup
Logs Tell Stories, Metrics Draw Lines
Logs record what happened once; metrics record how often things happen over time. This lesson assigns each job to its tool so incidents start with a graph and end with a line, not the reverse.
10 min reading
Objectives
- Send event detail to logs and countable signals to metrics, never both jobs to one
- Write log lines a stranger can follow without the source code open
- Name metrics with labels that answer questions instead of creating new ones
- Explain why high-cardinality labels break a metrics backend
Why this matters
An outage starts and the team opens a log stream with forty thousand lines per minute, grepping for a feeling. Twenty minutes later someone asks for the error rate graph and it does not exist, because every signal was logged as text and nothing was counted. The incident runs on vibes until the one person who knows the system arrives. Two tools, two jobs, decided beforehand.
Concepts
Logs are discrete events with context: one request failed with this id, this user, this downstream error, at this time. Each line stands alone and tells a small complete story: timestamp, level, component, request id, what was attempted, what resulted. Structured logs (key=value or JSON) make machines parse what humans still read; the M03 JSON discipline pays off here. Log volume is the cost center: debug in development, info plus above in production, and sampling or level switches for the flood.
Metrics are numbers over time with labels: request count by endpoint and status, latency histograms, queue depth, error ratios. They aggregate, alert and graph; they never explain single events. Cardinality is the budget: labels with bounded values (status code, region) are cheap, labels with unbounded values (user id, request id, timestamp) create a new time series per value until the backend falls over. If the value belongs to one request, it belongs in a log, never in a label.
Correlation ties them together. The request id in the log line matches the exemplar or trace reference on the metric spike; the graph says when and how much, the logs say which and why. Emit both from the same code path with the same identifiers, or the incident has two piles of evidence that refuse to join. Lesson 2 builds the indicators on top of this split.
Worked example
Checkout errors spike and nobody knows which downstream broke. The graph shows a 500 rate by endpoint climbing; exemplars point at request ids; the log lines for those ids share one downstream timeout:
ts=2026-10-02T13:01:11Z level=error svc=checkout req=9f2a downstream=pricing err="timeout after 2s" attempt=3
Expected reading: the metric proved the spike shape and scope, the shared request id joined graph to lines, and the lines name pricing timeouts, not application bugs. Fix the downstream contract instead of scaling the caller. Without the id on both sides, the team would still be grepping. Lab L19 replays this order: graph first, lines second, cause named third.
The common wrong move
Logging everything at debug in production just in case, then drowning the budget and missing the incident in the flood. Or the mirror: metrics with user-id labels that cardinality-crash the backend during the first real traffic. Both come from refusing the split. Decide per signal at write time: counted over time goes to metrics with bounded labels, singular context goes to logs with ids.
Lab and next step
Lab L19 hands you logs plus simple metrics from a sick service and requires the error source found in that order, with the join key quoted. Next, lesson 2 turns raw signals into service indicators worth paging on.
Quick check
An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Lesson feedback
No published feedback yet.
Log in and complete the lesson to leave feedback.
Exercise
Pick one service you run. List five signals it emits and classify each as log or metric with the reason. Find one high-cardinality label or one debug-in-production stream, fix it, and show the before/after volume or series count.
Pass criteria
Five signals classified with reasons; one cardinality or volume fault found and fixed with measured before/after.