DevOps Practitioner · Module 7: Applied observability and SLOs
Logs That Correlate: One Request ID Everywhere
Distributed failures scatter evidence across services. One request ID, generated at the edge and logged at every hop, stitches the scattered lines back into a single story.
10 min reading
Objectives
- Explain why a shared request ID beats timestamps for correlation
- Propagate the ID across service calls and log it at every hop
- Follow one failing request through logs with the ID as the thread
- Keep logs structured so machines can join what humans read
Why this matters
A checkout fails and three services log errors within the same second. Timestamps narrow it to a hundred candidate requests across the services, and the team spends an hour guessing which lines belong together. With a shared request ID the same incident takes one query: every line with the ID, ordered by time, is the story. Correlation is not a tool purchase, it is one field propagated everywhere, decided once and enforced by review.
Concepts
Generate the ID at the edge (gateway, ingress, first service) and pass it downstream on every call: headers between services, arguments to workers, fields in events. Each service logs it with every line for that request, alongside the local span of what happened. The L44 lab follows a failure across fixture logs using nothing but the shared ID, proving the mechanism on small data before anyone needs it on large.
Structure makes joining possible. Key-value or JSON lines with stable field names (request_id, service, duration_ms, outcome) let machines filter and group what humans then read. Free-text logs with the ID buried in prose work for one incident and fail at scale; the format is decided in the logging library or middleware, once, not per service by taste. Include the outcome and duration at the point of decision, not three hops later where context is gone.
Sample deliberately. Full debug logging of every request costs storage and attention; the discipline is info by default with the ID always present, debug scoped to the failing pattern, and errors carrying the ID plus the decision context. When an incident needs more, raise verbosity by selector (route, service, ID prefix), never globally and never permanently.
Worked example
A two-service fixture handles a request that fails in the second service. The learner greps both log files for the shared ID and assembles the timeline: edge received, downstream called, dependency timed out, error returned. Then the same failure without the ID takes ten times longer to assemble from timestamps alone. The comparison is quoted, not asserted.
Common wrong move
Logging the ID in some services and not others. Partial propagation breaks the thread exactly where the failure lives, because the failing hop is the one someone skipped. Enforce by middleware and review, not by asking every author to remember.
Quick check
An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Lesson feedback
No published feedback yet.
Log in and complete the lesson to leave feedback.
Exercise
On a two-service fixture, propagate a request ID, fail one request downstream, and assemble the timeline from both logs using only the ID.
Pass criteria
The record shows the ID at every hop in both logs and the assembled timeline with the failing hop named.