DevOps Practitioner · Module 7: Applied observability and SLOs
Traces: Following a Request Across Services
Logs say what happened; traces say where the time went and who called whom. A waterfall of spans turns a slow request from a mystery into a named hop with a measured duration.
10 min reading
Objectives
- Explain what a trace adds over correlated logs: timing and parentage
- Read a trace waterfall and name the slow span
- Propagate context so hops join one trace instead of starting many
- Sample traces to afford the visibility without drowning in spans
Why this matters
A page takes four seconds and every service reports healthy sub-second timings. The time hides between the measurements: serialization, retries, a downstream call nobody logged. Correlated logs list the events; only the trace shows the gaps, because spans record start, duration and parentage. Without traces the team optimizes the services that measure well while the real cost sits in the unmeasured spaces between them.
Concepts
A trace is a tree of spans for one request: the root at the edge, child spans per hop, each with a name, start time, duration and the parent it answers to. The waterfall renders the tree against time, and the slow span is visible at a glance: the wide bar nobody expected. Context propagation (trace headers passed like the request ID, by the same middleware) joins hops into one trace; a hop that drops context starts an orphan trace, and orphans are the trace version of the missing log line from lesson 2.
Spans carry attributes, not essays: route, status, peer service, the decisions made. Errors attach to the span where they happened with the exception and the ID. Instrumentation mixes automatic (framework middleware creating spans for inbound and outbound calls) with manual (wrapping the expensive function the team suspects). Start automatic everywhere, add manual where the waterfall shows a gap with no span.
Sampling keeps the bill sane. Head sampling decides at the edge (keep a fraction of traces); tail sampling keeps the interesting ones (errors, slow requests) after seeing the outcome. Small labs keep everything; production keeps a fraction plus all failures. The fixture numbers in this module's labs are labeled fixtures: they teach the reading skill, not the production budget.
Worked example
A three-hop fixture serves a slow request: the waterfall shows two fast hops and one wide downstream bar with a retry inside it. The learner names the slow span, reads its attributes for the retried call, and quotes the durations that add up to the total. Removing the retry in the fixture narrows the bar and the total follows.
Common wrong move
Adding tracing to one service and declaring observability done. A single service's spans are expensive logs; the value starts when context crosses boundaries. Instrument the edges first, then the hops that hurt.
Quick check
An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Lesson feedback
No published feedback yet.
Log in and complete the lesson to leave feedback.
Exercise
On a multi-hop fixture, read the trace waterfall of a slow request, name the slow span with its durations, and show the total explained by the parts.
Pass criteria
The record shows the waterfall reading, the named slow span with quoted durations, and the parts adding to the total.