DevOps Practitioner · Module 7: Applied observability and SLOs
SLIs, SLOs and Alerts That Page for Users
Not every red graph deserves a page. SLIs measure what users feel, SLOs promise how much goodness, error budgets ration the failures, and alerts fire on burn that endangers the promise, nothing else.
10 min reading
Objectives
- Define SLI (what you measure), SLO (the promise) and error budget (what you may spend)
- Compute an SLI and remaining budget from a given series
- Page on budget burn that threatens users, ticket the rest
- Justify every alert with the user pain it prevents
Why this matters
A pager fires forty times a week and the team stops hurrying: disk at eighty percent, CPU spikes, a certificate with months left. Then the checkout error rate triples at 3 AM and the alert sits in the same noisy queue, seen late. Alert fatigue is not a morale problem, it is a design failure: pages for machine symptoms instead of user pain. The fix is a budget: pages fire when the error budget burns too fast, because budget burn is user pain measured directly.
Concepts
An SLI is a ratio over valid events: successful responses over total requests, on-time deliveries over all deliveries. The SLO is the target (99.9 percent monthly); the error budget is the allowed failure share (0.1 percent), spendable on releases, experiments and risk. Burn rate is budget consumption speed: fast burn pages (users hurt now), slow burn tickets (trend to watch). The L45 lab computes SLI and remaining budget from a given fixture series and justifies the alert decision in writing.
Alert design follows the budget. Page when fast burn threatens the SLO within the response window, with the runbook attached and the symptom stated as user impact. Ticket slow burn and symptom metrics (CPU, disk, queue depth) for working hours; they predict, they do not page. Every alert carries its justification: which user pain, at what burn, with what first action. Alerts without justification get deleted in review, no matter who wrote them.
Budgets also govern delivery. Budget healthy means releases proceed; budget exhausted means freezes and reliability work until it recovers. This turns the reliability-versus-features fight into arithmetic both sides can read. Fixture numbers teach the arithmetic; production numbers set the policy, and the two are never confused in the record.
Worked example
A fixture series shows a service at 99.5 percent against a 99.9 SLO with two days left in the month. The learner computes the burned budget, projects exhaustion before month end at current burn, and justifies a page for the fast-burn window plus a ticket for the trend. Then a CPU-spike alert with no budget link is deleted in the same review, with the reason quoted.
Common wrong move
Paging on infrastructure symptoms with user-agnostic thresholds. Disk, CPU and certificate alerts have their place as tickets with runbooks; as pages they train the team to ignore the pager, and the training succeeds exactly when it matters most.
Quick check
An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Lesson feedback
No published feedback yet.
Log in and complete the lesson to leave feedback.
Exercise
From a given fixture series, compute the SLI and remaining error budget, project exhaustion, and write the page-or-ticket decision with its justification.
Pass criteria
The record shows the SLI math, the budget remaining with projection, and the justified alert decision quoting user impact.