The most detailed free FDE + DevOps library: 140+ lessons, 70+ labs and 80 long-form articles, in English and Turkish. Start learning →

DevOps Foundations · Module 7: Operations, observability and backup

Service Indicators Worth Paging On

Most alerts describe machine feelings, not user pain. This lesson builds indicators from the user's view so the pager fires when users hurt and stays silent otherwise.

10 min reading

Objectives

  • Define availability, latency and correctness indicators for a real service
  • Set alert thresholds from burn rate instead of vibes
  • Separate paging alerts from tickets with a written rule
  • Explain what each alert tells the responder to do first

Why this matters

A team owns forty alerts and ignores all of them, because thirty-nine fire on CPU feelings and disk predictions while users see errors nobody paged on. Alert fatigue is not a discipline problem; it is a signal problem. Every alert the responder ignores trains the next ignore, until the real page drowns with the noise. Fewer, user-anchored alerts beat broader coverage.

Concepts

Indicators start from user-visible outcomes. Availability: fraction of valid requests served without error. Latency: fraction served fast enough (a percentile with a bound, not an average that hides the tail). Correctness: fraction producing right results, where wrong-but-200 responses exist. Each is a ratio over a window with a target: 99.9 percent of requests under 300ms over 30 days. The target encodes the promise; the window encodes patience. Machine signals (CPU, memory, disk) are causes and capacity data, paged only when no user indicator covers the failure they predict.

Burn rate turns ratios into alerts. Fast burn (a large budget fraction gone in an hour) pages now; slow burn (the same over days) files a ticket for working hours. Both reference the same error budget the target implies: the allowed failure room the team may spend on velocity. An alert without a budget reference is an opinion; with one, it is arithmetic the responder can trust at 3am.

Every paging alert carries its first action. The alert text names the indicator, the scope, and the runbook link; the runbook's first three lines say what to check and what safe action buys time. An alert whose response is unknown is a ticket wearing a pager costume. Review the set quarterly: each alert must have fired usefully or be demoted, because silent alerts are untested code guarding production.

Worked example

Checkout 500s spike at 2am and nobody wakes; CPU alerts fire instead and get snoozed. The indicator work:

availability = 1 - (5xx / total), window 5m, target 99.9 fast burn: 2% budget in 1h -> page slow burn: 5% budget in 6h -> ticket

Expected reading: the 2am spike burns 3 percent of the monthly budget in an hour, so the fast-burn alert pages with the runbook link; the CPU wobble has no user indicator attached and stays a dashboard, not a page. Verify by replaying last month's incidents against the new rules: every user-visible event pages, every ignored CPU feeling stays silent. Lab L19's metric half is this in miniature: one graph that would have paged correctly.

The common wrong move

Averaging latency and alerting on the mean. The average hides the tail where users live: half the requests at 50ms and a tenth at 8 seconds averages to fine while a tenth of users suffer. Percentiles with bounds (p95, p99 under the target) expose what the mean buries. Any latency alert on an average is a comfort object, not an instrument.

Lab and next step

Lab L19's metrics half requires the paging-worthy graph identified and the noise graphs named as such. Next, lesson 3 covers the other half of operations: backups that restore, proven on a separate target.

Quick check

An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.

Lesson feedback

No published feedback yet.

Log in and complete the lesson to leave feedback.

Exercise

Write indicators for one service you run: availability, latency (percentile plus bound) and one correctness check, each with target and window. Derive one fast-burn page and one slow-burn ticket rule, and attach the first responder action to the page.

Pass criteria

Three user-anchored indicators with targets and windows; burn-rate page and ticket rules derived; first responder action written on the page.

Sources

Log in to track progressFree account: stores only your lesson progress and quiz results.