Writing runbooks your customer's team can run
Updated

The runbook test is simple: a stranger recovers the system without calling you. How to write to that bar.
Most runbooks describe the system. Useful runbooks prescribe the recovery. Write for the second.
The short answer
Each entry: symptom, how to confirm, exact fix commands with expected output, how to verify recovery, when to stop and escalate.
The entry template
Symptom: orders stop flowing, dashboard flatlines at 02:14 pattern. Confirm: run check-lag with source orders and expect lag_seconds under 300; if higher, this entry applies. Fix: 1. replay from checkpoint in dry-run mode and compare counts. 2. replay from checkpoint with apply. 3. Watch delivery_rate for 10 minutes. Verify: lag under 60s for 3 consecutive checks; error report shows zero unexplained drops. Stop line: if duplicates exceed 0.5 percent, halt and page the owner with the checkpoint id. Do not improvise beyond step 3.
Worked example: the three-symptom runbook
A fictional pipeline (fictional) ships with exactly three entries: stalled source, duplicate surge, credential expiry. Each tested by a teammate in staging: 12, 9 and 6 minutes. The fourth symptom discovered later gets the same template. Consistency is the feature.
Checklist: runbook quality
- Tested by a non-author on a staged symptom.
- Every command has its expected output beside it.
- Stop lines name a person, not "escalate appropriately."
- Reviewed after every real incident within a week.
Related reading
Straight answers
Frequently asked questions
What makes a runbook usable?
Symptoms first, exact commands, expected outputs, and a stop line that says when to escalate instead of improvising.
How long should it be?
Three symptoms fully covered beat thirty pages skimmed. Depth on the frequent, pointers for the rare.
Who tests it?
Someone who did not build the system, on a staged symptom, timed. Fix what slows them down.