Writing runbooks your customer's team can run

Updated

An engineer writing operations documentation on a laptop

The runbook test is simple: a stranger recovers the system without calling you. How to write to that bar.

Most runbooks describe the system. Useful runbooks prescribe the recovery. Write for the second.

The short answer

Each entry: symptom, how to confirm, exact fix commands with expected output, how to verify recovery, when to stop and escalate.

The entry template

Symptom: orders stop flowing, dashboard flatlines at 02:14 pattern. Confirm: run check-lag with source orders and expect lag_seconds under 300; if higher, this entry applies. Fix: 1. replay from checkpoint in dry-run mode and compare counts. 2. replay from checkpoint with apply. 3. Watch delivery_rate for 10 minutes. Verify: lag under 60s for 3 consecutive checks; error report shows zero unexplained drops. Stop line: if duplicates exceed 0.5 percent, halt and page the owner with the checkpoint id. Do not improvise beyond step 3.

Worked example: the three-symptom runbook

A fictional pipeline (fictional) ships with exactly three entries: stalled source, duplicate surge, credential expiry. Each tested by a teammate in staging: 12, 9 and 6 minutes. The fourth symptom discovered later gets the same template. Consistency is the feature.

Checklist: runbook quality

  1. Tested by a non-author on a staged symptom.
  2. Every command has its expected output beside it.
  3. Stop lines name a person, not "escalate appropriately."
  4. Reviewed after every real incident within a week.

Straight answers

Frequently asked questions

What makes a runbook usable?

Symptoms first, exact commands, expected outputs, and a stop line that says when to escalate instead of improvising.

How long should it be?

Three symptoms fully covered beat thirty pages skimmed. Depth on the frequent, pointers for the rare.

Who tests it?

Someone who did not build the system, on a staged symptom, timed. Fix what slows them down.

Bu sayfanın Türkçesi