Incident response: winning the first hour
Updated

The first hour decides the incident: one commander, a written timeline, and a go or no-go rollback call.
Incidents are lost in the first hour and reviewed in the second week. The difference between a 40 minute outage and a four hour one is rarely technical skill. It is structure: someone declares, someone decides, someone writes, and everyone else investigates.
The first-hour playbook
Declare early. A short message names the incident, the commander and the channel. Declaring is cheap; discovering at minute 50 that three people debugged the same thing separately is expensive.
One commander, many investigators. The commander does not debug. They assign, timebox, keep the timeline and make the rollback call. The best debugger in the room investigates; commanding wastes their skill and the incident.
Decide rollback on impact. Go or no-go within the first 30 minutes: is the blast radius growing, and is the fix uncertain? If yes to both, roll back first. Cause analysis happens after customers are safe.
Worked example: a fictional checkout outage
The context below is fictional. Fictional shop BrightCart (fictional) sees checkout errors climb after a noon deploy. Minute 5: the on-call declares the incident and takes command in the team channel. Minute 12: two investigators confirm errors track the new release, not traffic. Minute 20: blast radius still growing, fix uncertain, so the commander calls rollback to the previous tag. Minute 35: errors flat, customers checking out, timeline written so far. Next morning: the postmortem finds the caching bug on a calm branch with tests, and the only action items are a faster rollback path and a missing alert, not heroics.
Decision table: the rollback call
| Blast radius | Fix confidence | Call |
|---|---|---|
| Growing | Uncertain | Roll back now |
| Growing | Certain and minutes away | Fix forward with a deadline, rollback ready |
| Flat | Any | Investigate calmly, keep the timeline |
Checklist: a healthy first hour
- Incident declared with commander and channel named.
- Timeline written live, not reconstructed later.
- Rollback decision made explicitly, go or no-go with reasons.
- Customers informed once the direction is set, not before.
Related reading
- Rehearse the call: Go or No-Go on a Bad Release.
- Find the source fast: Error Source in Logs and Metrics.
- The program behind the practice: DevOps program.
Straight answers
Frequently asked questions
Who should lead the incident?
One commander who coordinates, not the deepest expert. Experts investigate; the commander decides and communicates.
When do you roll back?
When the blast radius is growing and the fix is uncertain. Rollback is a decision, not an admission of defeat.
What if we do not know the cause yet?
Decide on impact, not cause. Stop the bleeding on what you can see; root cause waits for daylight.