The most detailed free FDE + DevOps library: 140+ lessons, 70+ labs and 80 long-form articles, in English and Turkish. Start learning →

Incident response: winning the first hour

Updated

An engineer at a control panel in a control room

The first hour decides the incident: one commander, a written timeline, and a go or no-go rollback call.

Incidents are lost in the first hour and reviewed in the second week. The difference between a 40 minute outage and a four hour one is rarely technical skill. It is structure: someone declares, someone decides, someone writes, and everyone else investigates.

The first-hour playbook

Declare early. A short message names the incident, the commander and the channel. Declaring is cheap; discovering at minute 50 that three people debugged the same thing separately is expensive.

One commander, many investigators. The commander does not debug. They assign, timebox, keep the timeline and make the rollback call. The best debugger in the room investigates; commanding wastes their skill and the incident.

Decide rollback on impact. Go or no-go within the first 30 minutes: is the blast radius growing, and is the fix uncertain? If yes to both, roll back first. Cause analysis happens after customers are safe.

Worked example: a fictional checkout outage

The context below is fictional. Fictional shop BrightCart (fictional) sees checkout errors climb after a noon deploy. Minute 5: the on-call declares the incident and takes command in the team channel. Minute 12: two investigators confirm errors track the new release, not traffic. Minute 20: blast radius still growing, fix uncertain, so the commander calls rollback to the previous tag. Minute 35: errors flat, customers checking out, timeline written so far. Next morning: the postmortem finds the caching bug on a calm branch with tests, and the only action items are a faster rollback path and a missing alert, not heroics.

Decision table: the rollback call

Blast radiusFix confidenceCall
GrowingUncertainRoll back now
GrowingCertain and minutes awayFix forward with a deadline, rollback ready
FlatAnyInvestigate calmly, keep the timeline

Checklist: a healthy first hour

  1. Incident declared with commander and channel named.
  2. Timeline written live, not reconstructed later.
  3. Rollback decision made explicitly, go or no-go with reasons.
  4. Customers informed once the direction is set, not before.

Straight answers

Frequently asked questions

Who should lead the incident?

One commander who coordinates, not the deepest expert. Experts investigate; the commander decides and communicates.

When do you roll back?

When the blast radius is growing and the fix is uncertain. Rollback is a decision, not an admission of defeat.

What if we do not know the cause yet?

Decide on impact, not cause. Stop the bleeding on what you can see; root cause waits for daylight.

Bu sayfanın Türkçesi

Turn reading into a credential

This post is a free field note. Exams run at dated sittings in 15-seat classes; one price covers one attempt. All lessons are free.