The most detailed free FDE + DevOps library: 140+ lessons, 70+ labs and 80 long-form articles, in English and Turkish. Start learning →

DevOps Practitioner · Module 8: GitOps and incident response

Incident Response: Triage, Mitigate, Verify, Postmortem

Incidents are managed, not merely fixed. Triage names the user impact, mitigation restores service fast, verification proves it with evidence, and the postmortem converts the incident into prevention.

10 min reading

Objectives

  • Run the four phases in order: triage, mitigate, verify, learn
  • Mitigate with rollback-first before any root-cause work
  • Verify the fix with user-shaped evidence, not green dashboards alone
  • Write a postmortem that names causes and actions without blame

Why this matters

An outage pages at night and five engineers debug five theories in parallel while nobody tells users anything. Two fixes collide, the timeline is reconstructed from memory a week later, and the postmortem concludes human error with no action that would stop a recurrence. The failure was not the bug, it was the absence of a practiced shape: one lead, one timeline, mitigation first, learning last. This lesson is that shape, and the L48 lab runs it end to end on a short outage.

Concepts

Triage bounds the incident: what users feel, since when, how broad. Severity follows user impact, not technical drama. One lead coordinates; everyone else works a task or watches. Communication starts early with what is known and what is next, even when the known is small; silence manufactures rumors faster than incidents manufacture facts.

Mitigation restores service by the fastest recorded path: rollback (dp14l4), traffic shift, feature flag, scaled capacity. Diagnosis waits its turn; curiosity during user impact is billed to users. Verification proves restoration with user-shaped evidence (success rates, checkout completions from the M15 toolkit), because green pods with failing checkouts are the incident continuing quietly.

The postmortem is written within days, blameless by construction: what happened (timeline from the recorded log), why each decision made sense at the time, what the causes were (plural, systemic, never a name), and actions with owners and dates. Action items fix systems and guards (alerts, gates, runbooks); training issues get training, not punishment. Unwritten incidents recur; the writing is the prevention.

Runbooks carry the practiced motions: symptom, diagnosis steps with commands, the mitigation path, the verification check, escalation. The M07 runbook lesson returns at incident scale. A runbook that nobody has rehearsed is a rumor; L48 rehearses one.

Worked example

A short staged outage (bad config through the GitOps path) pages the learner. Triage names the impact and start time; rollback-first restores service with the time noted; verification quotes the recovered user signals; the postmortem names the gate waved through (M14's lesson) plus the missing canary, with two owned actions. The whole arc fits one sitting because the shape was practiced, not invented.

Common wrong move

Skipping verification because the graph looks better. Partial recovery with a green dashboard is the incident relapsing on a delay. User-shaped evidence or it did not recover.

Quick check

An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.

Lesson feedback

No published feedback yet.

Log in and complete the lesson to leave feedback.

Exercise

On a staged short outage, run triage to mitigation to verification with times noted, then write the blameless postmortem with owned actions.

Pass criteria

The record shows impact with times, the rollback-first restore, user-shaped verification evidence, and a postmortem with causes and dated owned actions.

Sources

Log in to track progressFree account: stores only your lesson progress and quiz results.