DevOps Practitioner · Module 8: GitOps and incident response · Lab
Resolve a Staged Outage: Triage to Postmortem
60 min hands-on · Advanced
A local demo stack with a staged bad-config outage (labeled drill, GitOps path); a timer for phase times. This is practice, never a real incident.
Local guide: run the steps below on your own machine in order, then check the validation list.
Objectives
- Triage user impact with start time before touching anything
- Mitigate rollback-first and verify with user-shaped evidence
- Write a blameless postmortem with systemic causes and owned actions
Step 1
Triage first
On the page, write impact, start time and severity from user signals before any fix. Start the timer; the triage note is the first deliverable.
Step 2
Mitigate then verify
Run the recorded rollback first and note restore time. Verify with user-shaped evidence (success rates, completed checkouts), not pod state alone.
Step 3
Learn in writing
Write the postmortem within the session: timeline from the notes, plural systemic causes (the gate waved through, the missing canary), and two actions with owners and dates. No names as causes.
How to confirm it worked
- Triage note with impact and start time before fixes
- Rollback-first restore with time noted
- User-shaped verification evidence quoted
- Postmortem with systemic causes and dated owned actions