The most detailed free FDE + DevOps library: 140+ lessons, 70+ labs and 80 long-form articles, in English and Turkish. Start learning →

DevOps Practitioner · Module 8: GitOps and incident response · Lab

Resolve a Staged Outage: Triage to Postmortem

60 min hands-on · Advanced

A local demo stack with a staged bad-config outage (labeled drill, GitOps path); a timer for phase times. This is practice, never a real incident.

Local guide: run the steps below on your own machine in order, then check the validation list.

Objectives

  • Triage user impact with start time before touching anything
  • Mitigate rollback-first and verify with user-shaped evidence
  • Write a blameless postmortem with systemic causes and owned actions
  1. Step 1

    Triage first

    On the page, write impact, start time and severity from user signals before any fix. Start the timer; the triage note is the first deliverable.

  2. Step 2

    Mitigate then verify

    Run the recorded rollback first and note restore time. Verify with user-shaped evidence (success rates, completed checkouts), not pod state alone.

  3. Step 3

    Learn in writing

    Write the postmortem within the session: timeline from the notes, plural systemic causes (the gate waved through, the missing canary), and two actions with owners and dates. No names as causes.

How to confirm it worked

  • Triage note with impact and start time before fixes
  • Rollback-first restore with time noted
  • User-shaped verification evidence quoted
  • Postmortem with systemic causes and dated owned actions