DevOps Practitioner · Module 6: Safe delivery and release engineering
Change Gates and Rollback Before Diagnosis
Gates put human judgment where it matters, and rollback-first puts users before curiosity. When the release breaks traffic, restore service with the recorded path, then study the failure with the pressure off.
10 min reading
Objectives
- Explain what a change gate checks and who can stop a release
- Place gates where judgment changes the outcome, not as confetti
- Roll back first when traffic breaks, diagnose second
- Record every gate decision so the history explains itself
Why this matters
Traffic breaks after a release and the team spends forty minutes reading logs while users fail, because nobody wants to roll back a change that might be innocent. The change was guilty enough: it arrived with the failure. Rollback-first is not blame, it is triage: restore the last known good in minutes, then diagnose with users safe and the evidence preserved. Every minute of innocent-until-proven-guilty debugging is a minute billed to users.
Concepts
A change gate is a checkpoint with a named approver and recorded criteria: tests green on the promoted artifact, canary analysis within bounds, schema phase verified, on-call staffed. Gates belong where human judgment changes the outcome: production entry for risky changes, contract phase for schema removal, traffic switch for blue-green. Gates between automated steps where nobody ever says no are confetti: they add queue time and teach approvers to click through, which corrupts the gates that matter. The M06 stage-gate lesson returns here at release scale.
Rollback-first needs a recorded path to work under pressure. Revision history (dp09l2), release revisions (dp11l4), previous variables (M13): whichever machinery the service uses, the path back is written before the release, not invented during it. The release checklist names the rollback command alongside the deploy command; the drill rehearses both.
Gate decisions are records, not chats. Who approved, on what evidence, at what time: the history explains after the fact why a risky release went ahead. Approvals in chat scroll away; approvals in the release record survive the postmortem. Denied gates record too, with the reason, so the same debate does not recur every release.
Worked example
A demo release breaks traffic after passing its gates. The learner runs the recorded rollback first (service restored, time noted), then diagnoses from the preserved evidence: a schema phase skipped under schedule pressure with the gate waved through in chat. The postmortem records both the technical cause and the gate failure, and the gate moves into the release record so waving it through stops being easy.
Common wrong move
Diagnosing forward while users fail because rollback feels like defeat. Rollback is the fastest diagnostic available: if the failure follows the release back, the release caused it; if it stays, something else did. Restore first, learn second.
Quick check
An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Lesson feedback
No published feedback yet.
Log in and complete the lesson to leave feedback.
Exercise
Break demo traffic with a release, run the recorded rollback first with the time noted, then diagnose from preserved evidence and record both causes.
Pass criteria
The record shows restore time from rollback, the technical cause, the process cause, and the gate fix that prevents recurrence.