The most detailed free FDE + DevOps library: 140+ lessons, 70+ labs and 80 long-form articles, in English and Turkish. Start learning →

DevOps Foundations · Module 8: Security and safe-change basics

Risk, Go/No-go and the Way Back

Every change is a bet; safe teams size the bet, check the table, and keep the exit marked. This lesson makes the delivery decision explicit and the way back practiced.

10 min reading

Objectives

  • State a change's risk in blast radius, reversibility and evidence
  • Run a go/no-go with named checks instead of gut feel
  • Roll back to the last working release and verify traffic
  • Explain when rollback beats rollforward with a rule, not hope

Why this matters

A bad release poisons checkout for forty minutes while the team debates whether to fix forward, because rollback was never rehearsed and nobody knows if the old version still works. The outage is not the bug; the outage is the missing way back. Rollback practiced monthly turns the same bug into a four-minute footnote. Delivery safety is reversibility plus rehearsal.

Concepts

Risk statements have three parts. Blast radius: who is affected and how much (one region, all users, writes or reads). Reversibility: how fast and how safely the change undoes (flag flip in seconds, redeploy in minutes, migration requiring expand/contract over days from the M14 preview). Evidence: what the change proved before production (gates from M06, load behavior, review depth). A change with wide blast radius, slow reversal and thin evidence is a big bet wearing a small-change costume; size the ceremony to the bet.

Go/no-go is a checklist with owners, run at the decision point. Indicators green (the M07 signals that matter), error budget sufficient for the bet size, rollback rehearsed for this shape of change, support staffed for the watch window, communication drafted for the failure case. Each line has a name beside it; anyone can call no-go, and no-go is a normal outcome, not a failure. Skipped checks get written down as accepted risk with an owner, never silently dropped.

Rollback beats rollforward by default rule: if the fix is not understood within the agreed minutes, or the blast radius is still growing, roll back first and diagnose on the safe version. Rollforward is for understood small faults with the fix in hand; everything else gets the known-good version serving traffic while humans think. The rule decides under pressure so pride does not: minutes and growth, written beforehand.

Delivery mechanics keep the way back open. Keep the last working release one command away (Helm rollback, previous digest redeploy, flag default); verify traffic after every move in both directions; never delete the old version until the new one proves itself through a full watch window. Database changes follow expand/contract so old and new code coexist; a migration the old version cannot read burns the bridge before crossing.

Worked example

Version 2.5.0 raises errors to 8 percent within ten minutes of full rollout. The call:

t+10: errors 8%, blast radius growing, cause unknown → rule says roll back. t+12: previous digest 2.4.3 serving, errors back under 0.5%, verified by graph. t+40: root cause found (schema assumption), fix planned as 2.5.1 with expand/contract. t+next window: 2.5.1 ships with staged rollout and the same go/no-go.

Expected reading: the rollback decision took two minutes because the rule predated the pressure; traffic verified the recovery, not hope; the fix followed the reversible path. Fix-forward would have gambled forty more minutes of user pain on an unknown cause. Lab L24 runs this exact drill against a bad new version with the clock running.

The common wrong move

Canary theater: a canary on 1 percent with nobody watching its indicators, promoted on a timer instead of evidence. The canary exists to prove safety with the M07 indicators before wider rollout; unwatched and auto-promoted it is a slow way to ship the same bug to everyone. Every stage gates on evidence, or the stages are M06 decoration again.

Lab and next step

Lab L24 runs the go/no-go plus rollback flow for a bad new version with measured recovery. This closes M08 and the Foundations level: Linux to delivery with hands on every layer. The path continues to Practitioner (W4) when it ships; until then, operate something small with these habits.

Quick check

An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.

Lesson feedback

No published feedback yet.

Log in and complete the lesson to leave feedback.

Exercise

Write the go/no-go checklist for one service you ship, with owners and the rollback-first rule in minutes. Rehearse a rollback to the previous version, time it, and verify traffic recovery on the graph.

Pass criteria

Checklist written with owners and the rollback rule; rehearsal timed with traffic-verified recovery recorded.

Sources

Log in to track progressFree account: stores only your lesson progress and quiz results.