The most detailed free FDE + DevOps library: 140+ lessons, 70+ labs and 80 long-form articles, in English and Turkish. Start learning →

DevOps Professional · Module 3: Migration, restore and disaster recovery

Restore versus Rollback: Different Tools, Different Moments

Rollback fixes bad code with good data underneath; restore fixes bad data whatever the code. Mixing them up destroys the good half while chasing the bad one. Decide from evidence, sequence carefully, record everything.

10 min reading

Objectives

  • Explain when rollback is right (bad release, good data) and when restore is right (bad data, any release)
  • Decide between them from evidence in minutes, not from habit
  • Sequence them safely when both are needed: restore data, then roll code
  • Record the decision and its evidence for the postmortem

Why this matters

A bad migration corrupts records and the team rolls back the code three times, each rollback perfectly restoring the application that reads the corrupted data perfectly. Hours pass before someone asks whether the data is the patient. Rollback was the reflex; restore was the cure, and the reflex burned the incident window. The decision between them is the highest-leverage minute of a data incident, and it must come from evidence, not from whichever runbook opened first.

Concepts

The decision rule fits one paragraph. If the data is good and the code is bad (new release serves errors, old data reads fine), roll back the code (M14's machinery). If the data is bad (corruption, bad migration, deleted rows), restore the data from the last known good inside RPO, whatever the code version. If both are bad, restore data first (to the point the old code understands), then roll code to match; reversing the order writes new-code shapes into restored data or old-code reads against new shapes.

Evidence decides in minutes. Serve reads from a copy: does the old code read clean (code fault), does every version read corrupt (data fault)? Check the deploy and data timelines for coincidence: failure starting with a release implicates code, with a migration or job implicates data. The check is a query and a timeline, not a debate; time-box it, because the decision delayed is darkness extended.

Sequencing matters when both fire. Restore to a separate target first, verify the data there, then cut traffic to the restored copy with the matching code version. Never restore over the only copy before verifying; never roll code during an unverified restore. Each step leaves a way back until the next step proves itself, the same checkpoint discipline as migrations.

Worked example

A fixture incident corrupts rows through a staged bad job while a new release also serves. The learner tests reads per version, finds data fault plus code fault, restores data to a separate target and verifies, then rolls code to match, verifying each step. The record shows the evidence, the order, and the reason each step preceded the next.

Common wrong move

Restoring over production without verifying on a separate target first. A bad restore destroys the evidence and the fallback together; the second attempt starts from less than the first. Separate target, verification, then cutover, every time.

Quick check

An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.

Lesson feedback

No published feedback yet.

Log in and complete the lesson to leave feedback.

Exercise

On a staged fixture incident with both bad data and bad code, decide from evidence, restore to a separate target and verify, then roll code to match.

Pass criteria

The record shows the evidence behind the decision, the verified separate-target restore, and the sequenced code rollback with reasons.

Sources

Log in to track progressFree account: stores only your lesson progress and quiz results.