DevOps Professional · Module 3: Migration, restore and disaster recovery
Consistency and Reconciliation Across the Move
Two systems overlap during every migration, and overlap breeds duplicates. Reconciliation queries find them, write-path fixes stop them, and continuous comparison proves the move is clean, not merely finished.
10 min reading
Objectives
- Explain what consistency means during dual-run: same reads, both sides
- Detect double-writes and duplicate records with reconciliation queries
- Fix the write path so each fact lands exactly once per system
- Keep reconciling until the differences stay zero, not just once
Why this matters
Dual-write runs for a week and the new system holds eleven percent more orders than the old. Retries wrote twice, a backfill overlapped the live stream, and nobody compared until cutover eve. The count gap is now an emergency with a deadline. Reconciliation run from day one would have caught the first duplicate within hours, when the cause was small and the fix was cheap. Comparison during the move is not overhead; it is the move's steering wheel.
Concepts
Consistency during migration means read-equivalence: the same query against either system returns the same answer. Reconciliation queries check this continuously: counts per window, checksums per shard, sampled deep comparison of full records. Differences classify into lag (new system behind, converges), duplication (same fact twice, needs dedup rules), and divergence (different facts, needs investigation). Lag is normal within bounds; the other two are defects with owners.
Double-writes come from retry without idempotency (M13's rule returns at data scale), overlapping backfill with live writes, and two writers that do not know each other. The L55 lab finds a staged double-write fault on a small dataset and fixes the write path: idempotency keys, backfill windows fenced from live traffic, single-writer discipline per fact. Dedup rules handle history (first wins, last wins, merge), stated per dataset, never assumed.
Reconciliation runs until retirement, not until the first clean report. One clean comparison proves one moment; the move needs a clean streak across load peaks, deploys and weekends. The dashboard shows the streak; cutover requires it. After retirement, one final reconciliation against the archive closes the books.
Worked example
A fixture dual-write produces duplicates from a staged retry fault. The learner's reconciliation query shows the count gap, classifies it as duplication, adds the idempotency key to the write path, dedups history by rule, and watches the streak go clean across repeated runs. The fault, the fix and the streak are all quoted.
Common wrong move
Reconciling once the night before cutover. A single clean comparison after weeks of blind dual-run proves nothing about the weeks; the gap found that night has no time left to fix. Compare from day one, cut on the streak.
Quick check
An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Lesson feedback
No published feedback yet.
Log in and complete the lesson to leave feedback.
Exercise
On a fixture dual-write with a staged duplication fault, find it with reconciliation queries, fix the write path, and show a clean streak before cut.
Pass criteria
The record shows the gap quoted, the fault classified, the write-path fix, and the quoted clean streak.