DevOps Professional · Module 3: Migration, restore and disaster recovery
RPO and RTO: Numbers Before Promises
Recovery promises are numbers or they are fiction. RPO sets how much data loss is survivable, RTO sets how long darkness may last, and only tested restores on separate targets turn either number into a promise.
10 min reading
Objectives
- Define RPO (how much data loss) and RTO (how long down) as business numbers
- Derive backup frequency from RPO and restore drills from RTO
- Test restores on a separate target on schedule, not on faith
- Refuse promises the drills do not support
Why this matters
The backup runs nightly and green, and the restore fails for fourteen months straight because nobody ever ran one: wrong credentials, moved paths, a format the current tool no longer reads. The outage that needs the backup discovers all of this at once. Untested backups are the industry's most common fiction, and the green backup dashboard is its cover art. The L57 lab diagnoses a broken backup and verifies the good one on a separate target; this lesson makes that drill the policy.
Concepts
RPO (recovery point objective) is the maximum survivable data loss, stated in time: fifteen minutes means backups (or replication lag) never older than fifteen minutes. RTO (recovery time objective) is the maximum survivable darkness: two hours means detected, decided, restored and verified inside two hours. Both are business numbers first (what does an hour of orders cost, what does a day offline do), then engineering budgets: backup frequency derives from RPO, restore speed plus decision time derives from RTO.
Restore drills prove the numbers on schedule. Full restore to a separate target (never the live system first), application verification against the restored data, timed end to end. Separate target matters: restoring over production tests nothing and risks everything, and a drill that shares fate with production shares its failures. Drill frequency follows the promise: monthly minimum for anything load-bearing, more often when the data path changes.
Broken drills are findings, not embarrassments. Each failure names its cause class (credentials, paths, format drift, capacity, runbook rot) and gets an owned fix with a re-drill date. The drill log is the evidence behind every RPO/RTO promise; a promise without a recent passing drill is a draft, and drafts are labeled as such to anyone who asks.
Worked example
A fixture backup pair runs: one backup restores cleanly to a separate target inside its RTO with data inside its RPO, quoted end to end; the other fails on a staged credential fault, diagnosed from the drill log, fixed, and re-drilled passing. The record shows both arcs: the proof and the repair.
Common wrong move
Counting backup success as recovery readiness. Backups succeed daily in systems that cannot restore; the dashboard measures the wrong half. Readiness is restores, timed, on separate targets, on schedule.
Quick check
An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Lesson feedback
No published feedback yet.
Log in and complete the lesson to leave feedback.
Exercise
Run a fixture restore drill to a separate target timed end to end, verify the data against RPO, and record the drill log with the next due date.
Pass criteria
The record shows the timed restore, the RPO verification quoted, and the drill log with findings and the next due date.