DevOps Foundations · Module 7: Operations, observability and backup
Backups That Actually Restore
Nobody cares that the backup job ran; everybody cares that the restore works. This lesson treats the restore as the product and the backup as its raw material.
10 min reading
Objectives
- Define what is backed up, how often, and how long copies live
- Restore to a separate wiped target and verify with checksums
- Explain RPO and RTO in your own service's numbers
- Test restores on a schedule instead of trusting backup jobs
Why this matters
A team discovers during a real incident that backups ran nightly for a year and restores fail on a corrupted archive from month two, because nobody ever restored anything. A backup nobody tested is a hope with a schedule. The restore is the only operation that matters, so it gets the calendar slot, the separate target, and the checksum, while the backup job is just the setup step.
Concepts
Scope first: what data, which databases and volumes, plus the configuration needed to use it (schema versions, credentials location, app version that reads the format). A data backup without its restore context is a puzzle for the worst day. Full plus incremental chains trade space for speed; the chain is only as good as its oldest link, so verify whole chains, not just the latest file. Retention encodes promises: daily for a week, weekly for a month, or whatever the recovery promises require, written down.
RPO and RTO are the contract numbers. Recovery Point Objective: how much data loss is acceptable, which sets backup frequency (hourly backups for a one-hour RPO). Recovery Time Objective: how long the restore may take, which sets tooling and practice (a four-hour RTO with an untested six-hour restore is fiction). Derive both from the service's needs, then prove the numbers with a timed restore instead of asserting them.
The restore drill is the product demo. Wipe a separate target, restore the fixtures, verify checksums against the source, and start the application against the restored data. Separate target matters: restoring over the original proves nothing about disaster and risks the only copy. Automate the drill where possible, calendar it where not, and keep the last three drill reports; a passing streak is the only honest backup metric.
Encryption and access close the loop. Backups concentrate every secret the service ever held, so they encrypt at rest with keys outside the backed-up system (a key stored next to the backup is decoration), and restore access is logged and limited. An unencrypted backup archive is a full data breach waiting for one stolen disk.
Worked example
Fixture data backs up nightly; the drill restores to a wiped directory and compares checksums. The first drill fails: the archive restores but the app rejects the schema because migrations ran since the backup and the restore procedure never recorded the schema version. Expected reading: the backup was fine and the restore broken, exactly the class of fault only drills find. The repair records schema version with every backup and replays migrations to the recorded point during restore. Verify with three consecutive green drills, each timed against the RTO. Lab L20 requires this full loop with checksum proof.
The common wrong move
Backing up the container or VM image instead of the data plus restore procedure. Images capture a moment of a running system with partial writes and no restore semantics; recovery becomes forensic boot repair instead of a practiced procedure. Back up data with its context, practice the restore, time it, and let images be build artifacts from the M06 pipeline, not recovery artifacts.
Lab and next step
Lab L20 backs up small fixture data, restores to a separate wiped target, and verifies checksums, timed against a stated RTO. Next, lesson 4 writes the runbook someone else can actually run, then proves it by handing it over.
Quick check
An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Lesson feedback
No published feedback yet.
Log in and complete the lesson to leave feedback.
Exercise
Define RPO and RTO for one dataset you own, then run a restore drill to a separate wiped target with checksum verification. Time it against your RTO and write the drill report with what broke.
Pass criteria
RPO/RTO stated in the service's numbers; drill run on a separate target with checksums; timed result compared to RTO with breakage recorded.