DevOps Practitioner · Module 5: Ansible and configuration management
Secrets and Change Safety on Machines You Own
Machines you own still deserve production habits. Secrets arrive by reference at run time, risky changes roll out in stages behind previews, and the last known good is one reapply away.
10 min reading
Objectives
- Keep secret content out of playbooks, variables and history
- Reference secrets from outside the repo at run time
- Gate risky changes with check mode, limits and staged rollouts
- Roll back a bad config change by reapplying the last known good
Why this matters
A password sits in a group variables file, committed, copied to every checkout. A new teammate clones the repo on a personal laptop and the secret now lives wherever that laptop travels. The M11 rule repeats here because the temptation repeats here: secret content in version control is a breach with a delayed announcement, whatever the tool. Reference at run time from the operator's environment or a vault agent; the repo holds names, never bytes.
Concepts
Secret handling has three motions. Store outside the repo (encrypted vault files with keys outside version control, or a secrets manager the runner reads at execution). Reference by name in playbooks and roles, so tasks and templates work with variables whose values arrive at run time. Rotate through overlap where the service allows it, then revoke at the source; deleting a vault entry whose password still works somewhere revoked nothing.
Change safety stacks four gates. Check mode with diff previews the blast radius. Limits restrict the run to one canary target first. Staged rollout moves group by group with verification between, exactly the delivery patience M14 teaches for releases. Reapply of the last known good is the rollback: because runs are idempotent, reapplying the previous variables restores the previous state without special machinery, and the L37 lab proves it.
Audit closes the loop. Every run logs who targeted what with which playbook version; the log is the answer to the post-incident question of what changed. Runs against shared lab machines get announced; runs against learner-owned targets still get recorded. Unlogged change is indistinguishable from sabotage after the fact.
Worked example
A demo role manages a service config with a secret reference. The learner runs it against a canary target with check mode first, applies, verifies, then rolls to the group. A bad variable is staged, the canary check catches it, and reapply of the previous variables restores service. The repo search at the end proves zero secret bytes throughout.
Common wrong move
Encrypting secrets into the repo and skipping the key, rotation and revocation story. The blob without the machinery is decoration that feels like safety. References plus practiced rotation first; machinery when scale demands it.
Quick check
An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Lesson feedback
No published feedback yet.
Log in and complete the lesson to leave feedback.
Exercise
Apply a demo role to a canary learner-owned target with check mode first, stage a bad variable, catch it at the canary, and restore by reapplying the previous variables.
Pass criteria
The record shows the canary preview, the caught fault, the restore with service verified, and a repo search proving zero secret bytes.