DevOps Practitioner · Module 8: GitOps and incident response
Drift, Manual Changes and the Way Back
Drift is information: someone believed the live change was necessary. The job is to decide quickly whether they were right, adopt it into git or revert it, and keep a way back that survives tooling outages.
10 min reading
Objectives
- Classify drift: hotfix, experiment, or accident, each with its way back
- Diagnose manual-change versus desired-state conflicts from the sync diff
- Adopt correct drift into git or revert it, with the decision recorded
- Keep an emergency path that works when git or the agent is down
Why this matters
An on-call engineer fixes production with a live edit at 2 AM and goes back to sleep. Morning brings a choice the team fumbles weekly: the agent wants to revert the fix, the engineer wants it kept, and the argument happens in chat while the sync diff sits unread. Drift disputes are routine; treating each as a crisis wastes the routine-ness. Classify, decide, record: the hotfix was right (adopt it), the experiment is over (revert it), the accident needs a guard (revert plus gate).
Concepts
Classification drives the response. Hotfix drift (live change that saved the night) gets adopted into git within the day, with the incident reference in the commit message. Experiment drift (a test someone left running) gets reverted and the owner notified. Accident drift (a typo applied live, a wrong context) gets reverted plus a guard: narrower access, a required preview, a checklist. The L47 lab diagnoses a staged conflict from the sync diff and walks one branch of each.
Adoption is a commit, not a sync-back click alone. The live state is read, translated into the source format (values, manifests), reviewed and merged; the agent then converges, confirming the adoption. Revert is the agent's normal motion with self-heal, or a deliberate sync where manual. Either way the decision and its reason enter the record: future drift of the same shape resolves by precedent instead of by argument.
The emergency path assumes the tooling is down. If git is unreachable or the agent is broken during an incident, responders fix live state directly (users before purity, always), then reconcile after: adopt or revert with the incident reference. This path is written, drilled and rare; an undocumented emergency path is just more drift with a story.
Worked example
A staged manual scale-up of a demo app conflicts with the git state. The learner reads the sync diff, classifies it as a hotfix (traffic spike documented in the scenario), adopts it into git with the reference, and watches the agent confirm. Then a staged accidental label change is reverted the same way, with the guard (a preview requirement) recorded.
Common wrong move
Auto-healing every drift without reading it. The agent reverts the 2 AM hotfix mid-incident and the outage returns wearing a new cause. Read first, especially during incidents; self-heal can wait minutes while humans classify.
Quick check
An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Lesson feedback
No published feedback yet.
Log in and complete the lesson to leave feedback.
Exercise
Stage a manual change against git state on a local setup, classify it from the sync diff, adopt or revert with the reason recorded, and quote the resolution.
Pass criteria
The record shows the sync diff, the classification with its reason, and the quoted converged state after adopt or revert.