DevOps Professional · Module 2: Reliability, capacity and resilience
Failure Domains and the Limits of HA
High availability is survival of named failures, not a feeling of safety. Spread across real domains, rehearse the failures you claim to survive, and state plainly what the rehearsal did not prove.
10 min reading
Objectives
- Explain failure domains: process, node, zone, region, provider
- Spread workloads across domains you actually have, not ones you imagine
- Measure recovery under controlled failure injection locally
- State the HA honesty limit: the lab does not prove production survival
Why this matters
An architecture diagram shows three zones and claims zone survival, but both replicas schedule onto one node pool in one zone because affinity was never set and the data lives on a single disk. The first zone event takes everything, and the diagram was fiction. HA claims need grounding: which domains exist here, what runs in each, what was actually killed in rehearsal. Unrehearsed HA is a hope with boxes and arrows.
Concepts
Failure domains nest: process (restart), node (drain, reboot), zone (power, network partition), region (rare, total), provider (rarest, existential). Spreading means anti-affinity across nodes minimum, zones where they exist, and data replication matching the compute spread (compute in three zones with data in one is one-zone HA wearing a costume). Each spread decision names the failures it survives and the ones it does not; the document is the claim, rehearsal is the evidence.
Failure injection rehearses locally and safely: kill processes, drain nodes, partition the lab network, corrupt the fixture data path. The L54 lab measures recovery (detection time, failover time, data loss window) under controlled injection and writes the numbers down. Chaos without a hypothesis is vandalism; every injection states what should happen, then checks. Blast radius stays in the lab: never inject into shared, borrowed or production systems, ever.
The honesty limit is load-bearing and printed in every record. Local rehearsal proves the mechanism (detection works, failover engages, recovery completes) and proves nothing about production survival: no real zones, no real partitions, no provider events, no correlated failures at scale. Claiming HA from lab numbers is fiction; carrying rehearsed mechanisms plus honest limits into production design is the actual skill.
Worked example
A fixture pair runs across two local failure domains with anti-affinity and replicated fixture data. The learner kills one domain, measures detection plus failover plus loss window, and restores. Then the same kill with data on one disk shows the loss the spread did not cover. Both numbers quoted, both limits stated, no survival claimed beyond what ran.
Common wrong move
Counting replicas as availability. Three replicas on one node, one zone, one disk is one failure domain with extra processes. Spread is about domains, not counts.
Quick check
An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Lesson feedback
No published feedback yet.
Log in and complete the lesson to leave feedback.
Exercise
On a local setup, kill one failure domain of a spread fixture pair, measure detection, failover and loss, and write the numbers with the honesty limit stated.
Pass criteria
The record shows the spread design with its named survivals, the quoted recovery numbers, and the explicit statement of what was not proven.