The most detailed free FDE + DevOps library: 140+ lessons, 70+ labs and 80 long-form articles, in English and Turkish. Start learning →

DevOps Professional · Module 2: Reliability, capacity and resilience

Failure Domains and the Limits of HA

High availability is survival of named failures, not a feeling of safety. Spread across real domains, rehearse the failures you claim to survive, and state plainly what the rehearsal did not prove.

10 min reading

Objectives

  • Explain failure domains: process, node, zone, region, provider
  • Spread workloads across domains you actually have, not ones you imagine
  • Measure recovery under controlled failure injection locally
  • State the HA honesty limit: the lab does not prove production survival

Why this matters

An architecture diagram shows three zones and claims zone survival, but both replicas schedule onto one node pool in one zone because affinity was never set and the data lives on a single disk. The first zone event takes everything, and the diagram was fiction. HA claims need grounding: which domains exist here, what runs in each, what was actually killed in rehearsal. Unrehearsed HA is a hope with boxes and arrows.

Concepts

Failure domains nest: process (restart), node (drain, reboot), zone (power, network partition), region (rare, total), provider (rarest, existential). Spreading means anti-affinity across nodes minimum, zones where they exist, and data replication matching the compute spread (compute in three zones with data in one is one-zone HA wearing a costume). Each spread decision names the failures it survives and the ones it does not; the document is the claim, rehearsal is the evidence.

Failure injection rehearses locally and safely: kill processes, drain nodes, partition the lab network, corrupt the fixture data path. The L54 lab measures recovery (detection time, failover time, data loss window) under controlled injection and writes the numbers down. Chaos without a hypothesis is vandalism; every injection states what should happen, then checks. Blast radius stays in the lab: never inject into shared, borrowed or production systems, ever.

The honesty limit is load-bearing and printed in every record. Local rehearsal proves the mechanism (detection works, failover engages, recovery completes) and proves nothing about production survival: no real zones, no real partitions, no provider events, no correlated failures at scale. Claiming HA from lab numbers is fiction; carrying rehearsed mechanisms plus honest limits into production design is the actual skill.

Worked example

A fixture pair runs across two local failure domains with anti-affinity and replicated fixture data. The learner kills one domain, measures detection plus failover plus loss window, and restores. Then the same kill with data on one disk shows the loss the spread did not cover. Both numbers quoted, both limits stated, no survival claimed beyond what ran.

Common wrong move

Counting replicas as availability. Three replicas on one node, one zone, one disk is one failure domain with extra processes. Spread is about domains, not counts.

Quick check

An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.

Lesson feedback

No published feedback yet.

Log in and complete the lesson to leave feedback.

Exercise

On a local setup, kill one failure domain of a spread fixture pair, measure detection, failover and loss, and write the numbers with the honesty limit stated.

Pass criteria

The record shows the spread design with its named survivals, the quoted recovery numbers, and the explicit statement of what was not proven.

Sources

Log in to track progressFree account: stores only your lesson progress and quiz results.