DevOps Professional · Module 2: Reliability, capacity and resilience
Retries and Backpressure: Bounded Help, Not Storms
Retries help the unlucky request and kill the struggling system. Bounded retries with backoff and jitter absorb transient faults; backpressure sheds the excess before the queue becomes the outage.
10 min reading
Objectives
- Explain how unbounded retries turn one failure into a storm
- Bound retries with budgets, backoff, jitter and deadlines
- Apply backpressure: shed, queue or degrade instead of collapsing
- Observe a bounded pair under fault and quote the contained behavior
Why this matters
A downstream slows, clients retry instantly and identically, and the retry flood triples the load on the already struggling service. The outage that would have been partial becomes total, caused not by the fault but by the help. Every retry storm shares one design: unbounded, immediate, synchronized retries with no shedding anywhere. The L53 lab bounds a local service pair and watches the storm stay small; this lesson explains each bound's job.
Concepts
Retries need four bounds. A retry budget caps the fraction of traffic that may be retried (a few percent, not per-request infinite). Backoff spaces attempts exponentially so the downstream gets breathing room. Jitter desynchronizes clients so retries stop arriving as a wall. A deadline (timeout plus overall attempt cap) ends the effort instead of letting one slow dependency hold every caller forever. Idempotency decides safety: retry only what repetition cannot corrupt.
Backpressure is the downstream's defense and the caller's honesty. Shed excess early with fast failure (better a quick error than a slow collapse), queue boundedly with visible depth and drop policy, degrade gracefully (cached, simplified, delayed) where the product allows. Circuit breakers formalize this: after enough failures, stop calling for a while, probe with a trickle, close when health returns. The breaker protects the caller from waiting and the downstream from drowning, at the cost of explicit fallback paths that must exist before they are needed.
Observe the pair under fault. The lab injects downstream slowness and quotes: retry rate within budget, p99 bounded by the deadline, shed count visible, breaker state transitions logged. Contained behavior is quoted numbers, not hoped adjectives.
Worked example
A local service pair runs with naive retries first: downstream fault, retry flood, both services saturated, quoted meltdown. Then the bounds land (budget, backoff with jitter, deadline, breaker with fallback): same fault, retry rate capped, errors fast and counted, recovery automatic when the fault clears. Same fault, designed behavior.
Common wrong move
Retrying non-idempotent operations on hope. Double charges, duplicate orders, repeated side effects: the retry succeeds technically and corrupts actually. Mark operations idempotent or do not retry them.
Quick check
An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Lesson feedback
No published feedback yet.
Log in and complete the lesson to leave feedback.
Exercise
Fault a local downstream, first observe naive-retry meltdown quoted, then bound the pair and quote the contained behavior with recovery.
Pass criteria
The record shows meltdown numbers, the four bounds with their settings, and contained numbers with automatic recovery quoted.