DevOps Professional · Module 2: Reliability, capacity and resilience · Lab
Bound Retry Storm Between a Local Service Pair
50 min hands-on · Core
A local machine with a caller and a downstream fixture service; fault injection limited to the fixture downstream. No shared systems touched.
Local guide: run the steps below on your own machine in order, then check the validation list.
Objectives
- Observe naive-retry meltdown with quoted numbers
- Apply budget, backoff with jitter, deadline and breaker
- Show contained behavior with automatic recovery quoted
Step 1
Melt it down naively
Run the pair with immediate unlimited retries, fault the downstream, and quote the meltdown: retry multiplication, saturation on both sides, recovery only when the fault clears.
Step 2
Bound the pair
Set a retry budget, exponential backoff with jitter, an overall deadline and a breaker with a fallback. Refault identically and quote retry rate within budget, bounded p99, visible sheds, breaker transitions.
Step 3
Prove recovery
Clear the fault and quote automatic recovery with times. Record the four bounds and their settings as the deliverable configuration.
How to confirm it worked
- Meltdown quoted with multiplication numbers
- Contained run quoted within budget with bounded p99
- Recovery quoted with times after fault clears
- Bound settings recorded as configuration