Wrong probes cut healthy traffic: tune them
Updated

Aggressive probes turn slow into dead: separate readiness from liveness and tune thresholds to real timings.
Health probes are load balancers with kill authority, and wrong settings turn a slow afternoon into a massacre. Liveness too tight kills containers that only needed patience; readiness too loose feeds traffic to the sick. Tuning starts from measurements, not defaults.
Separate, then measure
Two questions, two probes. Readiness: should this instance take traffic right now. Liveness: is this instance beyond saving. Different questions, different endpoints, different thresholds.
Thresholds from timings. Measure normal response times, then set periods as multiples with a slow-start allowance after deploys. Copy-pasted defaults encode somebody else's latency.
Fail safe under load. When in doubt, readiness sheds and liveness waits. Removing an endpoint sheds load; killing a container adds restart storms to the load.
Worked example: a fictional probe massacre
The context below is fictional. Fictional service ParcelTrack (fictional) sets liveness to fail after two slow responses. One slow dependency afternoon, liveness kills half the fleet; restarts thunder against the same slow dependency; the incident doubles.
The fix separates probes: strict readiness removes slow instances from rotation, patient liveness only restarts the truly stuck. Thresholds come from the week's timing data. Next slow afternoon: degraded but alive.
Checklist: probes that protect
- Readiness and liveness are different checks.
- Thresholds derived from measured timings.
- Slow-start allowance after every deploy.
- Load sheds via readiness, never via killing.
Related reading
- Hands on: Wrong Probe Cuts Traffic.
- The shutdown companion: Healthchecks and graceful shutdown.
Straight answers
Frequently asked questions
Same endpoint for both probes?
Almost never. Readiness answers can I take traffic; liveness answers am I dead. One endpoint serving both confuses slow with dead.
What thresholds are sane?
Derived from measured timings: slow-start allowance, then periods several times the normal response. Defaults are guesses about someone else's app.
Probes failing under load only?
Classic too-tight liveness. Loosen liveness, keep readiness strict: shed load by removing endpoints, not by killing containers.