The most detailed free FDE + DevOps library: 140+ lessons, 70+ labs and 80 long-form articles, in English and Turkish. Start learning →

Wrong probes cut healthy traffic: tune them

Updated

Satellite dishes on a rooftop against clouds

Aggressive probes turn slow into dead: separate readiness from liveness and tune thresholds to real timings.

Health probes are load balancers with kill authority, and wrong settings turn a slow afternoon into a massacre. Liveness too tight kills containers that only needed patience; readiness too loose feeds traffic to the sick. Tuning starts from measurements, not defaults.

Separate, then measure

Two questions, two probes. Readiness: should this instance take traffic right now. Liveness: is this instance beyond saving. Different questions, different endpoints, different thresholds.

Thresholds from timings. Measure normal response times, then set periods as multiples with a slow-start allowance after deploys. Copy-pasted defaults encode somebody else's latency.

Fail safe under load. When in doubt, readiness sheds and liveness waits. Removing an endpoint sheds load; killing a container adds restart storms to the load.

Worked example: a fictional probe massacre

The context below is fictional. Fictional service ParcelTrack (fictional) sets liveness to fail after two slow responses. One slow dependency afternoon, liveness kills half the fleet; restarts thunder against the same slow dependency; the incident doubles.

The fix separates probes: strict readiness removes slow instances from rotation, patient liveness only restarts the truly stuck. Thresholds come from the week's timing data. Next slow afternoon: degraded but alive.

Checklist: probes that protect

  1. Readiness and liveness are different checks.
  2. Thresholds derived from measured timings.
  3. Slow-start allowance after every deploy.
  4. Load sheds via readiness, never via killing.

Straight answers

Frequently asked questions

Same endpoint for both probes?

Almost never. Readiness answers can I take traffic; liveness answers am I dead. One endpoint serving both confuses slow with dead.

What thresholds are sane?

Derived from measured timings: slow-start allowance, then periods several times the normal response. Defaults are guesses about someone else's app.

Probes failing under load only?

Classic too-tight liveness. Loosen liveness, keep readiness strict: shed load by removing endpoints, not by killing containers.

Bu sayfanın Türkçesi

Turn reading into a credential

This post is a free field note. Exams run at dated sittings in 15-seat classes; one price covers one attempt. All lessons are free.