The most detailed free FDE + DevOps library: 140+ lessons, 70+ labs and 80 long-form articles, in English and Turkish. Start learning →

Healthchecks and graceful shutdown: die well

Updated

A doctor adjusting a stethoscope during a checkup

Dead-but-running is the worst container state: a real healthcheck plus SIGTERM handling fixes it.

The worst container state is not crashed. Crashed restarts. The worst is dead-but-running: the process answers the orchestrator while serving nothing, collecting traffic it cannot handle. Two mechanisms prevent it: a healthcheck that tells the truth, and shutdown code that finishes work fast.

Check truthfully, shut down fast

Healthchecks must test real work. A check that only proves the process exists blesses zombies. Check the dependency the container actually needs: can it answer, reach its data, accept one request.

Handle SIGTERM explicitly. The orchestrator asks nicely before it kills. Finish in-flight requests, close connections, then exit with a clear code. Log the shutdown path so restarts are explainable later.

Worked example: a fictional zombie checkout

The context below is fictional. Fictional container ParcelTrack (fictional) serves checkout with a healthcheck that hits the root path, which answers even when the payment connection is dead. After a payment outage, all containers stay green while checkout fails.

The fix has two halves: the check now performs one cheap payment-gateway ping, and the app handles SIGTERM by draining for five seconds then exiting. Next payment incident: sick containers leave rotation in seconds, deploys stop blocking on hung shutdowns.

Checklist: a container that dies well

  1. Healthcheck exercises the real dependency.
  2. SIGTERM drains briefly, then exits cleanly.
  3. Shutdown path logged with reason and duration.
  4. Liveness cannot kill a merely busy container.

Straight answers

Frequently asked questions

Liveness or readiness, which first?

Readiness first: stop sending traffic to sick containers. Liveness restarts the truly dead. Wrong liveness settings kill healthy containers under load.

Why does my container ignore SIGTERM?

The app never sees the signal: wrong PID 1, missing handler, or a shell wrapper swallowing it. Exec form plus an explicit handler fixes most cases.

How long may shutdown take?

Seconds, with a known timeout. Draining connections is good; draining forever blocks every deploy.

Bu sayfanın Türkçesi

Turn reading into a credential

This post is a free field note. Exams run at dated sittings in 15-seat classes; one price covers one attempt. All lessons are free.