DevOps Foundations · Module 5: Container basics
Health Checks, Signals and Resource Limits
A container that cannot report health, ignore SIGTERM, or fit its limits is a liability the scheduler cannot manage. This lesson makes containers good citizens under supervision.
10 min reading
Objectives
- Write healthchecks that test readiness, not just aliveness
- Handle SIGTERM so shutdowns are graceful instead of abrupt
- Set memory and CPU limits from measured numbers
- Explain what the orchestrator does on failed checks and OOM
Why this matters
Rolling updates send traffic to containers still warming up because no readiness check exists, and users meet 500s on every deploy. Shutdowns drop in-flight requests because the app ignores SIGTERM until SIGKILL. One container eats the host's memory and the OOM killer picks its victims. All three are supervision failures: the container never told the platform what healthy, done and enough mean.
Concepts
Healthchecks answer whether the container can serve, tested from inside on a real dependency path: an HTTP endpoint that queries the database beats a process-exists check that passes on a wedged app. Separate liveness (restart me if stuck) from readiness (send me traffic when warm); conflating them restarts containers that only needed thirty more seconds. Timeouts and retries on the check itself matter: a check slower than its interval queues probes until the container looks perpetually sick.
Signals end containers, and SIGTERM arrives first with a grace period before SIGKILL. Apps must catch TERM, stop accepting work, finish in-flight requests, then exit; the shell-form CMD trap from M01 lesson 2 applies here because PID 1 ignores signals by default in many runtimes. Exec-form CMD plus an init (tini or equivalent) or a proper signal handler turns ten-second kills into clean drains. Log the shutdown path: silent exits are undebuggable exits.
Limits encode measured reality. Memory limit with headroom above the observed peak plus growth; CPU shares for fairness, quotas for hard caps. No limit means one container can starve the host; a limit set from guesses means mysterious OOM kills at peak. Measure under load (the M18 method in miniature: small load, watch the peak, add margin), set, then verify the kill only fires beyond genuine overload. Requests/limits pairs in orchestrators schedule honestly only when the numbers reflect the app.
Worked example
Deploys cause user-visible errors for twenty seconds. The evidence:
HEALTHCHECK --interval=5s --timeout=3s --retries=2 CMD curl -f http://localhost:8080/ready || exit 1
with no /ready endpoint in the app, so the check always fails and the platform either never routes (stuck rollout) or routes blindly (errors for users). The repair has two halves: implement /ready to test the real dependency (database reachable, migrations applied), and set the check to match startup time (start-period covering the warm phase). Verify with a rolling update observed request by request: zero errors across the rollout, old containers draining before exit. Lab L15 proves graceful stop and limits with measurements.
The common wrong move
HEALTHCHECK CMD true, or no check with the platform defaulting to alive. Green dashboards over broken apps: traffic flows to containers that cannot serve, and the first signal of failure is user complaints. A healthcheck that cannot fail is not a check; delete it and feel the absence, then write a real one.
Lab and next step
Lab L15 verifies health, SIGTERM handling and resource limits with measured proof. From here the path continues to M06: the pipeline that builds, tests and ships these images.
Quick check
An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Lesson feedback
No published feedback yet.
Log in and complete the lesson to leave feedback.
Exercise
Add a real readiness endpoint or check to a container you run, deploy it with start-period covering warmup, and show one rolling update (or stop/start) with zero failed requests. Record the check definition and the proof.
Pass criteria
Check tests a real dependency path, not process existence; start-period covers measured warmup; zero-failure rollout shown with request evidence.