DevOps Practitioner · Module 1: Kubernetes workload management
Probes That Tell the Truth: Readiness, Liveness and Startup
Probes are the cluster's only window into whether your process is actually working. A probe on the wrong endpoint lies, and the cluster acts on the lie: traffic cut, pods restarted, rollouts stuck.
10 min reading
Objectives
- Assign each probe its job: readiness gates traffic, liveness restarts, startup buys warmup time
- Explain how a wrong readiness probe cuts traffic and a wrong liveness probe causes restart loops
- Probe a real dependency path instead of a process that is always alive
- Read probe failure events and name the misconfigured field
Why this matters
A service passes its checks for months, then a deploy routes all traffic to pods still warming their caches, and every request times out for three minutes. The readiness probe hit the root path, which answers 200 the instant the process starts, long before the service can do useful work. The probe told the truth about the process and lied about readiness, and the Service believed it. Every probe incident is this shape: the endpoint answers a different question than the cluster asks, and automation acts on the wrong answer at machine speed.
Concepts
Three probes, three questions. Readiness asks whether this pod should receive traffic right now; failing readiness removes the pod from Service endpoints without restarting it. Liveness asks whether the process is stuck beyond recovery; failing liveness restarts the container. Startup asks whether a slow-starting container is still warming; while startup fails, liveness stays paused, which protects warmups that exceed the liveness timeout. A pod with only liveness has no traffic gate; a pod with only readiness never recovers from deadlock. Most services need readiness plus liveness, and slow starters add startup.
Probe the dependency path, not the process. An HTTP check against a health endpoint that opens the database connection proves readiness; a TCP check against the listening port proves only that the socket exists. Timeouts and thresholds encode patience: too eager and normal jitter restarts healthy pods, too generous and real failures linger. Failure events name the probe, the endpoint and the counts, so diagnosis starts at the event text, not at the application log.
Liveness must never depend on an external service. If the database goes down and every pod's liveness checks the database, the cluster restarts the whole fleet instead of waiting, turning a downstream outage into a crash loop. Readiness may reflect dependencies (no traffic until the dependency answers); liveness reflects only the process itself.
Worked example
A demo service exposes /alive (always 200) and /ready (200 only after a thirty-second warmup with the fixture store reachable). The learner first wires both probes to /alive and watches a rollout send traffic to cold pods; then rewires readiness to /ready with a startup probe covering the warmup, and watches the same rollout wait. The L26 lab replays this as a repair: traffic cut by a wrong probe, fixed by probing the true path.
Common wrong move
Copying probe blocks between services without changing the endpoint. Every service gets the same /healthz check, half the services have no such route, and the cluster restarts pods that were serving fine. A probe is a claim about a specific application path; pasted claims are false until verified against the running code.
Quick check
An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Lesson feedback
No published feedback yet.
Log in and complete the lesson to leave feedback.
Exercise
On a local cluster, give a demo service a readiness probe on a path that answers before warmup completes, show traffic reaching a cold pod, then fix the probe to the true ready path and show the Service waiting.
Pass criteria
The record quotes the cold-traffic evidence, the corrected probe definition, and endpoint output showing no traffic before ready.