DevOps Practitioner · Module 1: Kubernetes workload management
Scheduling: Requests, Limits, Taints and Affinity
Scheduling is a capacity market plus a set of placement rules. Requests buy guaranteed seats, limits cap the feast, taints repel, affinity attracts, and Pending tells you exactly which rule or shortage said no.
10 min reading
Objectives
- Explain how requests reserve capacity and limits cap bursts
- Read a Pending pod and separate resource shortage from taint and affinity causes
- Apply taints with matching tolerations to reserve nodes for a workload class
- Use affinity to express placement preference without hard-coding node names
Why this matters
New pods sit Pending on a Friday and the team adds nodes until the bill doubles, but the pods stay Pending. The cause was never capacity: a node taint from a drained GPU pool repelled every pod without the matching toleration, and the scheduler said so in the pod events from the first minute. Pending is a decision with a written reason, not a mystery. Reading the scheduler's reason before spending money is the whole skill, and it takes one command.
Concepts
Requests are reservations: the scheduler places a pod only on a node whose unreserved capacity covers the sum of its containers' requests. Limits are ceilings enforced at runtime: a container bursting past its memory limit is killed, past its CPU limit is throttled. Requests without limits let one container starve its neighbors; limits without measured requests either waste seats or invite eviction. Measure first (lesson M05 established the habit), then set requests near measured typical use and limits with stated headroom.
Taints repel pods unless the pod tolerates them. Draining, dedicated pools and dirty nodes all express themselves as taints: NoSchedule keeps new pods off, NoExecute evicts the intolerant. Affinity is the mirror: node affinity attracts pods to labeled nodes (prefer or require), pod affinity and anti-affinity co-locate or spread relative to other pods. Prefer soft rules (preferredDuringScheduling) for spreading across zones in small clusters where hard rules would leave pods unschedulable.
Diagnose Pending in order: read the pod events for FailedScheduling and its reason (Insufficient cpu/memory, node taints, affinity mismatch); check node capacity and current reservations; then inspect taints and affinity terms. Three distinct causes, three distinct fixes: right-size requests, add matching tolerations, or relax the affinity term. The L27 lab drills exactly this separation so the fix matches the cause.
Worked example
A local two-node cluster receives three pods: one with requests larger than either node (Insufficient memory), one repelled by a taint the learner applies (Toleration missing), one with a hard node affinity to a label that exists nowhere (No nodes matched). The learner reads each FailedScheduling reason, fixes each cause differently, and watches all three schedule. Same symptom, three diagnoses, no new nodes bought.
Common wrong move
Setting requests to zero so pods schedule anywhere, then wondering why nodes run hot and the kernel kills processes. Zero requests is not free capacity, it is an unmeasured reservation that the scheduler cannot see. Unschedulable-looking clusters are usually mismeasured, not small.
Quick check
An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Lesson feedback
No published feedback yet.
Log in and complete the lesson to leave feedback.
Exercise
On a local cluster, create one Pending pod per cause (oversized request, missing toleration, impossible affinity), quote each scheduler reason, fix each, and show all running.
Pass criteria
The record holds three scheduler reasons quoted from events, three distinct fixes, and all pods running with no added nodes.