DevOps Professional · Module 2: Reliability, capacity and resilience
Capacity and Bottlenecks: Measure Before You Buy
Capacity arguments without measurements are opinions with invoices. Load the system small, find the real constraint, refute the plausible-but-wrong scaling plan with numbers, and buy only what the bottleneck justifies.
10 min reading
Objectives
- Explain why bottlenecks move: fixing one reveals the next constraint
- Measure a service under small load and name the actual bottleneck
- Refute a wrong scaling proposal with measured numbers
- Size capacity from the bottleneck with headroom, not from totals
Why this matters
Traffic slows and the proposal is double the fleet: twice the machines, twice the bill, approved in one meeting. The bottleneck was a single database connection pool, and doubling the app servers doubled the contention. The invoice doubled, the latency stayed. Bottlenecks hide behind plausible stories (more traffic needs more servers), and only measurement tells the pool from the fleet. The L52 lab measures a small bottleneck and refutes a wrong proposal on paper before money moves.
Concepts
Bottlenecks are the slowest serial constraint: the pool, the single queue, the lock, the disk, the downstream with a fixed rate. Utilization laws are simple tools: the resource nearest saturation as load rises is the current bottleneck, and fixing it moves the constraint elsewhere, which is progress, not failure. Measure with small, controlled load first: enough to saturate one thing while watching everything, in an environment where saturation is safe.
Refutation is a deliverable. The wrong proposal (double the fleet, cache everything, shard now) gets a measured answer: the numbers showing the constraint elsewhere, the cost of the wrong fix, the cost of the right one. This is not politics, it is engineering communication: a one-page brief with graphs beats a hallway debate every time, and it works upward as well as sideways.
Size from the bottleneck outward. Capacity equals the bottleneck's sustainable rate times headroom for spikes and growth, plus the next constraint's distance (how soon it becomes the bottleneck). Headroom is a policy number (peak plus growth plus incident margin), written down, not a feeling. Revisit on schedule or when the bottleneck moves, which it will, and say so in the sizing document so movement surprises nobody.
Worked example
A fixture service slows under small load. The learner profiles: app CPU idle, pool exhausted, downstream latency climbing with queue depth. The bottleneck is the pool, not the fleet. The wrong proposal (double app servers) is refuted with the pool numbers; the right fix (pool sizing plus a query fix) is sized from measured throughput with stated headroom. Same slow service, measured answer.
Common wrong move
Sizing from totals: total traffic divided by per-server capacity equals server count. Totals hide serialization, contention and downstream limits; the quotient buys machines that wait on the same bottleneck faster. Bottleneck first, arithmetic second, purchase last.
Quick check
An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Lesson feedback
No published feedback yet.
Log in and complete the lesson to leave feedback.
Exercise
Load a fixture service small, profile to the real bottleneck, refute a given wrong scaling proposal in a one-page brief, and size the right fix with headroom.
Pass criteria
The record shows the profile with the bottleneck named, the refutation brief with numbers, and the sized fix with headroom stated.