FDE Foundations · Module 10: Reliability
Capacity and Cost Operations
Reliability includes the bill and the peak. Size for observed peaks with headroom, treat spend as a monitored signal, and revisit capacity at every growth conversation.
9 min reading
Objectives
- Size for the peak, not the average
- Watch cost as an operational metric with alerts
- Plan headroom with the customer's growth in mind
Peak, not average
Traffic and data arrive unevenly: month-end invoice runs, Monday-morning queues, seasonal spikes. Measure the peak, not the mean, and keep headroom (the customer's tolerance decides how much: 2x is a common ask). The report that runs in 20 minutes today will run in 40 next quarter; capacity planning is a scheduled conversation, not a one-time sizing.
Cost as a signal
Track spend per day per component (compute, storage, model calls, egress). Alert on deviation from trend, not just on absolute ceilings: a 30 percent jump usually means a bug (a retry storm, a duplicated batch) before it means growth. The cost dashboard belongs in the same review as the SLO report.
The growth conversation
Quarterly, review with the customer: volumes, growth, and what breaks first at 2x, 5x. "At 5x, the nightly window overflows; the fix is partitioning, a two-week item" is exactly the kind of forward-looking statement that separates an operator from a firefighter.
Nobody remembers the outages you prevented. Write them down anyway: prevented incidents are the capacity plan's receipts.
Quick check
An optional 2-3 question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Exercise
Write the capacity plan for a fictional batch service: observed peak, headroom choice with reason, three cost alerts, and the one-line 5x statement for the customer review.
Pass criteria
Peak measured with a number, headroom justified, alerts on trend deviation as well as ceilings, and the 5x statement names the first thing to break and the fix.