Latency Budget
Updated · Tech checked
The maximum acceptable response time for a workflow, allocated across its stages (retrieval, generation, network).
Declared in the brief (e.g. p95 < 3s interactive), measured end-to-end, enforced by alerts. Spend it where users feel it - time-to-first-token often matters more than completion time; see LLM latency & cost.