Measuring Latency and Cost in LLM Applications: Budgets, Not Anecdotes
Updated · Tech checked
Measure per-workflow: latency as p50/p95/p99 against a declared budget, cost as tokens×price plus infrastructure, both tagged by feature and model version - reviewed weekly, alerted on drift, and reported to the customer in plain units.
The short answer
LLM apps die financially, not technically: a demo that costs €0.40/request is a monthly invoice that ends the project. The fix is instrumentation with budgets declared in advance - latency percentiles per workflow, cost per request and per month, tagged by feature and model version.
Latency: percentiles or it didn't happen
- Track p50 / p95 / p99 per user-facing workflow, end-to-end (retrieval + generation + network), not just model time.
- Streaming counts: time-to-first-token is often the UX number that matters.
- Declare the budget in the brief ("p95 < 4s for interactive; batch exempt") - the success metrics pattern applies to non-functional promises too.
Cost: three numbers
- Cost per request = (input tokens × in-price) + (output tokens × out-price) + infra share. Log it on every call.
- Cost per workflow execution (a single user job may chain calls - measure the chain).
- Monthly projection at observed volume, with the volume assumption printed. A projection without volume is a guess; label it.
Tagging discipline
Every call logs: feature, model_version, prompt_version, user_tier. Without tags you'll know the total doubled and nothing else. With tags, drift answers are one query: "v4 prompt doubled output tokens."
The control loop
- Alert on drift: cost/request +30% or p95 +30% vs 7-day baseline = investigate (model silently upgraded? prompt regression? retry storm?).
- Optimize in order: cache (content-hash), shorten prompts, tier models (cheap first, escalate on confidence), batch non-interactive work. Biggest lever is usually "do you need the big model at all" - the AI-was-wrong-tool pattern applied to budgets.
- Report monthly to the customer in their units: cost per resolved ticket, per document processed. Cost transparency is a trust feature - the production checklist makes it a gate.
Continue: POC-to-production checklist · Rollback runbook