Measuring outcomes, not output, at the customer
Updated

Shipped lines versus changed weeks: picking the metric the customer would defend, and reviewing it on a date.
Nobody renews a contract because you shipped 40 story points. They renew because Tuesday got shorter. Measure Tuesday.
The short answer
One outcome metric per slice, defined before go-live, reviewed on a calendar date. Output counts stay in the appendix.
Picking the metric
| Output (appendix) | Outcome (readout) |
|---|---|
| Pages shipped | Lookup minutes per case |
| Pipelines built | Hours of manual reconciliation removed weekly |
| Model accuracy | Escalations avoided per month |
| Dashboards delivered | Decisions made from the dashboard per week |
| Uptime percent | Failed customer operations per month |
Worked example: the metric that saved a project
A fictional integration (fictional) misses its latency target but cuts manual handling from 6 hours to 40 minutes weekly. The readout leads with the 5+ recovered hours, names the latency gap with a fix date, and the customer expands scope. Honest numbers compound; vanity numbers get audited.
Checklist: measurement discipline
- Metric, threshold, date and method in the brief.
- Baseline measured before the change, not reconstructed after.
- First review within 14 days of go-live.
- Misses reported with cause and next slice, never hidden.
Related reading
Straight answers
Frequently asked questions
What is the difference?
Output is what you shipped; outcome is what changed in the customer's week. Only outcomes survive the readout.
Who picks the metric?
Both sides, in the brief, before building. Metrics chosen after shipping get chosen to flatter.
What if the number misses?
Report it with the cause and the next slice. A missed metric with a plan beats a vanity metric every time.