Defining Success Metrics Before Building AI: The Anti-Hype Checklist
Updated · Tech checked
AI projects succeed when success is a number the business already tracks: task time, error rate, cost per ticket, deflection with quality floor. Define baseline, target, measurement source, and a kill threshold before writing a line of model code.
The short answer
"Improve support with AI" is not a metric. A metric has four parts: baseline (measured before you build), target (with a date), measurement source (the query/dashboard that produces it), and quality floor (the worst accuracy you'll tolerate while chasing the target). Add a kill threshold - the result that makes you turn it off - and you have an honest project.
The metric menu
| Business outcome | Example metric | Watch out for |
|---|---|---|
| Speed | Median handling time | Speeding up bad work |
| Cost | Cost per resolved ticket | Hiding human rework in another team |
| Volume | Deflection % (self-served) | Deflecting anger, not tickets |
| Quality | Error/misroute rate below floor | Gaming the audit sample |
| Revenue | Conversion on assisted flows | Correlation ≠ contribution |
Baselines: the part everyone skips
Measure the current process for 2-4 weeks (or mine logs) before building. "We think it's slow" is not a baseline. If the customer can't produce a baseline, your first deliverable is a measurement harness - genuinely, that's an FDE deliverable worth a week.
Setting floors and kill thresholds
Example: support routing assist -
- Target: misroute 40% → 12% in 60 days.
- Floor: never above 25% after day 30, or the assist is disabled pending redesign.
- Kill: if agent override rate exceeds 60% for two consecutive weeks, the assist isn't helping; stop and re-discover.
Write these into the brief. The RAG evaluation article covers the model-side sibling of this discipline.
Continue: Translation method · Production checklist