Case Study: When the Rules Engine Beat the AI
Updated
A fictional insurer's claims triage 'AI initiative' meets a deterministic classification problem - the FDE's best move is proving a 400-line rules service wins on accuracy, cost, latency and auditability, and shipping that instead.
Fictional training case, not a real engagement.
Situation
"Oakline Mutual" (fictional) wants AI triage for insurance claims: classify into 6 categories, route to teams. A consultant demo'd an LLM pipeline: 84% accuracy, ~€0.11/claim, 2.3s p95.
The delivery arc
- Data first: 24 months of labeled claims - 31,000 rows, 6 classes, 97.2% majority-class coverage of unambiguous cases.
- Baseline: a rules service from the classification memo (claim keywords, form codes, vendor codes): 96.8% accuracy on the historical set, €0/claim, 40ms, fully auditable.
- The LLM comparison, honestly: the model wins only on the ambiguous 2.8% - and mostly by guessing. Improvement plan: rules for the 97%, "needs human" triage for the ambiguous tail, LLM excluded (revisited only if volume or variance changes).
- Politics: the AI mandate came from the top; the FDE's job was making "no AI" a defensible answer - the comparison table, the audit trail argument, and the €0 marginal cost did it.
- Shipped: rules service + human queue for the tail; consultant's pipeline archived with a thank-you memo.
What learners should extract
- The strongest FDE move is sometimes not building the AI thing (portfolio case).
- Baselines aren't bureaucracy - they're decision protection.
- "Needs human" is a legitimate routing category, not a failure.
Practice version
The AI decisions practice task hands you a similar labeled dataset summary and asks for the recommendation memo.