Case Study: When the Rules Engine Beat the AI

Updated

A fictional insurer's claims triage 'AI initiative' meets a deterministic classification problem - the FDE's best move is proving a 400-line rules service wins on accuracy, cost, latency and auditability, and shipping that instead.

Fictional training case, not a real engagement.

Situation

"Oakline Mutual" (fictional) wants AI triage for insurance claims: classify into 6 categories, route to teams. A consultant demo'd an LLM pipeline: 84% accuracy, ~€0.11/claim, 2.3s p95.

The delivery arc

  1. Data first: 24 months of labeled claims - 31,000 rows, 6 classes, 97.2% majority-class coverage of unambiguous cases.
  2. Baseline: a rules service from the classification memo (claim keywords, form codes, vendor codes): 96.8% accuracy on the historical set, €0/claim, 40ms, fully auditable.
  3. The LLM comparison, honestly: the model wins only on the ambiguous 2.8% - and mostly by guessing. Improvement plan: rules for the 97%, "needs human" triage for the ambiguous tail, LLM excluded (revisited only if volume or variance changes).
  4. Politics: the AI mandate came from the top; the FDE's job was making "no AI" a defensible answer - the comparison table, the audit trail argument, and the €0 marginal cost did it.
  5. Shipped: rules service + human queue for the tail; consultant's pipeline archived with a thank-you memo.

What learners should extract

  • The strongest FDE move is sometimes not building the AI thing (portfolio case).
  • Baselines aren't bureaucracy - they're decision protection.
  • "Needs human" is a legitimate routing category, not a failure.

Practice version

The AI decisions practice task hands you a similar labeled dataset summary and asks for the recommendation memo.

We use Google Analytics to count visits. No ads, no cross-site tracking. Cookie Policy