FDE Foundations · Module 7: RAG & Retrieval
RAG Evaluation Scorecards
A RAG system without a scorecard is a demo. Build 30 to 50 real questions with reference answers, score retrieval and generation separately, and publish the numbers.
11 min reading
Objectives
- Build a question set with graded reference answers
- Score retrieval and generation separately
- Publish the scorecard and re-run it on every change
The question set
Collect real questions from operators, not invented ones. For each: the question, the reference answer, the target documents, and a grader (exact fields for factual answers, rubric lines for open ones). Thirty good questions beat three hundred synthetic ones; volume without realism measures nothing.
Two layers, two scores
Retrieval score: did the target documents appear in the candidate set (recall at k) and near the top (precision)? Generation score: is the answer correct, cited, complete, and free of fabrication? Judge generation with a rubric and, where stakes allow, a second pass by a model grader with human spot checks. Keep the two scores separate: they have different fixes and different owners.
The scorecard
One table, versioned: question count, retrieval recall, answer accuracy, citation coverage, fabrication incidents. Re-run on every change to chunking, models, or prompts, and publish it to the customer. The scorecard is also your regression alarm: a change that improves one number while sinking another gets caught before users catch it.
"It seems better" is not an evaluation. If the scorecard did not run, the change did not ship.
Quick check
An optional 2-3 question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Exercise
Build a mini scorecard for a fictional HR-policy assistant: ten questions (list titles), the metrics you will compute, plausible baseline numbers, and the re-run rule.
Pass criteria
Ten questions listed with topics, both retrieval and generation metrics defined, baseline numbers present, and a re-run trigger tied to specific change types.