Writing evaluation scorecards for retrieval systems
Updated

Answer quality you can trend: the question set, the rubric, and the weekly habit that keeps assistants honest.
Unevaluated assistants drift. A scorecard with the same questions every week turns drift into a visible line.
The short answer
Fixed question set, four-axis rubric, weekly runs with a named reviewer, trend in the readout. No new release without a score.
The rubric
| Axis | Scores | Fails when |
|---|---|---|
| Correctness | Answer matches the source | Contradicts the cited document |
| Citation | Every claim traces to a source | Uncited facts presented as known |
| Access respect | Zero forbidden rows surfaced | Any forbidden document cited |
| Refusal | Says unknown with a next step | Invents instead of declining |
Worked example: the scorecard that blocked a release
A fictional assistant (fictional) scores 34 of 40 in week one, with two access probes failed: executive-only documents surfaced to team users. The release waits; scopes tighten; week three scores 38 with zero access failures. The blocked release becomes the trust story. The permissions post fixes the scopes; the RAG decision post uses the trend to decide on training.
Checklist: eval discipline
- Same questions weekly, versioned with dates.
- Adversarial access probes in every run.
- Named reviewer signs the trend before releases.
- Failures name the fix and the re-test date.
Related reading
Straight answers
Frequently asked questions
How many eval questions are enough?
Start with 30 to 50: frequent questions, adversarial access probes, and known-hard cases. Grow from logs, not from imagination.
Who grades the answers?
A fixed rubric plus a named human reviewer weekly. Rubrics score citations, refusal behavior and access respect; humans confirm the trend.
What is a good pass threshold?
One you set before measuring, tied to the readout: for example 9 of 10 frequent questions cited and correctly scoped. Thresholds chosen after measuring flatter.