Writing evaluation scorecards for retrieval systems

Updated

Test papers being graded with a pen

Answer quality you can trend: the question set, the rubric, and the weekly habit that keeps assistants honest.

Unevaluated assistants drift. A scorecard with the same questions every week turns drift into a visible line.

The short answer

Fixed question set, four-axis rubric, weekly runs with a named reviewer, trend in the readout. No new release without a score.

The rubric

AxisScoresFails when
CorrectnessAnswer matches the sourceContradicts the cited document
CitationEvery claim traces to a sourceUncited facts presented as known
Access respectZero forbidden rows surfacedAny forbidden document cited
RefusalSays unknown with a next stepInvents instead of declining

Worked example: the scorecard that blocked a release

A fictional assistant (fictional) scores 34 of 40 in week one, with two access probes failed: executive-only documents surfaced to team users. The release waits; scopes tighten; week three scores 38 with zero access failures. The blocked release becomes the trust story. The permissions post fixes the scopes; the RAG decision post uses the trend to decide on training.

Checklist: eval discipline

  1. Same questions weekly, versioned with dates.
  2. Adversarial access probes in every run.
  3. Named reviewer signs the trend before releases.
  4. Failures name the fix and the re-test date.

Straight answers

Frequently asked questions

How many eval questions are enough?

Start with 30 to 50: frequent questions, adversarial access probes, and known-hard cases. Grow from logs, not from imagination.

Who grades the answers?

A fixed rubric plus a named human reviewer weekly. Rubrics score citations, refusal behavior and access respect; humans confirm the trend.

What is a good pass threshold?

One you set before measuring, tied to the readout: for example 9 of 10 frequent questions cited and correctly scoped. Thresholds chosen after measuring flatter.

Bu sayfanın Türkçesi