RAG Evaluation Before Deployment: Measure or It Isn't Ready

Updated · Tech checked

A retrieval system ships only after three measured layers: retrieval quality (does the right chunk surface), answer groundedness (does the response cite what supports it), and refusal correctness (does it stay silent out of scope) - each on a test set built from real queries, not vibes.

The short answer

"ChatGPT over our docs worked great in the demo" is not evaluation. Evaluation = fixed test set + defined metrics + threshold decisions written down before launch. Three layers, in order:

Layer 1 - Retrieval quality

  • Build the golden set: 50-200 real questions (mine support tickets/FAQ logs) each tagged with the chunk(s) that should be found.
  • Metrics: hit rate@k (was a golden chunk in top-k), MRR if ranking matters, and per-role breakdown when ACL filtering exists (filtering can silently gut recall).
  • Cheap first: BM25 baseline before any vectors. If your fancy pipeline can't beat keyword search on your data, you just saved the vector DB budget - that's a finding, publish it.

Layer 2 - Groundedness

For each answer: does every claim map to a retrieved, ACL-allowed source? Automate a first pass (citation coverage), then human-review a sample weekly. Track hallucination rate separately - it's the number executives actually fear.

Layer 3 - Refusal correctness

Queries outside scope, outside the user's permission, or with no answer in the corpus must refuse cleanly. Measure refusal precision/recall as a first-class metric. The worst RAG failure isn't a wrong answer - it's a confident wrong answer the user acts on.

The scorecard that goes in the brief

MetricBaseline (BM25)Pipeline v1Threshold
hit@562%81%≥75%
groundedness-88%≥85%
hallucination-1.4%≤2%
refusal precision-96%≥95%

Numbers illustrative - yours come from your set. The point: a table with thresholds agreed in advance, like the success metrics method.

Ongoing, not once

Prompts, models, and corpora drift. Re-run the set on every change (CI), sample production answers weekly, and treat metric drops as incidents. The POC-to-production checklist slots this in as items 1-6.

Continue: Permission-aware RAG project · Production delivery

We use Google Analytics to count visits. No ads, no cross-site tracking. Cookie Policy