RAG Evaluation Before Deployment: Measure or It Isn't Ready
Updated · Tech checked
A retrieval system ships only after three measured layers: retrieval quality (does the right chunk surface), answer groundedness (does the response cite what supports it), and refusal correctness (does it stay silent out of scope) - each on a test set built from real queries, not vibes.
The short answer
"ChatGPT over our docs worked great in the demo" is not evaluation. Evaluation = fixed test set + defined metrics + threshold decisions written down before launch. Three layers, in order:
Layer 1 - Retrieval quality
- Build the golden set: 50-200 real questions (mine support tickets/FAQ logs) each tagged with the chunk(s) that should be found.
- Metrics: hit rate@k (was a golden chunk in top-k), MRR if ranking matters, and per-role breakdown when ACL filtering exists (filtering can silently gut recall).
- Cheap first: BM25 baseline before any vectors. If your fancy pipeline can't beat keyword search on your data, you just saved the vector DB budget - that's a finding, publish it.
Layer 2 - Groundedness
For each answer: does every claim map to a retrieved, ACL-allowed source? Automate a first pass (citation coverage), then human-review a sample weekly. Track hallucination rate separately - it's the number executives actually fear.
Layer 3 - Refusal correctness
Queries outside scope, outside the user's permission, or with no answer in the corpus must refuse cleanly. Measure refusal precision/recall as a first-class metric. The worst RAG failure isn't a wrong answer - it's a confident wrong answer the user acts on.
The scorecard that goes in the brief
| Metric | Baseline (BM25) | Pipeline v1 | Threshold |
|---|---|---|---|
| hit@5 | 62% | 81% | ≥75% |
| groundedness | - | 88% | ≥85% |
| hallucination | - | 1.4% | ≤2% |
| refusal precision | - | 96% | ≥95% |
Numbers illustrative - yours come from your set. The point: a table with thresholds agreed in advance, like the success metrics method.
Ongoing, not once
Prompts, models, and corpora drift. Re-run the set on every change (CI), sample production answers weekly, and treat metric drops as incidents. The POC-to-production checklist slots this in as items 1-6.
Continue: Permission-aware RAG project · Production delivery