RAG or fine-tuning for customer data: an FDE decision
Updated

Start with retrieval, earn the right to train: the decision path for assistants over customer documents.
Customers ask for training because it sounds thorough. FDEs recommend retrieval first because it ships, cites and updates.
The short answer
Build retrieval with access control and evals first. Fine-tune only when measured logs prove a gap that retrieval cannot close.
The decision table
| Question | Retrieval wins when | Fine-tuning earns when |
|---|---|---|
| Freshness | Documents change weekly | Style is stable for months |
| Citations | Customer demands sources | Behavior matters more than sources |
| Access control | Per-team visibility exists | Single audience, stable corpus |
| Evidence | No evals yet | Evals show a persistent gap |
| Budget | Fixed monthly spend | Measured lift covers training cost |
Worked example: the wiki assistant
A fictional company (fictional) wants an assistant over 4,000 wiki documents with per-team visibility. Retrieval ships first: scoped search, cited answers, unknown-saying behavior, weekly evals. Logs show the assistant writes too verbosely for operators; a small style pass on prompts closes most of it. Fine-tuning waits until the eval trend flattens with a named gap. The evaluation post builds the scorecard; the capstone briefs carry the same access-control exercise for self-review.
Checklist: the decision path
- Retrieval with permissions ships before any training discussion.
- Evals run weekly with the same question set.
- The fine-tuning proposal names the logged gap and the success threshold.
- Cost, freshness and access stay in the brief, not in slides.
Related reading
Straight answers
Frequently asked questions
Should we fine-tune first?
Almost never. Retrieval over your documents answers most needs, stays fresh, and cites sources. Fine-tuning comes after retrieval hits its ceiling with evidence.
What breaks RAG at customers?
Access control, stale chunks and unmeasured quality. Fix permissions first, chunk with dates, and score answers before promising anything.
When is fine-tuning justified?
When logs show a stable style or vocabulary gap that retrieval plus prompts cannot close, with evals proving the gap and the fix.