RAG or fine-tuning for customer data: an FDE decision

Updated

Rows of archived books in a library

Start with retrieval, earn the right to train: the decision path for assistants over customer documents.

Customers ask for training because it sounds thorough. FDEs recommend retrieval first because it ships, cites and updates.

The short answer

Build retrieval with access control and evals first. Fine-tune only when measured logs prove a gap that retrieval cannot close.

The decision table

QuestionRetrieval wins whenFine-tuning earns when
FreshnessDocuments change weeklyStyle is stable for months
CitationsCustomer demands sourcesBehavior matters more than sources
Access controlPer-team visibility existsSingle audience, stable corpus
EvidenceNo evals yetEvals show a persistent gap
BudgetFixed monthly spendMeasured lift covers training cost

Worked example: the wiki assistant

A fictional company (fictional) wants an assistant over 4,000 wiki documents with per-team visibility. Retrieval ships first: scoped search, cited answers, unknown-saying behavior, weekly evals. Logs show the assistant writes too verbosely for operators; a small style pass on prompts closes most of it. Fine-tuning waits until the eval trend flattens with a named gap. The evaluation post builds the scorecard; the capstone briefs carry the same access-control exercise for self-review.

Checklist: the decision path

  1. Retrieval with permissions ships before any training discussion.
  2. Evals run weekly with the same question set.
  3. The fine-tuning proposal names the logged gap and the success threshold.
  4. Cost, freshness and access stay in the brief, not in slides.

Straight answers

Frequently asked questions

Should we fine-tune first?

Almost never. Retrieval over your documents answers most needs, stays fresh, and cites sources. Fine-tuning comes after retrieval hits its ceiling with evidence.

What breaks RAG at customers?

Access control, stale chunks and unmeasured quality. Fix permissions first, chunk with dates, and score answers before promising anything.

When is fine-tuning justified?

When logs show a stable style or vocabulary gap that retrieval plus prompts cannot close, with evals proving the gap and the fix.

Bu sayfanın Türkçesi