History Problem Core Idea Evaluation Results Impact Quiz Takeaways
Interactive Paper Explainer

Trained for the Noisy Corpus
RAFT

Domain RAG is an open-book exam with decoys everywhere. RAFT fine-tunes on questions where some retrieved documents are distractors — so the model learns to study the right pages and show its work.

Start Learning Read the Paper ↗
Train-time
Distractor docs
CoT
Training answers
Domain-RAG
Focus
2024
Tianjun Zhang et al.
History

The Open-Book Exam Nobody Trained For

Fine-tuning and RAG were treated as alternatives; RAFT merged them for domain adaptation.

2020-23
The two adaptation paths
RAG-prompting: fresh knowledge, no training, fragile to noise. Fine-tuning: domain fluency, but knowledge frozen into weights — time-critical facts go stale.
2023
RAG meets reality
Production domain corpora are messy: retrievers return distractors; models trust the wrong pages (and middles lose — entry #33).
Mar 2024
🚀 RAFT
Zhang et al.: fine-tune the model ON the RAG task itself — questions + oracle + distractor documents, chain-of-thought answers citing the right chunks.
2024+
Domain-RAG recipes
Medical/legal/finance assistants fine-tune with RAFT-style mixes; chunk-level citation training becomes the trust pattern.
Train on the Exam's Conditions

If deployment means 'answer from K retrieved documents, some irrelevant,' then training should look exactly like that. RAFT constructs finetuning examples where the context contains the oracle document plus distractors, and the target answer is a chain-of-thought that quotes the correct chunk first. The model learns three skills jointly: identify the supporting document, ignore the noise, and reason aloud from the evidence.

Chapter 01

Stale Weights, Noisy Retrieval

The two half-solutions that RAFT fused.

📚
The Adaptation Dilemma
  • Fine-tuning bakes knowledge in — but time-critical and private facts change faster than retraining cycles
  • RAG-prompting stays fresh — but the model was never trained to handle retrieved noise, distractors, or mixed relevance
  • Domain corpora punish naivety: retrieval returns the wrong pages alongside right ones
  • General-purpose models treat every retrieved document as equally trustworthy evidence
🎓
The RAFT Answer
  • Fine-tune on (question, oracle doc + distractor docs, CoT answer) triples — the deployment condition, reproduced in training
  • A fraction of examples use distractor-only contexts — training the model to trust its weights when retrieval fails
  • CoT answers open by quoting the correct chunk — citation discipline learned, not post-processed
  • Consistent gains over domain fine-tuning and general RAG prompting in open-book domain QA
Analogy — Studying with Decoys in the Binder

A student who only studies clean notes (fine-tuning) is lost the day the exam binder has 20 shuffled pages, 17 of them wrong. RAFT is the tutor who deliberately mixes decoys into the practice binder — and grades the answer only if the student cites the correct page and reasons from it. Exam day holds no surprises.

Chapter 02

The Training Recipe

Three design choices that define RAFT — each one mapping to a deployment skill.

1️⃣ Oracle + distractors
Each training context mixes the gold document with retrieved distractors from the domain corpus — the model must find the signal, not just read the stack.
2️⃣ Distractor-only fraction
A portion of examples contain NO oracle document — training graceful fallback: answer from parametric knowledge when retrieval comes up empty or poisoned.
3️⃣ CoT with chunk citation
Target answers are chain-of-thought that opens by quoting the correct chunk ("…according to document 3…") then reasons — attribution as a trained reflex.
Why This Beats the Alternatives

Domain fine-tuning alone answers from stale weights; general RAG prompting reads whatever arrives. RAFT trains the exact composite skill — selective reading under retrieval noise — and the paper shows it outperforms both in the open-book setting across domain benchmarks (with a smaller model often beating a larger general one). The chunk-citation CoT is not cosmetic: forcing the model to name its evidence source first measurably improves the reasoning that follows, the same discipline process supervision (entry #45) brings to math.

Interactive Demo — One Question, Four Training Philosophies

Tab through how each adaptation strategy answers the same domain question — and where each one breaks.

Chapter 03

The Evaluation Lens

What 'open-book domain QA' means and how RAFT measures progress.

Evaluation Setup

The broader lesson for practitioners: retrieval conditions are part of the training data contract — simulate deployment noise in fine-tuning or your RAG stack ships brittle.

Interactive Demo — Building a RAFT Training Example

How one (question, documents, answer) triple is constructed — the recipe you can apply to any domain corpus.

Chapter 05

The Composite Skill Wins

Fine-tuning FOR retrieval beats fine-tuning OR retrieval.

OPEN-BOOK QA
improves
consistently over RAG-prompting and domain SFT
DISTRACTOR ROBUSTNESS
trained
noise in context costs RAFT less at eval time
CITATIONS
learned
chunk-quoting CoT as the trained answer format
FALLBACK
graceful
distractor-only training handles retrieval misses
Interactive Demo — Distractor Count vs Accuracy

Press run for the robustness pattern: as distractor documents pile into the context, untrained RAG decays fastest — RAFT holds its line.

Legacy

Legacy — Deployment Conditions as Training Data

RAFT's principle now echoes through every serious domain-RAG build.

🧪 The simulation principle
'Train under the conditions you will serve under' became explicit doctrine — distractor-aware training, retrieval-noise augmentation, and failure-case curricula.
🏥 Domain assistant recipes
Medical, legal, and enterprise assistants fine-tune with RAFT-style oracle/distractor mixes — the pattern for private-domain RAG trust.
📎 Citation-first answers
Quoting the source chunk before reasoning migrated into production answer formats — ALCE's agenda (entry #32), trained-in rather than prompted.
⚖️ The SFT-vs-RAG truce
RAFT settled a fake war: fine-tuning and retrieval are composable — the model adapts to the corpus and the corpus keeps the model honest.
⚠️ What it did NOT solve
Needs domain QA data to construct examples; benefits hinge on realistic distractor simulation; multi-hop cross-document reasoning remains a frontier beyond its single-oracle setup.
🛤 Read next
The RAG curriculum: Self-RAG · RAGTruth · LoCoMo
Test Yourself

Quick Quiz

Check your understanding of the key concepts from RAFT.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ RAFT = fine-tuning ON the RAG task: questions with oracle + distractor documents, citation-first CoT answers.
✅ Distractor-only examples train graceful fallback when retrieval fails.
✅ Consistent open-book domain QA gains over domain SFT and general RAG prompting.
✅ Chunk-quoting answers make attribution a trained behavior — receipts by reflex.
✅ Robust to increasing distractor counts — selective reading survives noise.
✅ Read it as the reconciliation of fine-tuning and retrieval into one adaptation recipe.