History Problem Core Idea Results Results Impact Quiz Takeaways
Interactive Paper Explainer

Retrieve When You Need It
Self-RAG

Standard RAG retrieves on every query whether or not it helps. Self-RAG trains the model to decide — and to grade its own answers with reflection tokens for support and usefulness.

Start Learning Read the Paper ↗
4
Reflection token types
On-demand
Retrieval
ChatGPT-level
Factuality gains
2023
Asai et al.
History

From Always-On to On-Demand Retrieval

RAG fixed hallucination at the cost of versatility — Self-RAG made retrieval a decision.

2020-22
RAG's trade-off
Retrieve-and-generate (entry #29) grounds answers but retrieves indiscriminately — pop-culture queries get Wikipedia, code questions get web noise.
2022
The versatility cost
Forcing fixed retrieved passages into every prompt degrades chat quality and can inject unhelpful context — noted across the RAG literature.
Oct 2023
🚀 Self-RAG
Asai et al.: the model itself emits reflection tokens — retrieve-or-not, is-relevant, is-supported, is-useful — trained via a critic teacher and interleaved with generation.
2023+
Self-critique lineage
Constitutional AI-style self-evaluation, verifier tokens, and process supervision (entry #45) share the DNA: models grading their own outputs.
2024+
Adaptive RAG standard
Production systems route easy queries to parametric answers and hard ones to retrieval — Self-RAG's decision learned rather than hand-coded.
Four Tokens, Four Judgments

Self-RAG interleaves special reflection tokens with ordinary output: [Retrieve?] (is retrieval needed?), [IsRel] (is the passage relevant?), [IsSup] (is the claim supported by it?), and [IsUse] (is the response useful?). Training uses a critic model (GPT-4-annotated data distilled into the LLM) so the generator learns to make these calls itself — turning a fixed pipeline into a learned policy.

Chapter 01

RAG's Indiscriminate Habit

The costs nobody budgeted: retrieval for everything, relevance un-checked, quality un-graded.

🔁
Always-Retrieve RAG
  • Retrieval fires on every query regardless of need — 'hello' fetches web pages
  • A fixed number of passages injected whether relevant or not — context pollution
  • No mechanism checks whether generations are actually supported by retrieved evidence
  • Diminished versatility: chatty and creative queries suffer under forced grounding
🪞
The Self-RAG Answer
  • [Retrieve?] token: the model decides, per segment, whether to fetch
  • [IsRel] / [IsSup] / [IsUse]: self-graded relevance, support, and utility
  • Critic-distilled training: GPT-4-labeled reflection data teaches the generator
  • Results: outperforms ChatGPT and Llama2-chat on open-domain QA, factuality, and citation generation
Analogy — The Researcher with Judgment

Naive RAG is an intern who Googles everything — including their own name — and pastes whatever loads first. Self-RAG is a researcher who knows the difference between trivia they own and questions needing the library, pulls sources when warranted, checks quotes before using them, and asks 'did I actually answer the question?' before hitting send.

Chapter 02

The Token Grammar

The reflection vocabulary interleaved with generation — the whole control system in one card.

🔍 [Retrieve?]
Per generation segment: does this next claim need evidence, or does parametric knowledge suffice? Adaptive branching.
🧲 [IsRel]
Is the retrieved passage relevant to the prompt at hand? Filters context pollution before it enters generation.
✅ [IsSup]
Is the generated segment supported by the retrieved passage? — the grounding grade, per claim.
⭐ [IsUse]
Is the response useful as an answer to the question? — the quality grade, beyond mere support.
Training: critic distillation
  • A strong critic (GPT-4) labels reflection tokens on sampled outputs and retrieved passages
  • The generator is finetuned on the combined (text + reflection) sequences
  • At inference the generator emits its own reflection tokens — no critic in the loop
Inference: branching by token
  • [Retrieve?] = no → pure parametric continuation
  • [Retrieve?] = yes → fetch passages, then [IsRel] filters, [IsSup] grades each continuation
  • Best continuation picked by weighted reflection scores (support/utility)
  • Citations fall out naturally: supported segments carry their passages
Interactive Demo — Same Query, Two Policies

Tab through query types under always-on RAG vs Self-RAG — watch the retrieval decision adapt.

Chapter 03

The Evidence

Where Self-RAG lands: the 2023 leaderboard beats that made it famous.

Reported Results

Self-RAG reframed the RAG stack from a fixed pipeline into a policy — and reflection tokens became a reusable pattern (process supervision, verifier heads, self-consistency critics).

Interactive Demo — One Answer, Fully Reflected

Trace a generation with every reflection token firing — the full self-policing loop in five steps.

Chapter 05

Grounded and Versatile

The dual win the always-retrieve camp said was impossible.

OPEN-DOMAIN QA
beats ChatGPT
and Llama2-chat with the same task framing
FACTUALITY
improves
short-form and long-form grounding metrics
CITATIONS
emerge free
supported segments carry their evidence
ADAPTIVITY
learned
retrieval fired only when reflection says so
Interactive Demo — Remove One Token, Watch It Break

Each reflection token guards one failure mode. Press reveal to see the ablation pattern.

Legacy

Legacy — The Learned RAG Policy

Reflection tokens became a reusable pattern far beyond retrieval.

🎛 Adaptive retrieval standard
Production RAG stacks now route by query type — the decision Self-RAG moved from hand-coded heuristics into the model itself.
🪞 Self-critique as infrastructure
The critic-distillation recipe (strong model labels judgment, generator internalizes it) reappears in verifiers, process reward models (entry #45), and self-correction loops.
📎 Attribution for free
Per-claim support grades make citations an output of the generation process itself — ALCE's (entry #32) agenda, implemented inside the decoder.
⚖️ The versatility-accuracy framing
Self-RAG made explicit that RAG systems have TWO jobs (be right AND stay usable) — the trade-off later benchmarks evaluate directly.
⚠️ What it did NOT solve
Reflection quality inherits the critic's biases (GPT-4-labeled); token overhead complicates serving; and reflection can be confidently wrong — self-grading is a signal, not a guarantee.
🛤 Read next
The family: RAG · RAFT · RAGTruth
Test Yourself

Quick Quiz

Check your understanding of the key concepts from Self-RAG.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Self-RAG = reflection tokens: [Retrieve?] [IsRel] [IsSup] [IsUse] — a learned RAG policy.
✅ Critic distillation teaches the generator to self-reflect; inference runs without the critic.
✅ Outperforms ChatGPT and Llama2-chat on open-domain QA, factuality, and citations.
✅ Retrieval becomes per-segment and adaptive — no more context pollution on easy queries.
✅ Per-claim support grades make grounding inspectable and citations intrinsic.
✅ Read it as the bridge from RAG pipelines to self-policing generation.