Before RAG had a name, REALM put a differentiable retriever inside masked-LM pre-training itself — the model learns which Wikipedia documents to fetch, not just how to fill masks.
The 2019-20 question: should facts live inside the network or in a corpus it can consult?
The difference from every earlier retrieve-then-read system: REALM's retriever is inside the pre-training objective. Masked-LM likelihood is computed as a marginal over retrieved documents — so gradient flows tell the model which retrievals help predict masked tokens. The corpus effectively becomes an external, inspectable extension of the parameters, consulted during pre-training, fine-tuning, and inference alike.
The parameter-memory problem REALM was built to escape.
Parameter-only models sit a closed-book exam: everything must be memorized before the test, and the syllabus is baked in. REALM walks in with the textbook allowed — and crucially, practiced studying WITH the book open during training, so it learned exactly which pages answer which questions.
The equation that makes retrieval trainable — and its clever computational shortcut.
How a query and a document meet inside the encoder — the template every later RAG system copied.
REALM concatenates the masked sentence with each retrieved document and lets BERT attend over both — the model must copy or infer the answer token from the document to score well. The retrieval corpus is Wikipedia chunked into ~100-word documents; the same retriever serves pre-training, fine-tuning (Natural Questions, WebQuestions, TriviaQA), and inference unchanged. The headline result: better open-domain QA accuracy than T5-XXL with 11× more parameters — the first clean demonstration that retrieved documents can substitute for parameter count on knowledge tasks.
The headline: retrieval buys back what parameters would otherwise have to memorize.
REALM's structural ideas outlived its specific architecture.
Check your understanding of the key concepts from REALM.
Everything you need to remember about this paper.