A visual, step-by-step guide to the survey that gave the field a shared vocabulary — what hallucination is, its two types, where it enters the data → training → inference pipeline, and how to detect and mitigate it.
Fluent generation is old news — fluently wrong generation is the modern disease. Five years of scattered alarms led to one survey that finally drew the map.
Hallucination is not random noise. It is the by-product of training objectives that reward plausible text: a model trained to continue likely sequences will happily continue with likely-but-false content. Fixing it means intervening at specific pipeline stages — not wishing it away.
By 2022, every NLG subfield had met hallucination — and each was describing a different animal. The survey's first job was to prove they were all studying the same disease.
Each subfield touched one part of the animal: the summarization group felt the trunk, the dialogue group the ears, the MT group the legs — and each wrote a paper about a different creature. The survey walks all the way around the elephant and hands everyone the same map: the animal is one animal.
The survey's central gift to the field: a two-word taxonomy that works for any NLG task. The question is always the same — what does the output do to the source?
The output contradicts the source. The evidence to convict it is already in front of you.
Fell → rose, decline → increase. Checkable against the source, provably wrong.
The output adds content that cannot be verified from the source. The source can't settle whether it is true — it might even be true.
The acquisition rumor appears nowhere in the report — invented, or imported from outside. Unverifiable from the source.
The survey refuses the lazy answer ("models just make things up"). It traces hallucination to three distinct stages of the pipeline — and each stage has its own signature failure.
Autoregressive models condition each new token on everything already generated. Once a wrong detail is emitted, it becomes context — the model has no built-in eraser, so it elaborates the error coherently instead of correcting it. The result reads as detailed and confident, and every detail inherits the first mistake.
Before you fix hallucination, you have to find it. The survey organizes detection into four families — sorted by what they inspect, and how much access to the model they need.
| Family | Looks at | Catches best | Needs |
|---|---|---|---|
| Retrieval-based | External evidence vs. the output | Extrinsic hallucinations | Retriever + trusted corpus |
| Uncertainty-based | Variance across sampled answers | Unsupported, low-consensus claims | Sampling budget — black-box OK |
| Internal-state | Hidden states / logits | Errors before they are emitted | White-box model access |
| Post-hoc verification | Claim-level entailment checks | Both types, with editing | Verifier model or annotators |
No single family wins everywhere — the survey's message is that the right detector depends on access, budget, and which hallucination type you fear.
Because causes live at three stages, so do fixes. The survey's mitigation map runs from the corpus all the way to the decode step — and the strongest systems pull levers from more than one stage at once.
Fixes the stage-1 cause: stop teaching the model that inventing is normal.
Fixes the stage-2 cause: make training reward the right thing (InstructGPT Guide ↗).
Fixes the stage-3 cause: act at inference time — no retraining required.
A survey rarely changes what people build. This one did: its vocabulary and cause → fix map became the field's common infrastructure.
A survey maps; it does not experiment. No new model, no benchmark run, no guaranteed fix ships with this paper. Every mitigation it catalogues reduces hallucination in some setting — none eliminates it. Detecting extrinsic claims still needs an external source of truth, which does not always exist. And the field it unified is now scaling faster than the map can be redrawn.
Before this survey, every NLG sub-community had its own word for the same failure. Ji et al. merged them into one operational map: a definition anchored to the source, causes pinned to pipeline stages, and detection and mitigation slotted against each stage — the reference frame the whole LLM era still uses.
Check your understanding of the key concepts from the hallucination survey.
Everything you need to remember about this survey.