History Problem Taxonomy Causes Detection Mitigation Impact Deep Dive Quiz
Interactive Paper Explainer

The Anatomy of Hallucination
A Field Guide for NLG + LLMs

A visual, step-by-step guide to the survey that gave the field a shared vocabulary — what hallucination is, its two types, where it enters the data → training → inference pipeline, and how to detect and mitigate it.

Start Learning Read the Paper ↗
2
Hallucination Types
3
Pipeline Stages
10
Authors
6+
NLG Tasks Mapped
History

How Hallucination Became Everyone's Problem

Fluent generation is old news — fluently wrong generation is the modern disease. Five years of scattered alarms led to one survey that finally drew the map.

2017
Fluent nonsense appears
Seq2seq and early Transformer models make MT and abstractive summarization fluent — and researchers notice they sometimes generate fluent text with no connection to the source.
2019
Dialogue joins in
Task-oriented dialogue studies show systems confidently inventing slots, entities, and API results that never existed in the knowledge base.
2020
LLMs + RAG
The few-shot LLM era (GPT-3 Guide ↗) scales hallucination along with fluency; RAG (RAG Guide ↗) proposes grounding generation in retrieved documents.
2021
TruthfulQA
Lin, Hilton & Evans show that imitating human text means imitating human falsehoods — models fail questions humans also fail (TruthfulQA Guide ↗).
2022 · Feb
🚀 The unifying survey
Ji et al. publish "Survey of Hallucination in Natural Language Generation" (arXiv 2202.03629) — one definition, one taxonomy, and a cause → detection → mitigation map for the whole field.
2023
Detection gets operational
The survey lands in ACM Computing Surveys; SelfCheckGPT and FActScore turn hallucination detection into measurable, deployable pipelines (FActScore Guide ↗).
Key Insight

Hallucination is not random noise. It is the by-product of training objectives that reward plausible text: a model trained to continue likely sequences will happily continue with likely-but-false content. Fixing it means intervening at specific pipeline stages — not wishing it away.

THE TWO FAILING MODES IN THE DEFINITION
nonsensical — output is broken on its own
unfaithful — output drifts from the source
Both count as hallucination. The unfaithful mode is the one that scales with LLMs — and the one this survey is famous for mapping.
Chapter 01

One Symptom, Six Research Silos

By 2022, every NLG subfield had met hallucination — and each was describing a different animal. The survey's first job was to prove they were all studying the same disease.

🌀
The Fragmented Field
  • Summaries invent facts; chatbots invent API results; MT invents content; VQA describes objects that are not there
  • Each task community uses its own definition and its own tests
  • A fix that works for translation is never tried on chatbots
  • "Hallucination" itself means different things in different papers
  • Insights stay stranded — the field keeps rediscovering the same failure modes
🗺️
The Survey's Answer
  • One definition: nonsensical or unfaithful to the provided source
  • One taxonomy — intrinsic vs extrinsic — that works for every task
  • Causes traced to three stages: data → training → inference
  • Detection and mitigation organized by what they actually do
  • Classic NLG tasks and the new LLMs in a single frame
Analogy — The Blind Researchers and the Elephant

Each subfield touched one part of the animal: the summarization group felt the trunk, the dialogue group the ears, the MT group the legs — and each wrote a paper about a different creature. The survey walks all the way around the elephant and hands everyone the same map: the animal is one animal.

one failure mode, many local names → fabrication unfaithfulness confabulation nonsense off-target output
Chapter 02

The Core Idea — Intrinsic vs Extrinsic

The survey's central gift to the field: a two-word taxonomy that works for any NLG task. The question is always the same — what does the output do to the source?

Hallucination = output that is nonsensical or unfaithful to the provided source
Nonsensical
broken on its own
The output contradicts itself or degenerates — you can flag it without any source at all.
Unfaithful
drifts from the source
The output does not stay true to the input it was supposed to be grounded in.
Intrinsic
contradicts the source
Provable from the source alone: "revenue fell" in, "revenue rose" out.
Extrinsic
unverifiable from source
Brings in outside information the source neither confirms nor denies — truth unknown.
Intrinsic Hallucination

The output contradicts the source. The evidence to convict it is already in front of you.

SOURCE — EARNINGS REPORT
"Q3 revenue fell 12%, the third straight quarterly decline."

GENERATED SUMMARY
"Q3 revenue rose 12%, the third straight quarterly increase."

Fell → rose, decline → increase. Checkable against the source, provably wrong.

Extrinsic Hallucination

The output adds content that cannot be verified from the source. The source can't settle whether it is true — it might even be true.

SOURCE — EARNINGS REPORT
"Q3 revenue fell 12%, the third straight quarterly decline."

GENERATED SUMMARY
"The decline is expected to reverse once the rumored acquisition of Vertex Labs closes."

The acquisition rumor appears nowhere in the report — invented, or imported from outside. Unverifiable from the source.

Interactive Demo — Intrinsic or Extrinsic?

A source, a generated claim — you are the classifier. Four rounds across four different NLG settings: summarization, dialogue, machine translation, and retrieval-augmented generation.

Round 1 / 4 Summarization Score 0/0
Chapter 03

Where It Comes From — Three Stages to Poison

The survey refuses the lazy answer ("models just make things up"). It traces hallucination to three distinct stages of the pipeline — and each stage has its own signature failure.

STAGE 1 · DATA
The corpus teaches inventing
  • ▸ Source–counterfactual misalignment: training pairs whose target text is not actually supported by the input. The model literally learns that unsupported details are fine.
  • ▸ Social bias + duplicates: duplicated, biased passages get memorized — and later regurgitated with full confidence.
STAGE 2 · TRAINING
The objective rewards fluency, not truth
  • ▸ Exposure bias: teacher forcing feeds only gold prefixes, so the model never sees its own errors — and never learns to recover from them.
  • ▸ Representation gaps: imperfectly encoded knowledge is still emitted with the same confidence as solid knowledge.
  • ▸ Hallucination snowballing: an early wrong detail becomes context — and every later sentence builds on it (demo below).
STAGE 3 · INFERENCE
The decoding roll can leave the rails
  • ▸ Random sampling: the higher the decoding randomness (temperature), the more often low-probability — and off-source — tokens get picked.
  • ▸ Confident decoding: nothing in the sampler knows which continuations are supported; unlikely does not mean flagged.
The Mechanism — Hallucination Snowballing

Autoregressive models condition each new token on everything already generated. Once a wrong detail is emitted, it becomes context — the model has no built-in eraser, so it elaborates the error coherently instead of correcting it. The result reads as detailed and confident, and every detail inherits the first mistake.

Interactive Demo — Hallucination Snowball

Watch a fluent answer build itself on top of one invented number. Sentence 1 contains a single wrong detail (red); every later sentence (amber) conditions on it. Press start and watch the error compound.

generation log · one question, one answer
Chapter 04

Catching It — Four Families of Detection

Before you fix hallucination, you have to find it. The survey organizes detection into four families — sorted by what they inspect, and how much access to the model they need.

🔎
Retrieval-augmented checking
Retrieve external evidence at check time, then test each generated claim against it. Strong on extrinsic hallucinations — as long as the retriever and the corpus are trustworthy.
🎲
Uncertainty-based
Sample the model repeatedly: if answers disagree (self-consistency) or semantic entropy runs high, the claim is probably unsupported. Works even on black-box APIs — at the cost of a sampling budget.
🧠
Internal-state probing
Look inside: hidden states and logits often look different just before a hallucinated token is emitted. Needs white-box access — but it can catch hallucinations before they leave the model.
🛠️
Post-hoc verification
Verify-then-edit: a second pass (another model, or an NLI-style entailment checker) scores claims and rewrites the flagged spans. Powerful, but the verifier has its own error rate.
Detection Families — What Each One Needs
FamilyLooks atCatches bestNeeds
Retrieval-basedExternal evidence vs. the outputExtrinsic hallucinationsRetriever + trusted corpus
Uncertainty-basedVariance across sampled answersUnsupported, low-consensus claimsSampling budget — black-box OK
Internal-stateHidden states / logitsErrors before they are emittedWhite-box model access
Post-hoc verificationClaim-level entailment checksBoth types, with editingVerifier model or annotators

No single family wins everywhere — the survey's message is that the right detector depends on access, budget, and which hallucination type you fear.

Chapter 05

Fixing It — Every Stage Gets a Lever

Because causes live at three stages, so do fixes. The survey's mitigation map runs from the corpus all the way to the decode step — and the strongest systems pull levers from more than one stage at once.

Data-Level Mitigation
  • ▸ Cleaning: remove misaligned pairs, duplicates, and biased content from the corpus
  • ▸ Grounding: build training pairs whose targets are provable from their inputs

Fixes the stage-1 cause: stop teaching the model that inventing is normal.

Training-Level Mitigation
  • ▸ RLHF: human preferences reward faithful outputs and penalize unsupported claims
  • ▸ Grounded objectives: optimize faithfulness to the source, not just fluency

Fixes the stage-2 cause: make training reward the right thing (InstructGPT Guide ↗).

Generation & Post-Hoc Mitigation
  • ▸ RAG: anchor decoding in freshly retrieved evidence
  • ▸ Contrastive decoding: steer away from the hallucination-prone distribution
  • ▸ Self-refinement / verification: draft → check claims → edit or regenerate

Fixes the stage-3 cause: act at inference time — no retraining required.

Interactive Demo — Fix the Pipeline

The survey's core diagnosis: hallucination has different causes at different stages, so the right fix depends on where you intervene. Click a mitigation below to see which stage it targets and why.

STAGE 1 · DATA
Poisoned corpus
Misaligned source–counterfactual pairs; biased and duplicated passages.
STAGE 2 · TRAINING
Fluency-only objective
Exposure bias, representation gaps, and snowballing errors.
STAGE 3 · INFERENCE
Ungrounded decoding
Random sampling picks unsupported, low-probability paths.
0 / 5 fixes placed
Legacy

Impact — The Shared Map

A survey rarely changes what people build. This one did: its vocabulary and cause → fix map became the field's common infrastructure.

📖 Standard vocabulary
"Intrinsic / extrinsic" became the default way papers, benchmarks, and error analyses describe hallucination — the survey's taxonomy is now common language.
📊 Benchmark wave
The detection framing inspired a wave of measurable evaluation — HaluEval, FActScore, and long-form factuality benchmarks.
🔗 RAG as the default
Grounding generation in retrieval moved from a 2020 research idea to the standard production mitigation — exactly the inference-stage lever the survey mapped.
🔄 Self-correction research
Self-refinement and verification loops grew into a whole research line — from self-consistency to SelfCheckGPT-style checkers.
🎯 RLHF as grounding
Human-preference training became the default hallucination lever in LLM pipelines — rewarding truthfulness, not just plausibility.
⚠️ Still open
The survey's open questions are still open: better benchmarks, interpretability, and evaluation of long-form factuality. Hallucination is reduced — not solved.
What It Did NOT Solve

A survey maps; it does not experiment. No new model, no benchmark run, no guaranteed fix ships with this paper. Every mitigation it catalogues reduces hallucination in some setting — none eliminates it. Detecting extrinsic claims still needs an external source of truth, which does not always exist. And the field it unified is now scaling faster than the map can be redrawn.

Deep Dive

The Field Map: One Vocabulary for a Field-Wide Problem

Before this survey, every NLG sub-community had its own word for the same failure. Ji et al. merged them into one operational map: a definition anchored to the source, causes pinned to pipeline stages, and detection and mitigation slotted against each stage — the reference frame the whole LLM era still uses.

📐
The Definition — Faithful to What?
  • Intrinsic hallucination: the output contradicts the provided source
  • Extrinsic hallucination: output can't be verified from the source — it imports outside claims
  • Crucially source-relative: an extrinsic claim may be true in the world — the system still can't justify it
  • This is why RAG systems can hallucinate with a correct document in hand
🧩
Why One Fix Can't Exist
  • Causes live at three different stages: data (misaligned/biased pairs), training (exposure bias, snowballing), inference (sampling)
  • Detection families assume different access: retrieval needs a corpus, probing needs weights, uncertainty needs samples
  • Summarization, dialogue, MT, QA, data-to-text each break grounding differently
  • So mitigation is a stage-by-stage program, not a switch — the survey's real thesis
Interactive Demo — Walk the Pipeline

Select a pipeline stage to see what breaks there, which detection families can see it, and which mitigation levers exist. This is the survey's mental model, live.

VERDICT
A taxonomy is infrastructure
The survey's lasting value isn't any single finding — it's that later work could say "extrinsic hallucination in the RAG setting" and be understood instantly. PaperMap's hallucination track is organized exactly along this map: detection without access (SelfCheckGPT), post-hoc verification (FActScore), and grounded generation (RAGTruth).
❄️ Snowballing
One early hallucination raises the probability of the next — errors self-reinforce within a single generation. It's why long-form factual precision decays with length, a fact FActScore later quantified.
🔁 Exposure bias
Models train on gold prefixes but consume their own outputs at test time — distribution drift compounds over tokens, and drift is where fabrication lives.
🔍 Four detection families
Retrieval-based, uncertainty-based, internal-state probing, post-hoc verification — each buys a different trust guarantee at a different access level to the model.
🏗️ Mitigation ladder
Data cleaning + grounded construction; RLHF + grounded objectives; RAG, contrastive decoding, self-verification at inference. Stack them — no single rung reaches trustworthy.
Test Yourself

Quick Quiz

Check your understanding of the key concepts from the hallucination survey.

Reference

Key Takeaways

Everything you need to remember about this survey.

✅ Definition: hallucination is generated content that is nonsensical or unfaithful to the provided source.
✅ Taxonomy: intrinsic hallucinations contradict the source; extrinsic ones cannot be verified from it — they bring in outside information.
✅ Causes map to three stages: data (misaligned + biased pairs), training (exposure bias, representation gaps, snowballing), inference (random sampling).
✅ Detection comes in four families: retrieval-based, uncertainty-based, internal-state probing, and post-hoc verification.
✅ Mitigation has a lever at every stage: data cleaning + grounding, RLHF + grounded objectives, RAG + contrastive decoding + self-verification.
✅ The survey spans the NLG landscape — summarization, dialogue, MT, generative QA, data-to-text, visual-language — and the LLM era, with one vocabulary.