History Problem Core Idea Consequences Results Impact Quiz Takeaways
Interactive Paper Explainer

Hallucination Is
Mathematically Required
Why LMs Hallucinate

The 2025 theory answer: hallucinations are errors in binary fact classification — and even a perfectly calibrated model that knows fact A is 10× more likely than fact B must still err at predictable, non-zero rates.

Start Learning Read the Paper ↗
19%+
Lower bound (Zipfian)
40%+
On long generations
2025
Kalai et al.
0
Magic fixes
History

From Symptom to Diagnosis

Four years of measuring hallucination; one paper explaining why it cannot fully be trained away.

2021-23
The measurement era
TruthfulQA, SelfCheckGPT, FActScore, HaluEval (entries #59-62) — detection and benchmarks flourish; the cause stays folkloric: 'the model is just guessing'.
2023-24
Folk theories abound
RLHF punishment of "I don't know", greedy decoding, sampling temperature — each captures a facet, none explains the phenomenon structurally.
Sep 2025
🚀 Why LMs Hallucinate
Kalai, Vempala & Zhang: hallucination as binary-classification error under calibration — with provable lower bounds: ~19% error rate for calibrated predictors on Zipfian-distributed facts, 40%+ for long generations.
2025+
Theory-aware mitigation
The field reframes: if error is structural, the honest defenses are verification and abstention — not more data alone (the V-category's whole design brief).
The Binary-Classification Reduction

Take any statement class where a model must decide TRUE or FALSE about a specific fact ('The Eiffel Tower is in Paris'). Under calibration — the model predicts fact frequencies correctly — a fact that is true but rare competes with false-but-plausible alternatives whose predicted rates sit just below it. The theorem's shape: if a model assigns rates p1 > p2 to two competing statements, it must make errors at a rate bounded below by a function of p1/(p1+p2) — no amount of scale removes the errors when the distribution of facts is heavy-tailed. On Zipfian distributions like natural-world facts, the bound lands around 19%; for long generations where any single error taints the output, the error rate climbs above 40%.

Chapter 01

Blame the Training? Blame the Math

The paper's reframing: hallucination is not a bug to patch but a statistical necessity under the field's own objectives.

🕳
The Symptom Mindset
  • Hallucination treated as a training artifact: more data, better RLHF, stricter decoding — cure assumed to exist
  • RLHF-style objectives demonstrably reward confident guessing over admitting uncertainty (the paper's training-pipeline analysis)
  • No explanation for why even well-calibrated, well-trained models keep hallucinating at stubborn rates
  • Mitigation work proceeds without knowing the floor — effort spent below an unknown bound
📐
The Theoretical Answer
  • Hallucination = errors in binary classification of statements as fact vs fiction
  • Calibration analysis: correctly-predicted fact frequencies STILL force errors at provable rates
  • Lower bounds: ~19% error for calibrated predictors on Zipfian-distributed knowledge; 40%+ for long-form generation
  • Implication: verification and abstention are load-bearing, not optional extras
Analogy — The Exam Where Guessing Is Mandatory

A student is graded only on answers, never on 'I don't know' — so rational students guess. Now make the questions follow a Zipf-like rarity curve: rare facts mostly false-sounding but sometimes true. Even a student with perfect knowledge of how often each answer is right (calibration!) still must commit on every question — and the grading curve mathematically guarantees a minimum wrong-answer rate. The paper's point: you built the exam, the curve, and the rules — the errors are your spec, not the student's character.

Chapter 02

The Statistical Argument

The paper's machinery, from calibration to the bound.

The setup
  • Binary classification: for a prompt, a statement is either the fact or a plausible impostor
  • Calibration: the model's predicted probability of each fact matches its true frequency
  • Training and evaluation reward confident completion over abstention — the pipeline's contribution (RLHF analyses included)
  • Hallucination = choosing the impostor; the question is the minimum rate of such errors
The bound (intuition)
  • If fact F is predicted at rate p and competing impostors occupy the remaining mass, a calibrated model MUST sometimes pick an impostor
  • Heavy-tailed (Zipfian) fact distributions make p's small and errors unavoidable at ~19%+
  • Long generations: per-statement errors compound — any error taints the whole output → 40%+ of generations contain one
  • Calibration is the BEST case: miscalibrated models do worse
What the Training Pipeline Adds

The theory separates two causes. Statistical: even ideal calibration leaves the bound intact. Procedural: the field's own training and evaluation reward guessing over acknowledging uncertainty — RLHF raters penalize hedging, benchmarks score confident outputs, and the "I don't know" response is trained out. The pipeline pushes models to the bound and beyond; the bound guarantees it never reaches zero. The two-layer diagnosis is what makes the paper actionable rather than fatalistic.

Interactive Demo — Where the Error Comes From

Tab through the components of the argument — calibration, the impostor, the forced choice, the bound.

Chapter 03

The Consequences

If the floor is real, what follows for engineering?

🔍 Verification is structural
If errors are provably non-zero, checking outputs (retrieval grounding, tests, verifiers — entries #32, #34, #45) is not belt-and-suspenders: it is the only path below the floor.
🤐 Abstention must be rewarded
If guessing is forced by objective design, then honest uncertainty needs positive reward — calibration-aware training and abstention-allowing evals follow directly.
📏 Benchmark recalibration
"Zero hallucination" claims are category errors; the right metric is distance-from-bound, per fact distribution and generation length.
⚖️ Deployment math
19% (Zipfian facts) and 40%+ (long generations) are numbers risk models can use: hallucination is a priced constant, not a surprise.
Interactive Demo — One Fact, One Impostor

Watch a calibrated model face a rare fact — and be forced into a mathematically guaranteed error rate.

Chapter 05

The Floor Under the Industry

Numbers that turn a vibe ('models make things up') into a constant.

calibrated rates p1 > p2  →  hallucination rate ≥ f( p1 / (p1 + p2) )  ·  Zipf ⇒ ~19%
p1, p2
Fact rates
The model's predicted frequencies for a fact and its competing impostor — calibration means these are correct.
binary choice
The forced decision
The model must output one statement; there is no abstain option in standard decoding objectives.
Zipf
The distribution
Natural-world facts follow power-law frequencies — most facts are rare, so p's are small and errors concentrate.
~19% / 40%+
The floors
Error-rate lower bounds: single facts on Zipfian distributions, and long-form generation where any error taints the output.
ZIPFIAN FACTS
≥ ~19%
error lower bound for calibrated models
LONG GENERATIONS
40%+
share of outputs containing an error
BLAME SOURCE 1
statistics
calibration cannot reach zero
BLAME SOURCE 2
training
objectives reward guessing over abstention
Interactive Demo — So Nothing Works?

The bound sounds fatal. Press reveal for what the theory actually prescribes.

Legacy

Legacy — The Constellation Recentered

The hallucination literature now orbits a theorem.

📐 From folklore to theorem
The VI-category's scattered mechanisms (detection, benchmarks, corpora — entries #59-64) gained a common mathematical frame: errors in binary classification under calibrated prediction.
🛡 Verification's promotion
If the floor is provable, RAG-grounding, verifiers, and abstention stop being 'mitigations' and become the only structurally sound defenses — Category III's whole thesis, now with proofs.
🎚 Objective redesign
The training-pipeline analysis (guessing rewarded over uncertainty) gave RLHF-era labs a concrete target: calibration-aware reward and abstention-friendly evaluation.
🧮 Deployment risk math
19%/40%+ turned hallucination from an embarrassment into a priced constant — numbers auditors and product teams can actually plan around.
⚠️ What it did NOT solve
The bounds are distribution-dependent idealizations; real models are also miscalibrated, poorly abstaining, and adversarially probed — the floor is the BEST case, so practice is worse; and verification itself inherits error rates (the verifier's floor).
🛤 Read next
The response stack: TruthfulQA · Self-RAG · FActScore
Test Yourself

Quick Quiz

Check your understanding of the key concepts from Why LMs Hallucinate.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Hallucination = binary-classification error: fact vs impostor, chosen under forced completion.
✅ Even perfectly calibrated models carry provable error floors — statistics, not character.
✅ Zipfian fact distributions: ~19% error floor; long generations: 40%+ error-containing outputs.
✅ Training/eval reward guessing over abstention — the procedural cause stacked on the statistical one.
✅ The sound defenses: verification, grounding, and honestly-rewarded abstention — not magic data scale.
✅ Read it as the theorem that reorganized the entire hallucination category around a floor.