History Problem Core Idea Findings Results Impact Quiz Takeaways
Interactive Paper Explainer

Show Your Sources
ALCE

The first benchmark to treat citations as a first-class output: LLMs must answer AND attribute each claim to retrieved evidence — with automatic metrics for citation recall and precision.

Start Learning Read the Paper ↗
3
Dataset families
3
Metric dimensions
50%
ELI5 claims unsupported
2023
Gao et al.
History

From Confident to Accountable

The attribution problem: generated text needs receipts, and nobody had a benchmark for receipt quality.

2020-22
Hallucination goes mainstream
GPT-3-class models produce fluent, unverifiable text — 'trust me' as an output format (entries #59-63 measure the disease).
2022
Attribution as a research plea
Papers call for citations and grounded generation, but evaluation is manual, small-scale, and non-comparable.
May 2023
🚀 ALCE
Gao et al.: a reproducible benchmark — questions + retrieval corpora + automatic metrics (citation recall/precision) — making citation generation measurable.
2023+
The attribution stack
Inline citations, RAG with quotes, attribution evals in production guardrails — and the 'verifiability' axis of trust (entry #32 → the VI category).
2025
Attribution at deployment
Search-grounded assistants cite sources per claim — ALCE's framing, industrialized.
Citations Are a Measurable Skill

ALCE reformulates attribution as two checkable quantities: citation recall — what fraction of the generated statements are actually supported by their cited passages — and citation precision — what fraction of cited passages are actually necessary to support the statements. Between them, they separate the failure modes that 'it looks cited' hides: decorative citations, evidence-free claims, and quote-mining.

Chapter 01

Fluent, Receiptless

The trust problem ALCE formalized: outputs that cite neither their evidence nor their limits.

🗣
The Attribution Gap
  • LLM answers arrive with zero receipts — readers cannot verify claims without redoing the search
  • Early citation attempts were evaluated by hand: small studies, irreproducible, incomparable across models
  • Pipelines built on commercial search engines — the retrieval layer itself was unauditable
  • Hallucination metrics measure wrongness, not the verifiability that readers actually need
🧾
The ALCE Answer
  • Generation with inline citations: each claim span linked to retrieved document spans
  • Automatic metrics: citation recall (support) and citation precision (necessity)
  • Reproducible setup: public retrieval corpora, standard LLM APIs — no proprietary search dependency
  • Three task families spanning short-form QA, long-form synthesis, and conversational search
Analogy — The Courtroom Brief

An uncited LLM answer is oral argument — persuasive, unaccountable. An ALCE answer is a filed brief: every assertion footnoted to evidence a judge can pull. And the metrics are the clerk's audit — did every claim cite something that actually supports it (recall), and is every citation load-bearing rather than decorative (precision)?

Chapter 02

The Benchmark Anatomy

Questions, corpora, and the two metrics — the whole apparatus in one card.

Datasets and corpora
  • ASQA — ambiguous questions needing long, multi-facet answers (Wikipedia evidence)
  • ELI5 — long-form explain-like-I-am-5 answers (web retrieval corpus)
  • QAMPARI — multi-answer questions where many supporting passages must be cited together
  • Retrieval corpora shipped with the benchmark — the search layer is fixed and public, not a private engine
The two automatic metrics
  • Citation recall: fraction of generated statements supported by their cited passages — measures under-attribution
  • Citation precision: fraction of citations that support the statements they are attached to — measures citation padding
  • Support judgment: NLI-style entailment between the cited passage and the statement
  • Metric dimensions: fluency, correctness, citation quality — all strongly correlated with human judgements (per the paper)
  • Headline stat: on ELI5, even the best models lack complete citation support 50% of the time
Interactive Demo — The Citation Quality Spectrum

Tab through the same claim with different citation behaviors — from decorative to fully supported.

Chapter 03

What the Models Revealed

The benchmark's first findings — and the taxonomy of citation failure it exposed.

Findings (from the paper)

The subtle contribution: ALCE made citation quality a leaderboard-able property, which moved it from a research aspiration to a trainable objective.

Interactive Demo — Scoring One Generated Answer

Walk the automatic evaluation: split the answer into statements, check each against its citations, compute recall and precision.

Chapter 05

Correct AND Cited

Two axes, both measurable — the quadrant ALCE forced the field to see.

METRIC 1
recall
are cited passages actually supporting the claims?
METRIC 2
precision
are citations load-bearing, not decorative?
CORPORA
public
reproducible retrieval — no private search dependency
SKILL
learnable
in-context examples improve attribution substantially
Interactive Demo — Why Not Just Measure Hallucination?

Hallucination metrics exist (entry #60-62). Press reveal for why attribution needs its own benchmark.

Legacy

Legacy — The Receipt Standard

Citation generation went from plea to property: measurable, trainable, and eventually contractual.

📊 The evaluation genre
ALCE's recall/precision framing became the standard vocabulary of attribution evals — later search-grounded assistants report descendants of exactly these two numbers.
🔍 Grounded-generation training
In-context citation demonstrations mattering so much kicked off attribution-aware finetuning — receipts as a learned output format (the RAG with quotes pattern).
🤝 The reader-writer contract
ALCE formalized what production RAG owes users: per-claim evidence. Verification cost moved from reader to writer — the quiet governance win.
🧪 Reproducible retrieval
Shipping fixed corpora instead of relying on commercial search made attribution research comparable across labs for the first time.
⚠️ What it did NOT solve
Entailment checking is imperfect (NLI models err on long/technical spans); citation quality ≠ overall truthfulness (right claims can be badly cited); and multimodal attribution stayed out of scope.
🛤 Read next
The trust stack: RAGTruth · FActScore · Self-RAG
Test Yourself

Quick Quiz

Check your understanding of the key concepts from ALCE.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ ALCE = first automatic benchmark for LLM citation evaluation: questions + public corpora + metrics.
✅ Citation recall measures support; citation precision measures load-bearing-ness — two distinct failure modes.
✅ Support is entailment-based, so citation theater and quote-mining are actually caught.
✅ Even strong 2023 LLMs under-attribute: fluent answers, weak receipts.
✅ In-context examples of cited output improve attribution substantially — a learnable skill.
✅ Read it as the benchmark that turned 'show your sources' into a measurable, trainable property.