The first benchmark to treat citations as a first-class output: LLMs must answer AND attribute each claim to retrieved evidence — with automatic metrics for citation recall and precision.
The attribution problem: generated text needs receipts, and nobody had a benchmark for receipt quality.
ALCE reformulates attribution as two checkable quantities: citation recall — what fraction of the generated statements are actually supported by their cited passages — and citation precision — what fraction of cited passages are actually necessary to support the statements. Between them, they separate the failure modes that 'it looks cited' hides: decorative citations, evidence-free claims, and quote-mining.
The trust problem ALCE formalized: outputs that cite neither their evidence nor their limits.
An uncited LLM answer is oral argument — persuasive, unaccountable. An ALCE answer is a filed brief: every assertion footnoted to evidence a judge can pull. And the metrics are the clerk's audit — did every claim cite something that actually supports it (recall), and is every citation load-bearing rather than decorative (precision)?
Questions, corpora, and the two metrics — the whole apparatus in one card.
The benchmark's first findings — and the taxonomy of citation failure it exposed.
The subtle contribution: ALCE made citation quality a leaderboard-able property, which moved it from a research aspiration to a trainable objective.
Two axes, both measurable — the quadrant ALCE forced the field to see.
Citation generation went from plea to property: measurable, trainable, and eventually contractual.
Check your understanding of the key concepts from ALCE.
Everything you need to remember about this paper.