A visual, step-by-step guide to FActScore — the EMNLP 2023 metric that breaks a
long-form generation into atomic facts, checks each one against a knowledge source
like Wikipedia, and turns "is this model factual?" into a number you can audit.
FActScore is the meeting point of two old ideas — fine-grained content units from summarization, and claim verification from fact-checking — applied to a new problem: long-form generations from LLMs.
2004
Summarization Content Units (Nenkova & Passonneau)
Evaluating summaries by splitting them into small content units — the intellectual ancestor of the atomic fact.
2018
FEVER (Thorne et al.)
Fact verification at scale: check short human-written claims against Wikipedia evidence.
2020
QA/NLI metrics for summaries
FactCC, QAGS, FEQA: automatic consistency metrics that compare a summary to its source via questions or entailment.
2021–22
Attribution & entity checks
Entity-level precision and citation checking (Rashkin et al. 2021; Lee et al. 2022; Gao et al. 2022) — but only for short texts or citations.
2023 · May
🚀 FActScore (Min et al.)
Atomic facts + a chosen knowledge source. Fine-grained factual precision for long-form generation — published at EMNLP 2023.
2023 →
The factuality-eval wave
Long-form factuality becomes a standard eval: follow-ups like SAFE and FELM automate atomic-fact checking further; factuality sections spread through model cards and leaderboards.
Key Insight
One number per generation hides a distribution of errors. A typical 150-word bio contains 26–41 atomic facts. Two bios can both "look good" while one supports 66.7% of its facts and the other only 10.0%. FActScore makes that gap visible — and every individual failure findable.
ONE PROMPT, TWO SCORES (FIGURE 1 OF THE PAPER)
"Tell me a bio of Bridget Moynahan." ChatGPT: 9 facts → 6 ✓3 ✗ → 66.7% StableLM: 10 facts → 1 ✓9 ✗ → 10.0%
A human preference vote would just say "ChatGPT's bio is better."
Chapter 01
The Problem with Coarse Verdicts
Long-form generations mix supported facts, fabrications, and filler in one fluent paragraph. Judging the whole paragraph with one verdict — by human or by model — throws away exactly the information you need.
⚖️
The Old Way: One Verdict
Human A/B preference: "which bio is better?" — slow, subjective, ~$4 per generation
One score for ~150 words: no idea where the errors are
Fluent, confident prose masks fabrications — and humans are biased toward fluency
Binary judgments can't rank models that are close, or diagnose the gap
Strict sentence-level rules collapse everything: one error can sink a whole "unsupported" rating
🔬
FActScore: Grade Every Fact
Decompose the generation into short, single-claim atomic facts
Verify each atomic fact against a knowledge source you trust (English Wikipedia here)
Score = the percentage of atomic facts that are supported
An automated estimator reproduces human labels within a <2% error rate
Every error is visible — an itemized to-do list for fixing the model
Interactive Demo — Coarse vs Atomic
A real example from Figure 1 of the paper: the same prompt, "Tell me a bio of Bridget Moynahan.", answered by ChatGPT and by StableLM. Toggle between how a coarse eval sees the answers and what atomic decomposition reveals.
Why Fine-Grained Wins
Imagine grading a 30-question exam by skimming it once and writing "B+". That is preference-based eval. Now imagine marking every answer, counting the correct ones, and handing back the list of misses. That is FActScore — and because the grader can be an LM with retrieval, it costs a fraction of human annotation while staying within ~2 points of human scores.
Chapter 02
The Core Idea — Atomic Facts + a Trusted Source
Two moves define the metric: make the unit of evaluation small enough to verify (the atomic fact), and make truth relative to a knowledge source you choose — not to a global oracle.
FActScore = (1 / |A|) · Σ 1[aᵢ is supported by C]
A
Atomic facts
The generation is split into short sentences, each conveying one piece of information.
1[·]
Verifier
1 if the atomic fact is supported by the knowledge source C, else 0.
|A|
All facts count
Every atomic fact — supported or not — sits in the denominator. No free rides.
E[·]
Averaged
The score is the expectation over prompts, counting only generations where the model responds.
Key Idea 1 — The Atomic Fact
An atomic fact is a short sentence conveying one piece of information. It is more fundamental than a sentence: one sentence can bundle several verifiable claims, and partial support is common. Decomposition makes each claim separately checkable — the direct descendant of summarization content units (2004).
ONE SENTENCE → FIVE ATOMIC FACTS
"Marie Curie was a Polish-French physicist and chemist who won the Nobel Prize in Physics in 1903." → Marie Curie was Polish. → Marie Curie was French. → She was a physicist. → She was a chemist. → She won the 1903 Nobel Prize in Physics.
Key Idea 2 — Truth Relative to a Source
Instead of asking whether an atomic fact is globally true, FActScore asks whether it is supported by a given knowledge source C — the one end users consider reliable. The paper uses English Wikipedia for people biographies: objective, specific, reasonably self-consistent, with good coverage. Change the source, and the same generation can score differently — that is a feature, not a bug: factuality is audited against a citable reference.
✓ Supported
The knowledge source contains evidence for the atomic fact — possibly after a bit of inference.
✗ Not-supported
No evidence in the source — the fabrication the metric exists to catch.
✱ Irrelevant
Not related to the prompt — e.g., the model wrote about a different person entirely.
The Third Label in Action (Table 7 of the paper)
Asked for a bio of John Estes, one model produced text about William Estes, an actor on the CBS drama Blue Bloods. Every atomic fact in that answer — "William Estes is an American", "William Estes is an actor" — was labeled Irrelevant: individually checkable, but about the wrong person, so they should be removed rather than counted as factual wins or losses.
Chapter 03
The Pipeline — Decompose, Retrieve, Verify
From "Tell me a bio of X" to a single, auditable percentage. Five steps, each simple enough to inspect on its own.
Pipeline Overview
✍️
1 · Generate
Prompt "Tell me a bio of <entity>" — the model writes ~110–155 words, taken as-is.
✂️
2 · Decompose
Split into sentences; InstructGPT breaks each into atomic facts (humans revise, for gold labels).
📚
3 · Retrieve
For each fact, a GTR retriever (T5-based) pulls the most relevant Wikipedia passages.
🧑⚖️
4 · Verify
Passages + fact + "True or False?" go to an LM evaluator; it compares the probabilities of True/False.
🧮
5 · Score
% of atomic facts supported = FActScore, averaged over prompts where the model responds.
Interactive Demo — The Atomic Fact Splitter
A real example from the paper (Table 7): one sentence of a model-generated bio. First decompose it into atomic facts — then verify each against Wikipedia and watch the FActScore build up, fact by fact.
PROMPT ▸ Tell me a bio of Ylona Garcia. GENERATION ▸ Ylona Garcia has since appeared in various TV shows such as ASAP (All-Star Sunday Afternoon Party), Wansapanataym Presents: Annika PINTAsera and Maalaala Mo Kaya.
Step 1: decompose the generation into atomic facts.
FActScore— / —
Four Verifier Variants (Table 3)
No-context LM: just "<atomic-fact> True or False?" — closed-book. Often overestimates, sometimes flips the model ranking.
Retrieve → LM: retrieved passages + fact + "True or False?" — retrieval consistently improves error rate and ranking.
NP (nonparametric probability): mask each token of the fact, score it with a masked LM over Wikipedia, threshold the average.
Retrieve → LM + NP: ensemble — Supported only if both agree. With Inst-LLaMA 7B this reaches error rates as low as 0.4–1.4%.
Why Retrieval Helps the Verifier
The evaluator LM has not memorized every fact — asking it closed-book mostly measures what it already believes. Showing it retrieved Wikipedia passages grounds the judgment in evidence, the same way a human checker reads the article first. With ChatGPT as the evaluator, Retrieve → LM keeps the error rate to ~5% and preserves the true ranking between evaluated models; the paper also tried question-generation prompting (as in QAGS/FEQA) and found True/False prompting more reliable, because generated questions are hard to control.
Chapter 04
Rare Entities, More Hallucination
FActScore made a long-suspected pattern measurable: factual precision tracks how often the model saw the entity during pretraining. The rarer the person, the more fabrication.
Interactive Demo — Frequency Effect Explorer
Pick an entity-frequency tier. The paper buckets people by a freqValue — the max of Wikipedia occurrence count and pageview count — into five tiers, and finds FActScore falls monotonically as rarity increases. Bar values below are illustrative approximations of Figure 2's trend.
Why Rarity Hurts
An LM's parametric memory is a compression of its pretraining data. A globally famous person appears thousands of times — the model stores dense, reliable facts. A niche figure appears a handful of times, so the model reconstructs the bio from patterns: plausible professions, plausible cities, plausible spouses. Those reconstructions are exactly the Not-supported atomic facts FActScore catches. The effect agrees with earlier frequency studies (Kandpal et al.), and is consistent across every model evaluated.
Position in the Generation Matters Too
Figure 2 (bottom) shows the later part of a generation has significantly worse precision. Early facts — nationality, profession — are the most repeated in pretraining data; as the bio reaches for specific dates, relationships, and career steps, the model is on thinner ice. Two risk factors, one lesson: test the long tail, and don't trust the tail end of long outputs.
ILLUSTRATIVE OF FIGURE 2 (BOTTOM): PRECISION ACROSS A GENERATION
early
·
·
·
late
Even Search Isn't Immune
A striking secondary finding: retrieval-augmented PerplexityAI also drops as entities get rarer — a relative drop of about 50% at the atomic level from the most frequent to the most rare tier. Retrieval at generation time helps a lot, but it does not fully solve the long tail: retrieval itself is harder when the entity is obscure, and the model still paraphrases (or over-summarizes) what it finds.
Chapter 05
The Scoreboard
First a human evaluation of three commercial systems; then the automated estimator unleashed on 13 subjects — 12 LMs plus human-written Wikipedia bios — across 500 entities each.
Human Evaluation (Table 1 — 183 entities, expert labels)
System
FActScore
Atomic facts / bio
Responds
InstructGPT (text-davinci-003)
42.5
26.3
99.5%
ChatGPT
58.3
34.7
85.8%
PerplexityAI (search-augmented)
71.5
40.8
90.7%
Human-written bio (Wikipedia, estimated)
~87
29.0
88.8%
The punchline: even PerplexityAI — with a commercial search engine and the right Wikipedia page within reach — fails to support ~28% of its atomic facts (11% explicitly not-supported, 15% irrelevant). Note that ChatGPT and PerplexityAI abstain from answering 14% and 9% of prompts, which quietly boosts their precision; InstructGPT almost never abstains. The human row comes from the 500-entity estimated study (Table 5 + Figure 3).
At Scale: 13 Subjects, 500 Entities Each (Figure 3, ChatGPT-with-retrieval estimator)
HUMAN (WIKIPEDIA BIOS)
~87%
FActScore
The bar every model misses by a wide margin
GPT-4
~72%
FActScore
Comparable to ChatGPT — but abstains less, says far more
~83% of its atomic facts unsupported, vs ~13% for human-written bios
Estimated scores from the paper's Figure 3 (approximate figure reads, marked ~). The two best estimator variants agree on the ranking with a Pearson correlation of 0.99.
Head-to-Head Findings
All LMs are substantially less factual than humans — against claims that LMs were approaching human performance, even on an easy task like writing a bio.
GPT-4 ≈ ChatGPT in precision, but GPT-4 abstains less (12% vs 16%) and packs in far more facts (61 vs 37 per bio).
Within a family, size tracks precision: Alpaca 65B > 13B > 7B, and Vicuna 13B > 7B.
Same size, different recipe, different score: at 7B, Alpaca/Vicuna (~40%) beat MPT-Chat (~30%) and StableLM (~17%).
Vicuna says ~3× more than Alpaca (51 vs 17 facts per bio) at similar precision — one reason Alpaca never abstains while Vicuna sometimes does.
What FActScore Did NOT Solve
Precision, not recall: a bio that says almost nothing can score ~100%. The paper's own Mary I example is highly precise yet skips how she returned to the line of succession — so it's a poor bio.
One source, one language: English Wikipedia; your trusted source may differ, and coverage gaps become false "not-supported" labels.
Equal weights: a wrong birth year costs exactly as much as a wrong middle initial.
Verification costs compute: each atomic fact needs retrieval + LM calls — cheap versus $26K of human labor, but not free.
Legacy
Impact — Factuality Becomes Measurable
FActScore gave the field a shared unit for long-form hallucination: the unsupported atomic fact. Everything that followed — better verifiers, factuality leaderboards, targeted decoding fixes — builds on that move.
📦 pip install factscore
The metric shipped as an open-source package, released together with the paper's human annotations — atomic-fact checking became a drop-in evaluation for anyone.
📊 Long-form factuality, quantified
"ChatGPT achieves 58%" turned a vague worry into a measurable gap. Fine-grained factuality sections spread through model cards and eval suites after this paper.
🔎 Retrieval, twice vindicated
Search-augmented PerplexityAI topped the human evaluation (71.5%), and retrieval also powered the best automated verifier — evidence that grounded generation and grounded checking both beat closed-book.
🧭 Risk factors mapped
Entity frequency and position-in-generation effects told eval designers where hallucinations live: the long tail of entities and the tail end of outputs.
💸 Evaluation economics
6,500 generations from 13 subjects, estimated automatically with <2% error — annotation that would have cost $26K at the paper's $4-per-generation human rate.
🧬 The lineage
Follow-ups automated and extended the recipe further — search-based verifiers (e.g., SAFE), factuality benchmarks (e.g., FELM), and multilingual/long-form variants — with FActScore as the reference point.
Deep Dive
Decompose, Verify, Count
FActScore changed the unit of measurement. Instead of asking "is this bio true?", it splits generations into atomic facts and asks "what fraction can be individually verified?" — turning factual precision from a judgment into a ratio with a long tail.
🔬
The Atomic Advantage
Granularity locates the lie: a 90% FActScore bio tells you which 10% is fabricated
Decomposition + retrieval + a True/False judge matches human labels within a <2% error rate
One number comparable across systems: InstructGPT 42.5% · ChatGPT 58.3% · search-augmented PerplexityAI 71.5%
Human-written bios set the ceiling: ~87% — even people don't fully verify
⚠️
Precision Is Not the Whole Story
A vague-but-true bio scores ~100% — the metric doesn't reward informativeness (no recall term)
Risk rises for rare entities: Wikipedia thins out, the verifier starves
Facts late in a generation are more likely fabricated — position is a risk factor
The estimator itself costs retrievals + LLM verdicts per atom — precision is bought per fact
Interactive Demo — The Atom Verifier
A one-sentence biography decomposes into atoms exactly as the paper does. Each atom is checked against retrieved Wikipedia evidence — green supported, red unsupported. The counter is FActScore being computed in front of you. Then see where real models land on the same yardstick.
"Marie Curie was a Polish-born physicist who conducted pioneering research on radioactivity, won two Nobel Prizes, discovered the elements polonium and radium, and won the Nobel Prize in Economics in 1911."
FActScore = supported atoms / total atoms
VERDICT
Fluency masks fabrication — granularity unmasks it
The 42.5% InstructGPT bios read flawlessly; that's what made the number shocking. By counting at the atom level, FActScore showed that fluent long-form generation is a precision problem that worsens with length and rarity — and that retrieval (RAG, Track II) and sampling checks (SelfCheckGPT) are attacking the same failure from two sides. Where this story continues: HaluEval tests whether models can recognize the fabrication; RAGTruth moves the whole question into the RAG pipeline.
🧩 What counts as an atom
One verifiable claim, stripped of hedging: "Curie was born in Poland", "she discovered radium". Decomposition quality bounds the whole metric — bad splits make unanswerable atoms.
📉 The long tail
At 500 entities × 13 subjects, best public models sat at 40–55%, GPT-4 ~72%, StableLM ~17%. The distribution — not the median — is the finding: everyone fails somewhere.
⏳ Position risk
Atoms late in a generation verify worse — consistent with snowballing from the hallucination survey. Length is not free even for top models.
🛠️ It shipped as tooling
pip install factscore — the methodology became an eval harness teams run before release, which is the highest compliment a metric can receive.
Test Yourself
Quick Quiz
Check your understanding of the key concepts from the FActScore paper.
Reference
Key Takeaways
Everything you need to remember about this paper.
✅ FActScore = the percentage of a generation's atomic facts supported by a knowledge source you choose (Wikipedia in the paper).
✅ Human evaluation (183 bios): InstructGPT 42.5% · ChatGPT 58.3% · search-augmented PerplexityAI 71.5% — fluent does not mean factual.
✅ Estimated at scale (500 entities × 13 subjects): human bios ~87%, GPT-4 ~72%, best public models ~40–55%, StableLM ~17%.
✅ The automated estimator — GTR retrieval + True/False prompting of a strong LM — matches human labels within a <2% error rate.
✅ Hallucination risk rises for rare entities and for facts late in a generation; retrieval helps but doesn't erase the long tail.
✅ It measures precision, not recall — a vague-but-true bio can score ~100%. Code + human annotations: pip install factscore.