History Problem Method Variants Pipeline Results Impact Deep Dive Quiz
Interactive Paper Explainer

The Model Checks Itself
SelfCheckGPT

A visual, step-by-step guide to the paper that taught LLMs to fact-check themselves — sample the same model many times, and watch hallucinations contradict their own alternatives. No external database, no token probabilities, no extra annotators.

Start Learning Read the Paper ↗
5
Scoring Variants
N=20
Samples Drawn
93.4
Best Sentence AUC-PR
0
External Databases
History

From Confident Errors to Self-Checking

SelfCheckGPT arrived after three years of hand-wringing about hallucinations. Each earlier fix needed something extra — internals, databases, annotators. This one needed only the model itself.

2020
GPT-3: fluent but unfaithful
175B parameters make GPT-3 stunningly fluent — and equally confident when inventing facts. Fluency stops being a signal of truth.
2021–22
External fact-checking pipelines
Claim detection → evidence retrieval → verdict. Powerful, but bolted onto external databases with complex modules — and datasets built by perturbing real text (Liu et al., 2022) may not reflect true LLM hallucination.
2022
Sampling + self-evaluation
Self-consistency (Wang et al.) votes over sampled reasoning paths; Kadavath et al. show LLMs can estimate their own correctness — but many of these methods need token probabilities or careful prompting.
2023 · Mar
🚀 SelfCheckGPT preprint
Manakul, Liusie & Gales (University of Cambridge): sample N stochastic responses, measure consistency, flag hallucinations — zero-resource and black-box.
2023 · Dec
EMNLP 2023 + dataset release
Published at the EMNLP 2023 main conference, with the wiki_bio_gpt3_hallucination dataset (238 GPT-3 passages, 1,908 annotated sentences) and open-source code.
2024 →
The self-verification wave
Sampling-based detection and self-checking became a standard tool in the LLM-trust toolbox, with a long tail of follow-ups chasing cheaper and sharper consistency checks.
Key Insight

A language model that truly knows a concept has a peaked probability distribution: every sample lands near the same facts. A model that doesn't know has a flat distribution — each sample drifts to a different invention. Sampling is a probe into what the model actually memorized.

KNOWN VS UNKNOWN CONCEPT (ILLUSTRATIVE)
"Lionel Messi is a _" → footballer (peaked — model knows)
"John Smith is a _" → flat distribution (model guesses)
A guess sampled five times gives five different answers. That drift is the hallucination signal.
Chapter 01

The Problem — Fluent but Unfaithful

Every pre-2023 hallucination detector demanded something you don't have when you're talking to a chatbot through an API. SelfCheckGPT's whole bet is to need nothing but the conversation itself.

🧪
Detecting Hallucinations the Old Way
  • White-box: truth classifiers over hidden states (Azaria & Mitchell, 2023) need model internals — never exposed by APIs — plus labelled training data
  • Grey-box: token probabilities and entropy correlate with factuality, but many APIs (e.g. ChatGPT at the time) return text only
  • Fact verification: retrieve evidence from external databases — complex multi-stage pipelines, and facts can only be checked against what the database covers
  • Proxy models: run a second open-source LLM to score the first — its probabilities transfer poorly (LLaMA-30B correlations were weak)
  • Benchmarks: datasets built by perturbing factual text may not reflect how LLMs actually hallucinate
🔁
SelfCheckGPT's Bet
  • Needs only what you already have: the model and its own outputs
  • Black-box: works through any text-only API — no logits, no internals
  • Zero-resource: no external database, no retrieval stack, no annotators in the loop
  • Sentence-level scores in [0,1] — find which sentence is the lie, not just that the passage is bad
  • Free by-product: average the sentence scores and you can rank whole passages by factuality
The Cross-Examination Card

Ask a student the same question five separate times. If they know, the story stays stable — details wiggle, facts don't. If they're confabulating, every retell invents something new. An LLM is no different.

KNOWN TOPIC
Q: "When was the telescope invented?"
A¹: "1608, in the Netherlands."
A²: "1608 — Dutch spectacle-makers."
A³: "1608, credited to Lipperhey."
→ consistent → likely known
UNKNOWN TOPIC
Q: "When was the Felton Scope invented?"
A¹: "1874, in Austria."
A²: "1912, in Britain."
A³: "1820, in France."
→ divergent → likely hallucinated
Chapter 02

The Core Idea — Sample and Compare

Generate the main response deterministically, then draw N stochastic samples of the same prompt. For every sentence in the response, measure how well the samples support it. Unsupported sentences are the hallucinations.

𝒮BERT(i) = 1 − (1/N) Σn=1N maxk BERTScore( ri , skn )
SelfCheckGPT with BERTScore — Equation (1) of the paper. The other variants swap the comparison, not the recipe.
ri
Sentence i of the response
The main output R is split into sentences; each one is assessed on its own.
skn
Sentence k of sample n
One of the N stochastic re-generations of the same prompt (S¹…Sᴺ).
maxk
Best match per sample
Compare ri against every sentence in sample n and keep the most similar one.
1 − avg
Inconsistency score
Flip the average similarity: 0.0 = fully supported (factual), 1.0 = no sample agrees (hallucinated).
Black-box, Literally

The main response is generated at temperature 0.0 with beam search; the N=20 samples use the same prompt at temperature 1.0 — random sampling. If sampling were deterministic, every "sample" would be identical and the probe would be useless. In the paper's WikiBio setup the prompt is simply: "This is a Wikipedia passage about {concept}:"

What the Score Measures

Each sentence gets 𝒮(i) ∈ [0.0, 1.0] — 0.0 = grounded in the model's knowledge, 1.0 = contradicted by every sample. Sentences can then be thresholded individually, and a passage-level factuality score (the average) ranks whole passages — the paper shows both work.

selfcheckgpt · judge console — interactive demo 1
One prompt, five samples. Toggle the concept and watch what sampling reveals: a known fact keeps re-appearing; a hallucinated fact dissolves into five different stories.
CONSISTENCY ACROSS SAMPLES awaiting samples…
Illustrative demo — the paper's real runs draw N=20 samples per passage with GPT-3 (text-davinci-003) and score every sentence.
Chapter 03

Five Ways to Measure Consistency

The recipe is fixed: response + N samples. The free choice is how to decide whether a sentence is "consistent" with the samples. The paper ships five interchangeable scorers — four classic ones plus asking an LLM directly.

① BERTScore — similarity

Embed every sentence (RoBERTa-Large backbone) and compare the response sentence with the most similar sentence in each sample. If even the best matches are dissimilar, nothing in the samples supports the sentence.

score = 1 − avg( best BERTScore per sample )
② QA — question the facts

Generate multiple-choice questions about the sentence (MQAG framework), then let an answering system answer each question using each sample. If the samples give different answers than the response — inconsistent.

score = #answer-mismatches / #questions
③ n-gram — word counts

Train a tiny n-gram language model on the samples themselves (response tokens get +1 smoothing). A sentence full of words the samples never use gets a low probability. The cheapest variant needs no extra model at all: unigram (max) — find the sentence's rarest word and count its occurrences.

score ∝ −log P(sentence | samples)
④ NLI — entailment

Treat each sample as a premise and the sentence as a hypothesis; a DeBERTa-v3-large model fine-tuned on MNLI estimates P(contradiction), using only the entailment and contradiction logits. Average over all samples.

score = avg( P(contradict | sample, sentence) )
⑤ Prompt — ask the model directly

Skip the machinery: paste a sample into an LLM (GPT-3 or ChatGPT) plus the sentence and ask "Is the sentence supported by the context above? Answer Yes or No." Map {Yes: 0.0, No: 1.0, N/A: 0.5} and average. GPT-3 answered Yes/No 98% of the time — but less capable models (text-curie-001, LLaMA) failed at this consistency assessment, so the judge needs to be strong.

Interactive Demo — Variant Showdown

Two sentences from one response — one factual, one hallucinated. Pick a variant to see exactly how it turns samples into a score, and where it lands on the paper's leaderboard.

SENTENCE-LEVEL AUC-PR · NON-FACTUAL DETECTION (WIKIBIO, TABLE 2)
Chapter 04

The Pipeline, Step by Step

Six steps, one model, zero external resources. This is the full recipe as run on the WikiBio GPT-3 dataset — and the same recipe you could run against any API today.

STEP 1 · ASK ONCE
Prompt the model at temperature 0.0 with beam search → the main response R, e.g. a full biography passage.
STEP 2 · SAMPLE N=20
Same prompt, temperature 1.0 → stochastic samples S¹…S²⁰. Each is a fresh roll of the model's dice.
STEP 3 · SPLIT
Break R into sentences r₁ … r_I. Detection happens per sentence — that's what makes the output actionable.
STEP 4 · SCORE
For every ri, compare against all 20 samples with any variant → inconsistency score 𝒮(i) ∈ [0,1].
STEP 5 · THRESHOLD
Sentences whose score crosses the cut are flagged: ⚠ likely hallucination. Flag, rewrite, or drop them.
STEP 6 · RANK
Average the sentence scores → a passage-level factuality score. Rank thousands of passages without any labelling.
More Samples, Better — Smoothly

The paper varies the sample count: performance rises smoothly with N and shows diminishing gains — and the n-gram variant needs the most samples before it plateaus. Sampling cost buys accuracy on a predictable curve, so you pick your own point on it.

Dataset Glance — WikiBio GPT-3
238
passages generated
1,908
annotated sentences
184.7
tokens / passage (±36.9)
20
samples per passage
Chapter 05

Results — Detecting the Drift

GPT-3 wrote Wikipedia-style biographies for 238 real people; humans labelled every sentence. The headline: the model hallucinated a lot — and sampling caught it.

How Bad Was It? (Human Annotation)
39.9%
major-inaccurate — fully hallucinated sentences
33.1%
minor-inaccurate — related but partly wrong
27.0%
accurate sentences
~73%
of all sentences at least partly non-factual

Inter-annotator agreement (Cohen's κ): 0.748 for the 2-class task, 0.595 for 3-class. The passage-level histogram even shows a sharp peak at score 1.0 — "total hallucination", where the whole biography is fabricated.

w/ BERTSCORE
81.96
sentence-level AUC-PR
w/ QA
84.26
sentence-level AUC-PR
w/ UNIGRAM (MAX)
85.63
sentence-level AUC-PR
w/ NLI
92.50
sentence-level AUC-PR
w/ PROMPT
93.42
sentence-level AUC-PR — best overall
Table 2 (excerpt) — Sentence-level Detection & Passage-level Ranking
MethodAccessAUC-PR (NonFact)PearsonSpearman
Random baseline—72.96——
GPT-3 Avg(−log p)grey-box83.2157.0453.93
GPT-3 Max(−log p)grey-box87.5157.8355.69
LLaMA-30B Avg(−log p)black-box proxy75.4321.7220.20
SelfCheckGPT w/ BERTScoreblack-box81.9658.1855.90
SelfCheckGPT w/ QAblack-box84.2661.0759.29
SelfCheckGPT w/ Unigram (max)black-box85.6364.7164.91
SelfCheckGPT w/ NLIblack-box92.5074.1473.78
SelfCheckGPT w/ Promptblack-box93.4278.3278.30

Pearson / Spearman = correlation with human passage-level factuality judgements. SelfCheckGPT beats the grey-box probability baselines while needing only text — and proxy-LLM probabilities transfer badly.

Interactive Demo — Sentence-Level Scanner

A five-sentence passage (illustrative). Run the check: each sentence is scored against the samples, and anything inconsistent gets flagged — just like Step 5 of the pipeline.

Scores are illustrative; the real pipeline scores each sentence against N=20 sampled responses.
🔍 Grey-box is strong — if you can get it
GPT-3's own token probabilities detect factual sentences at 53.97 AUC-PR (random: 27.04). The catch: most APIs never expose them.
⚡ The cheapest variant over-delivers
Unigram (max) just looks for the rarest word in the sentence across all 20 samples — "if a token appears once, it's likely non-factual" — and still scores 85.63.
🪞 The model can self-check
An ablation with only 4 samples: GPT-3 prompted to judge its own sentences beats the unigram method — and ChatGPT is slightly better still.
📚 Samples vs stored knowledge
Using the real WikiBio reference text instead of self-samples: BERTScore/QA perform comparably or better with self-samples; NLI/Prompt benefit from retrieval — but external databases can't cover every use case.
What SelfCheckGPT Did NOT Solve
Legacy

Impact — Sampling as a Lie Detector

SelfCheckGPT showed that the cheapest possible probe — asking again — is a competitive hallucination detector. Its fingerprints are all over modern LLM-trust tooling.

💬 Chat reliability
A practical pre-trust filter: sample, compare, and only surface answers that survive their own cross-examination.
🌊 Sampling-based evaluation wave
From self-consistency to semantic uncertainty — sampling turned from a decoding trick into a measurement instrument for model knowledge.
🔌 Black-box-friendly design
Released at the dawn of the ChatGPT era, its API-only assumption aged perfectly: every closed model since is a black box.
📦 Open dataset + code
wiki_bio_gpt3_hallucination (238 annotated passages) and the selfcheckgpt toolkit became shared testbeds for hallucination research.
➕ Complements retrieval
Retrieval-grounded checking covers what a database knows; self-checking covers everything else — the paper shows the two can combine.
💸 The cost trade-off
Sampling is simple but never free — follow-up work keeps chasing the same accuracy with fewer samples and cheaper judges.
Deep Dive

Consistency Is a Truth Signal

SelfCheckGPT's insight fits in one line: a model that knows a fact re-derives it under every resampling; a model that's fabricating invents a different story each time. Sampling — the cheapest thing an LLM does — becomes a lie detector.

⚫
What Black-Box Purity Buys
  • No knowledge base, no logits, no white-box access, no extra annotators
  • Works through any API: sample N times, compare, score
  • Five interchangeable scorers — BERTScore, QA, n-gram, NLI, LLM-prompt — best: Prompt 93.42 and NLI 92.50 AUC-PR
  • Sentence-level scores 𝒮(i) ∈ [0,1] rank where the passage lies, not just that it lies
💸
The Bill, and the Blind Spot
  • Detection costs N extra generations per passage (N=20 in the paper) — accuracy bought with compute
  • Consistency is not truth: a model can repeat a popular myth with high agreement — the TruthfulQA trap
  • Divergence needs the fact to be borderline-known; obscure truths sampled badly look like lies
  • Returns diminish smoothly with N — there is no cheap setting that also works
Interactive Demo — The Sampling Spectroscope

Pick a claim, then resample the model. Each card is one sampled continuation voting on the claim; the agreement meter computes the hallucination score the way SelfCheckGPT does: 1 minus average support. One run is a detector; the difference between runs is the whole method.

VERDICT
On WikiBio, GPT-3 was at least partly non-factual in ~73% of annotated sentences
A black-box detector that runs on 73% of long-form content being partly fabricated is not a research toy — it's a production gate. The paper's later SelfCheck-RAG line uses exactly this gate to filter generations before they reach users. The caveat stands: agreement measures the model's knowledge boundary, not the world's truth — pair it with grounded checking (FActScore) when the stakes are real.
🧠 The knowledge boundary
Samples converge on facts inside training distribution and scatter outside it. Divergence is a map of what the model doesn't actually know — which is why it doubles as an uncertainty estimate.
🔀 Scorer interchangeability
The sampling mechanism is fixed; the comparison layer is swappable. NLI and LLM-prompt won AUC-PR (92.50 / 93.42), but a cheap n-gram scorer still works when compute is scarce.
📈 The N-curve
Performance climbs smoothly from N=1 to N=20 with diminishing returns — a dial you turn per dollar. Benchmarks like HaluEval exist precisely to price that dial.
🪞 Consistent ≠ true
An imitative falsehood (the TruthfulQA story) samples with high agreement. Consistency detects fabrication, not deception — know which one you're guarding against.
Test Yourself

Quick Quiz

Check your understanding of the key concepts from the SelfCheckGPT paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Known concept → consistent samples; hallucinated fact → divergent, contradictory samples. That single asymmetry is the detector.
✅ Zero-resource & black-box: only the model's own sampled text is used — no database, no logits, no extra annotators.
✅ Sentence-level scores 𝒮(i) ∈ [0,1]: high = inconsistent with the samples = likely hallucinated; averaging gives passage-level ranking.
✅ Five interchangeable scorers: BERTScore, QA, n-gram, NLI, LLM-prompt — NLI (92.50) and Prompt (93.42) AUC-PR led the table.
✅ On WikiBio, GPT-3 was at least partly non-factual in ~73% of annotated sentences; the dataset (238 passages, 1,908 sentences) is public.
✅ The price is sampling: performance climbs smoothly with N (20 in the paper) with diminishing returns — accuracy bought with compute.