A visual, step-by-step guide to the paper that taught LLMs to fact-check themselves — sample the same model many times, and watch hallucinations contradict their own alternatives. No external database, no token probabilities, no extra annotators.
SelfCheckGPT arrived after three years of hand-wringing about hallucinations. Each earlier fix needed something extra — internals, databases, annotators. This one needed only the model itself.
A language model that truly knows a concept has a peaked probability distribution: every sample lands near the same facts. A model that doesn't know has a flat distribution — each sample drifts to a different invention. Sampling is a probe into what the model actually memorized.
Every pre-2023 hallucination detector demanded something you don't have when you're talking to a chatbot through an API. SelfCheckGPT's whole bet is to need nothing but the conversation itself.
Ask a student the same question five separate times. If they know, the story stays stable — details wiggle, facts don't. If they're confabulating, every retell invents something new. An LLM is no different.
Generate the main response deterministically, then draw N stochastic samples of the same prompt. For every sentence in the response, measure how well the samples support it. Unsupported sentences are the hallucinations.
The main response is generated at temperature 0.0 with beam search; the N=20 samples use the same prompt at temperature 1.0 — random sampling. If sampling were deterministic, every "sample" would be identical and the probe would be useless. In the paper's WikiBio setup the prompt is simply: "This is a Wikipedia passage about {concept}:"
Each sentence gets 𝒮(i) ∈ [0.0, 1.0] — 0.0 = grounded in the model's knowledge, 1.0 = contradicted by every sample. Sentences can then be thresholded individually, and a passage-level factuality score (the average) ranks whole passages — the paper shows both work.
The recipe is fixed: response + N samples. The free choice is how to decide whether a sentence is "consistent" with the samples. The paper ships five interchangeable scorers — four classic ones plus asking an LLM directly.
Embed every sentence (RoBERTa-Large backbone) and compare the response sentence with the most similar sentence in each sample. If even the best matches are dissimilar, nothing in the samples supports the sentence.
Generate multiple-choice questions about the sentence (MQAG framework), then let an answering system answer each question using each sample. If the samples give different answers than the response — inconsistent.
Train a tiny n-gram language model on the samples themselves (response tokens get +1 smoothing). A sentence full of words the samples never use gets a low probability. The cheapest variant needs no extra model at all: unigram (max) — find the sentence's rarest word and count its occurrences.
Treat each sample as a premise and the sentence as a hypothesis; a DeBERTa-v3-large model fine-tuned on MNLI estimates P(contradiction), using only the entailment and contradiction logits. Average over all samples.
Skip the machinery: paste a sample into an LLM (GPT-3 or ChatGPT) plus the sentence and ask "Is the sentence supported by the context above? Answer Yes or No." Map {Yes: 0.0, No: 1.0, N/A: 0.5} and average. GPT-3 answered Yes/No 98% of the time — but less capable models (text-curie-001, LLaMA) failed at this consistency assessment, so the judge needs to be strong.
Six steps, one model, zero external resources. This is the full recipe as run on the WikiBio GPT-3 dataset — and the same recipe you could run against any API today.
The paper varies the sample count: performance rises smoothly with N and shows diminishing gains — and the n-gram variant needs the most samples before it plateaus. Sampling cost buys accuracy on a predictable curve, so you pick your own point on it.
GPT-3 wrote Wikipedia-style biographies for 238 real people; humans labelled every sentence. The headline: the model hallucinated a lot — and sampling caught it.
Inter-annotator agreement (Cohen's κ): 0.748 for the 2-class task, 0.595 for 3-class. The passage-level histogram even shows a sharp peak at score 1.0 — "total hallucination", where the whole biography is fabricated.
| Method | Access | AUC-PR (NonFact) | Pearson | Spearman |
|---|---|---|---|---|
| Random baseline | — | 72.96 | — | — |
| GPT-3 Avg(−log p) | grey-box | 83.21 | 57.04 | 53.93 |
| GPT-3 Max(−log p) | grey-box | 87.51 | 57.83 | 55.69 |
| LLaMA-30B Avg(−log p) | black-box proxy | 75.43 | 21.72 | 20.20 |
| SelfCheckGPT w/ BERTScore | black-box | 81.96 | 58.18 | 55.90 |
| SelfCheckGPT w/ QA | black-box | 84.26 | 61.07 | 59.29 |
| SelfCheckGPT w/ Unigram (max) | black-box | 85.63 | 64.71 | 64.91 |
| SelfCheckGPT w/ NLI | black-box | 92.50 | 74.14 | 73.78 |
| SelfCheckGPT w/ Prompt | black-box | 93.42 | 78.32 | 78.30 |
Pearson / Spearman = correlation with human passage-level factuality judgements. SelfCheckGPT beats the grey-box probability baselines while needing only text — and proxy-LLM probabilities transfer badly.
SelfCheckGPT showed that the cheapest possible probe — asking again — is a competitive hallucination detector. Its fingerprints are all over modern LLM-trust tooling.
SelfCheckGPT's insight fits in one line: a model that knows a fact re-derives it under every resampling; a model that's fabricating invents a different story each time. Sampling — the cheapest thing an LLM does — becomes a lie detector.
Check your understanding of the key concepts from the SelfCheckGPT paper.
Everything you need to remember about this paper.