History Problem Core Idea Why It Works Results Cost Impact Quiz
Interactive Paper Explainer

Sample, Vote, Repeat
Self-Consistency

Chain-of-thought prompting asks the model a question once — and prays the single reasoning path is right. Self-consistency asks it 40 times, then takes the vote. Complex reasoning jumps by up to +17.9 points without changing the model at all.

Start Learning Read the Paper ↗
+17.9
GSM8K Points (56.5 → 74.4)
40
Sampled Reasoning Paths (k)
0
Training Steps Required
2022
Year on arXiv (ICLR 2023)
History

From One Path to Many

Self-consistency landed three months after chain-of-thought prompting — and removed the last weakness that paper left behind: the model only ever got one attempt.

2022 · Jan
Chain-of-Thought Prompting (Wei et al.)
A few worked examples in the prompt make large models reason step by step — read the CoT guide. But decoding still returns exactly one reasoning path.
2022 · early
Greedy decoding is the default
Prompting papers read the single path the model produces. Everyone evaluates argmax — one roll of the dice, one answer, no retry.
2022 · Mar
🚀 Self-Consistency (Wang et al., Google)
Sample diverse reasoning paths with temperature, extract the final answers, and marginalize: take the majority vote. GSM8K jumps +17.9 points with zero training.
2022–23
Voting goes universal
Follow-up work extends the idea beyond discrete answers — universal self-consistency adapts voting to open-ended generation.
2023 →
A standard tool
Self-consistency becomes a core component of reasoning benchmarks and a filter for generating and cleaning training data.
2024 →
The test-time compute era
Majority-voting ideas echo in o1-style and R1-style reasoning models, which spend extra inference compute on purpose — thinking longer becomes a scaling axis.
Key Insight

A language model with temperature is a distribution over reasoning paths, not a single answer. Greedy decoding reads one path and commits. Self-consistency samples that distribution and marginalizes the paths out: the answer is whichever value most paths agree on.

TWO DECODING RULES, SAME MODEL
greedy:  take the single most-likely path → its answer
SC:      sample k diverse paths → vote over their answers
Chapter 01

The Problem — One Roll of the Dice

Chain-of-thought prompting made reasoning visible. But the standard decoder takes exactly one shot at it — and a single early slip decides the final answer.

☝️
One Greedy Path
  • Greedy decoding commits to the single most-likely reasoning path
  • One slip — a misread number, a bad division — ruins every step after it
  • Most problems admit many valid routes; greedy bets everything on one
  • No signal for whether the final answer is reliable
🗳️
Many Paths, One Vote
  • Sample k = 40 diverse reasoning paths with temperature
  • Wrong paths derail differently — their answers scatter
  • Right paths take different routes but converge on the same answer
  • The vote is robust — and its margin is a free confidence estimate
Anatomy of a Doomed Path
8 × 8 = 64 ✓
64 − 28 = 36 ✓
36 ÷ 3 jars = 12 ✗
→ final answer: 12

Two correct steps, then a misread — 4 jars becomes 3 jars — and every remaining step inherits the error. Greedy decoding ships "12" with exactly the confidence it would ship "9". There is no second opinion, no retry, and no signal that anything went wrong.

Chapter 02

The Core Idea — Sample & Marginalize

Self-consistency changes the decoding rule and nothing else: same model, same weights, same chain-of-thought prompt. Instead of greedily taking one reasoning path, sample many and marginalize them out with a vote.

The Recipe — Four Steps, Zero Training
✍️
1 · Prompt Once
The usual chain-of-thought prompt with worked exemplars — unchanged.
🌡️
2 · Sample k Paths
Decode k diverse reasoning paths (up to 40) with temperature instead of greedy argmax.
🔎
3 · Extract Answers
Parse the final answer from each path — a number, an option, a short token.
🗳️
4 · Take the Vote
Output the most frequent final answer — the mode of the sampled paths.
P(a | q)  =  Σi  P(a | pathi, q) · P(pathi | q)
approximated by the vote:  pick argmaxa Σi 1[ answer(pathi) = a ]
pathi
Reasoning Path
One of the k sampled chains of thought for question q — each a fully worked solution.
P(pathi | q)
Path Prior
How likely the sampler is to generate path i. Temperature spreads this mass over diverse routes.
Σi
Marginalization
Sum out the reasoning path: the answer should not depend on which route the model happened to draw.
argmax vote
Empirical Mode
With k samples, vote counts approximate the marginal — the most frequent final answer wins.
Interactive Demo — Vote With Me

A grade-school problem, a batch of model samples, and one vote. Choose how many reasoning paths to draw — then watch the answers land and the histogram settle.

Q. Nadia has 8 bags with 8 marbles in each. She gives 28 marbles to her cousin, then splits the rest equally into 4 jars. How many marbles are in each jar?
sample

Illustrative — paths are precomputed samples in the spirit of the paper: different routes, same question. Green = lands on the true answer (9); red = derailed.

Chapter 03

Why the Vote Works

Voting over random samples sounds too simple to work. It works because of an asymmetry the paper puts at the center: correctness is self-consistent, errors are not.

✅ Correctness Converges

A right answer can be reached by genuinely different derivations — direct arithmetic, per-jar reasoning, working backwards from a guess and checking it. Dozens of routes, one destination. The vote banks exactly this convergence: each correct path casts a ballot for the same value.

❌ Errors Scatter

A wrong path usually fails for its own local reason — a misread jar count, a hallucinated premise, a slipped division. Each derailment lands somewhere different, so no single wrong answer ever collects a majority. If 70% of paths are right and the 30% wrong mass splits three ways, the truth wins with room to spare.

Analogy — The 40 Engineers

Ask 40 engineers to solve the same problem independently: the consensus is usually right even when individuals err — not because engineers are infallible, but because correct solutions coincide while mistakes are personal. Self-consistency gives a language model the same privilege: temperature buys independence between paths, and the vote collects the wisdom. Bonus: the vote margin is a free confidence estimate — a narrow vote means an uncertain model, before you ever see the label.

Interactive Demo — Watch the Vote Stabilize

Start at k = 1 — a single sample can be flat wrong. Drag up and watch the correct answer (9) take a stable majority while the wrong answers starve.

k=110203040

Illustrative — a fixed GSM8K-style answer distribution: ~70% of paths land on 9, the rest scatter across 12, 3 and 59. Watch the vote share of the correct answer stabilize as k grows.

Chapter 04

Results — +17.9 Points, Same Weights

The paper evaluates self-consistency on arithmetic, commonsense and symbolic reasoning. The pattern repeats everywhere: large, consistent gains with zero training and zero changes to the model or prompt.

GSM8K · PaLM-540B + SC
74.4
with 40 sampled paths (majority vote)
+17.9 absolute over greedy CoT
GSM8K · CoT (greedy)
56.5
same model, same prompt, single path
the baseline the vote beats
TYPICAL GAIN · OTHER BENCHMARKS
+8 to +18
absolute points, AQuA / SVAMP / StrategyQA and more
arithmetic, commonsense & symbolic
TRAINING REQUIRED
0
steps, new parameters, new prompts — none
works with standard samplers out of the box
Chain-of-Thought Decoding — Greedy vs Self-Consistency
BenchmarkReasoning TypeCoT (greedy)CoT + Self-Consistency
GSM8K · PaLM-540Bgrade-school math56.574.4  (+17.9)
AQuAmath word problems (multiple choice)—+8 to +18 pts (typical)
SVAMPelementary math word problems—+8 to +18 pts (typical)
StrategyQAimplicit multi-step commonsense—+8 to +18 pts (typical)

Only the GSM8K headline (PaLM-540B, k = 40, same prompt, same weights) is quoted exactly. The paper reports consistent absolute gains in roughly the +8 to +18 point range across its arithmetic, commonsense and symbolic benchmarks — exact deltas vary by model, prompt and k, so they are summarized here as a range. On several benchmarks, self-consistency also beat systems that had been fine-tuned for the task.

Bigger Than Training

With self-consistency, PaLM-540B outperformed or matched much larger models — and on several benchmarks, models that had been fine-tuned for the task — without a single gradient step. The reasoning ability was already in the weights; greedy decoding was simply throwing most of it away.

Robust Across Samplers

The gains do not hinge on one lucky sampler: temperature, top-k and nucleus sampling all worked. What matters is diversity between paths — enough for errors to decorrelate, not so much that the chains turn into nonsense.

Chapter 05

The Catch — k× the Compute

Self-consistency buys accuracy with test-time compute. That trade is the paper's most modern idea: accuracy as a function of the inference budget, not just of parameters.

Interactive Demo — Pay k×, Gain a Saturating Curve

Drag the number of sampled paths: cost grows linearly, accuracy saturates. Then flip the toggle to see how little the 40th path actually buys.

k = 10
INFERENCE COST
10×
ACCURACY (GSM8K-STYLE)
69.4

Illustrative saturating curve, anchored to the paper's GSM8K headline (56.5 → 74.4 at k = 40). The linear cost line is exact.

💸 k× the tokens
Sampling k = 40 paths means decoding roughly 40× the tokens of a single greedy pass — cost and latency scale linearly with k. Self-consistency is a test-time compute budget, spent on purpose.
📉 Flattening returns
Accuracy rises fast, then flattens: most of the gain typically arrives well before k = 40. Treat k as a dial and stop where the curve stops paying for itself.
📝 Needs a discrete answer
The vote requires extractable final answers — a number, an option. The paper scopes self-consistency to problems with a discrete answer; open-ended generation, where several defensible outputs exist, breaks the mechanism and needed follow-up work.
🔌 Drop-in decoding
No training, no model changes, no special prompt engineering. Pairs with any chain-of-thought exemplars and standard samplers — temperature, top-k, nucleus.
Legacy

Impact — Test-Time Compute

Self-consistency was the first clean demonstration that inference compute is a scaling axis of its own. The reasoning models of 2024 and beyond are built on that insight.

📈 The test-time compute era
The vote was the proof of concept: more inference compute → more accuracy, same weights. o1-style and DeepSeek-R1-style reasoning models later made this structural — thinking longer, on purpose.
🗳️ Voting goes default
Majority voting and best-of-k selection became standard practice in benchmark evaluation, inference APIs, and anywhere a discrete answer is needed.
🏭 Data factories
Sample many paths, keep the consistent ones: self-consistency became a quality filter for generating synthetic reasoning data and a supervision signal for training.
🧠 Ensembles, made rigorous
An old intuition — many weak predictors beating one — formalized for a single model's sampled reasoning paths. One model, many minds.
🎁 Confidence for free
The vote margin doubles as an uncertainty estimate: a wide majority means a confident model, a narrow vote means "sample more or give up" — with no extra machinery.
🔗 Why models can do this at all
Multi-step in-context reasoning rests on circuitry like induction heads, which copy and extend patterns from the prompt. For the mechanistic complement to this decoding story, read our Induction Heads guide.
Test Yourself

Quick Quiz

Five questions on sampling, voting, and the price you pay.

Reference

Key Takeaways

Everything worth remembering about self-consistency.

✅ Swap the decoding rule: sample k diverse chain-of-thought paths, then take the majority vote over final answers.
✅ It is marginalization: P(a|q) = Σᵢ P(a|pathᵢ,q)·P(pathᵢ|q) — the vote approximates the sum with k samples.
✅ Why it works: correct answers converge from many routes; errors scatter — plurality picks the truth.
✅ GSM8K: 56.5 → 74.4 (+17.9 absolute) with PaLM-540B — same weights, same prompt, zero training.
✅ The cost is k× inference compute; gains saturate, so most value arrives well before k = 40.
✅ Limits: needs a discrete final answer — open-ended generation breaks the vote.