Chain-of-thought prompting asks the model a question once — and prays the single reasoning path is right. Self-consistency asks it 40 times, then takes the vote. Complex reasoning jumps by up to +17.9 points without changing the model at all.
Self-consistency landed three months after chain-of-thought prompting — and removed the last weakness that paper left behind: the model only ever got one attempt.
A language model with temperature is a distribution over reasoning paths, not a single answer. Greedy decoding reads one path and commits. Self-consistency samples that distribution and marginalizes the paths out: the answer is whichever value most paths agree on.
Chain-of-thought prompting made reasoning visible. But the standard decoder takes exactly one shot at it — and a single early slip decides the final answer.
Two correct steps, then a misread — 4 jars becomes 3 jars — and every remaining step inherits the error. Greedy decoding ships "12" with exactly the confidence it would ship "9". There is no second opinion, no retry, and no signal that anything went wrong.
Self-consistency changes the decoding rule and nothing else: same model, same weights, same chain-of-thought prompt. Instead of greedily taking one reasoning path, sample many and marginalize them out with a vote.
Voting over random samples sounds too simple to work. It works because of an asymmetry the paper puts at the center: correctness is self-consistent, errors are not.
A right answer can be reached by genuinely different derivations — direct arithmetic, per-jar reasoning, working backwards from a guess and checking it. Dozens of routes, one destination. The vote banks exactly this convergence: each correct path casts a ballot for the same value.
A wrong path usually fails for its own local reason — a misread jar count, a hallucinated premise, a slipped division. Each derailment lands somewhere different, so no single wrong answer ever collects a majority. If 70% of paths are right and the 30% wrong mass splits three ways, the truth wins with room to spare.
Ask 40 engineers to solve the same problem independently: the consensus is usually right even when individuals err — not because engineers are infallible, but because correct solutions coincide while mistakes are personal. Self-consistency gives a language model the same privilege: temperature buys independence between paths, and the vote collects the wisdom. Bonus: the vote margin is a free confidence estimate — a narrow vote means an uncertain model, before you ever see the label.
Start at k = 1 — a single sample can be flat wrong. Drag up and watch the correct answer (9) take a stable majority while the wrong answers starve.
Illustrative — a fixed GSM8K-style answer distribution: ~70% of paths land on 9, the rest scatter across 12, 3 and 59. Watch the vote share of the correct answer stabilize as k grows.
The paper evaluates self-consistency on arithmetic, commonsense and symbolic reasoning. The pattern repeats everywhere: large, consistent gains with zero training and zero changes to the model or prompt.
| Benchmark | Reasoning Type | CoT (greedy) | CoT + Self-Consistency |
|---|---|---|---|
| GSM8K · PaLM-540B | grade-school math | 56.5 | 74.4 (+17.9) |
| AQuA | math word problems (multiple choice) | — | +8 to +18 pts (typical) |
| SVAMP | elementary math word problems | — | +8 to +18 pts (typical) |
| StrategyQA | implicit multi-step commonsense | — | +8 to +18 pts (typical) |
Only the GSM8K headline (PaLM-540B, k = 40, same prompt, same weights) is quoted exactly. The paper reports consistent absolute gains in roughly the +8 to +18 point range across its arithmetic, commonsense and symbolic benchmarks — exact deltas vary by model, prompt and k, so they are summarized here as a range. On several benchmarks, self-consistency also beat systems that had been fine-tuned for the task.
With self-consistency, PaLM-540B outperformed or matched much larger models — and on several benchmarks, models that had been fine-tuned for the task — without a single gradient step. The reasoning ability was already in the weights; greedy decoding was simply throwing most of it away.
The gains do not hinge on one lucky sampler: temperature, top-k and nucleus sampling all worked. What matters is diversity between paths — enough for errors to decorrelate, not so much that the chains turn into nonsense.
Self-consistency buys accuracy with test-time compute. That trade is the paper's most modern idea: accuracy as a function of the inference budget, not just of parameters.
Drag the number of sampled paths: cost grows linearly, accuracy saturates. Then flip the toggle to see how little the 40th path actually buys.
Illustrative saturating curve, anchored to the paper's GSM8K headline (56.5 → 74.4 at k = 40). The linear cost line is exact.
Self-consistency was the first clean demonstration that inference compute is a scaling axis of its own. The reasoning models of 2024 and beyond are built on that insight.
Five questions on sampling, voting, and the price you pay.
Everything worth remembering about self-consistency.