A prompting technique so simple it sounds like a trick — show the model worked examples whose answers include intermediate reasoning steps, and multi-step reasoning performance jumps. But only once models get big enough. No training, no new architecture: just a better prompt.
Chain-of-thought didn't appear out of nowhere — it is the missing piece of a five-year story about how much you can get a language model to do purely through its prompt.
A language model generates one token at a time — so tokens are working memory. Standard prompting spends no tokens on thinking before the answer. Chain-of-thought spends them on the intermediate results a multi-step problem needs, and a model that gets every step right arrives at the right answer.
Standard prompting asks the model to answer immediately — question in, answer out. Multi-step problems need intermediate results that a single forward pass can't produce, and fine-tuning on task data is expensive.
Try multiplying 47 × 83 entirely in your head, committing to the first number that feels right. That is standard prompting: fluent, confident, and often wrong. Now do it on paper, one line at a time — 47 × 3 = 141, 47 × 8 = 376, carry, add — each step is easy, and a mistake would be visible on the page. A math teacher demands working for exactly the reasons CoT works: the scratch space is the thinking. The paper's move is simply that the "paper" is more tokens in the model's own output — and you lend it the paper by putting reasoning inside the few-shot exemplars.
Keep GPT-3's few-shot format. Change exactly one thing: the exemplars' answers now contain the reasoning chain before the final answer. The model mimics the format — and the reasoning comes along with it.
The exemplar teaches the output format — and nothing about how to get there.
The exemplar teaches the output format and the shape of the reasoning.
CoT is not a free lunch for every model. Gains only appear at roughly ~100B parameters — and below ~10B, chain-of-thought often makes performance worse: small models produce fluent but illogical chains of thought.
Arithmetic, commonsense, symbolic. Chain-of-thought prompting with PaLM-540B beats standard prompting across the board — with one honest caveat: the gains live in the largest models.
| Benchmark | Standard | With CoT | What it is |
|---|---|---|---|
| GSM8K | 17.9 | 56.9 (+39.0) | multi-step math word problems |
| SVAMP | 69.4 | 79.0 (+9.6) | simpler word problems, varying structure |
| AQuA | 25.2 | 35.8 (+10.6) | multiple-choice algebra |
| MAWPS | 79.2 | 93.3 (+14.2) | word problems, varied templates |
All values from the paper's Table 2. PaLM-540B + CoT on GSM8K reached state-of-the-art territory, surpassing even fine-tuned GPT-3 with a verifier (~55% prior best) — from prompting alone. The jump is biggest exactly where problems are hardest: multi-step GSM8K.
On commonsense benchmarks (CSQA, StrategyQA, and BIG-bench evaluation sets like sports understanding), CoT helps most at the largest scale. PaLM-540B with CoT hit 75.6% on StrategyQA — outperforming the prior fine-tuned state of the art (69.4%) — and beat an unaided sports enthusiast on sports understanding (95.4% vs 84%). Being honest: the gain on CSQA was minimal, and for GPT-3 some commonsense tasks did not improve at all.
On toy symbolic tasks — coin flips (state tracking) and last-letter concatenation — PaLM-540B with CoT reaches near-100% solve rates, including on out-of-distribution lengths longer than any exemplar. Standard prompting fails outright on last-letter concatenation (near 0%). One honest wrinkle: list reversal defeated two of the paper's co-authors' prompts — only a carefully engineered chain from a third solved it. Prompts are not all equal.
Every big number above belongs to a 540B-parameter model. The same chain-of-thought prompt applied to an 8B model does nothing or actively hurts — the paper found CoT "actually hurts performance for most models smaller than 10B parameters". The ability to follow and produce a correct reasoning chain emerges with scale. If your model is small, CoT is not your lever (Chapter 03); if it is huge, it is nearly free performance.
Maybe the gains come merely from "having computation in the prompt"? The paper strips the chain down to bare equations, pure computation, and reasoning-after-the-fact — every stripped-down version fails. Natural-language reasoning itself is the active ingredient.
The paper that made "reasoning" an LLM benchmark axis. Everything that followed — sampling chains, distilling them, training on them with RL — builds on this one change to the prompt.
Check your understanding of the key concepts from the Chain-of-Thought prompting paper.
Everything you need to remember about this paper.