History Problem Core Idea Emergence Results Ablations Impact Quiz
Interactive Paper Explainer

Thinking Step by Step
Chain-of-Thought

A prompting technique so simple it sounds like a trick — show the model worked examples whose answers include intermediate reasoning steps, and multi-step reasoning performance jumps. But only once models get big enough. No training, no new architecture: just a better prompt.

Start Learning Read the Paper ↗
+39
GSM8K points · PaLM-540B (17.9 → 56.9)
540B
Backbone model · PaLM
3
Reasoning domains · math, commonsense, symbolic
2022
Year Published
History

From Prompting to Thinking

Chain-of-thought didn't appear out of nowhere — it is the missing piece of a five-year story about how much you can get a language model to do purely through its prompt.

2017 – 2019
Prompting, but fine-tuned
The BERT/GPT era adapts models to each task by fine-tuning. Prompts exist, but only inside an expensive, task-specific training pipeline.
2020
GPT-3: the prompt is the context
Few-shot learning with no gradient updates — a handful of Q → A examples in the context window is enough. But the exemplars show only final answers.
2021
Scratchpads (Nye et al.)
Letting models write out intermediate computation steps during training-time tasks dramatically helps — early evidence that intermediate steps matter.
2022 · Jan
🚀 Chain-of-thought prompting (Wei et al.)
Keep the few-shot format; add reasoning steps to the exemplars' answers. Gains emerge at scale — PaLM-540B solves 56.9% of GSM8K from prompting alone, beating fine-tuned GPT-3 with a verifier.
2022 · May
Self-consistency & zero-shot CoT
Wang et al. sample many chains and majority-vote the answers; Kojima et al. show a single sentence — "Let's think step by step" — works with no exemplars at all.
2024 →
o1, DeepSeek-R1: reasoning as the product
Models trained with reinforcement learning to generate long chains of thought before answering — thinking, scaled into a product.
Key Insight

A language model generates one token at a time — so tokens are working memory. Standard prompting spends no tokens on thinking before the answer. Chain-of-thought spends them on the intermediate results a multi-step problem needs, and a model that gets every step right arrives at the right answer.

SAME PROBLEM, TWO TOKEN BUDGETS
A: 59 ← one guess, no room to compute
A: 23 − 20 = 3
   3 + 6 = 9 ← steps as working memory
   So the answer is 9.
The cafeteria problem: 23 apples, 20 used, 6 bought → 9.
Chapter 01

The Problem with Instant Answers

Standard prompting asks the model to answer immediately — question in, answer out. Multi-step problems need intermediate results that a single forward pass can't produce, and fine-tuning on task data is expensive.

🗣️
Answer First, Think Never
  • Standard few-shot exemplars jump straight from Q to the final answer
  • No working space — the answer must arrive in one forward pass
  • Errors compound silently: a wrong intermediate value is never surfaced
  • Fine-tuning on task data is expensive and must be redone per task
  • Small models fail multi-step problems either way
📝
CoT's Solution: Room to Reason
  • Intermediate steps act as working memory before the answer
  • Few-shot only — eight exemplars, zero training, zero gradient updates
  • Errors become visible and debuggable inside the chain
  • Unlocks at scale: same prompt, more parameters, more reasoning
  • One technique covers arithmetic, commonsense, and symbolic tasks
Analogy — Show Your Work

Try multiplying 47 × 83 entirely in your head, committing to the first number that feels right. That is standard prompting: fluent, confident, and often wrong. Now do it on paper, one line at a time — 47 × 3 = 141, 47 × 8 = 376, carry, add — each step is easy, and a mistake would be visible on the page. A math teacher demands working for exactly the reasons CoT works: the scratch space is the thinking. The paper's move is simply that the "paper" is more tokens in the model's own output — and you lend it the paper by putting reasoning inside the few-shot exemplars.

Chapter 02

The Core Idea — Reasoning in the Exemplars

Keep GPT-3's few-shot format. Change exactly one thing: the exemplars' answers now contain the reasoning chain before the final answer. The model mimics the format — and the reasoning comes along with it.

Exemplar: Q: {problem}
A: {step · step · step · …} → "So the answer is X"
Q:
The question
Kept verbatim — nothing about the input changes.
A: + steps
Reasoning chain
Natural-language sentences with the arithmetic inline, line by line, before any answer.
"The answer is"
Fixed closer
A consistent final-line format the model learns to emit — easy to parse, easy to score.
8
Exemplars
One set of eight hand-written Q-and-reasoned-A pairs reused across benchmarks (four for AQuA).
Standard Exemplar
Q: Olivia has $23. She bought five bagels for $3 each. How much money does she have left? A: 8.

The exemplar teaches the output format — and nothing about how to get there.

Chain-of-Thought Exemplar (from the paper's prompt)
Q: Olivia has $23. She bought five bagels for $3 each. How much money does she have left? A: Olivia had 23 dollars. 5 bagels for 3 dollars each will be 5 x 3 = 15 dollars. So she has 23 - 15 dollars left. 23 - 15 is 8. The answer is 8.

The exemplar teaches the output format and the shape of the reasoning.

Interactive Demo — Standard vs CoT: Watch the Answer Change

A GSM8K-style problem and two prompts. Press a button to see what each produces. (Outputs are precomputed in the spirit of the paper's examples — the point is how the answer changes, not a live model.)

Q: The cafeteria had 23 apples. If they used 20 to make lunch and bought 6 more, how many apples do they have?
Chapter 03

Emergence — the Scale Gate

CoT is not a free lunch for every model. Gains only appear at roughly ~100B parameters — and below ~10B, chain-of-thought often makes performance worse: small models produce fluent but illogical chains of thought.

BELOW ~10B
worse
fluent but wrong chains — CoT hurts
~100B · GPT-3 175B, LaMDA 137B
modest
gains appear — CoT starts to help
540B · PaLM
+39
GSM8K points over standard prompting
Interactive Demo — The Scale Gate on GSM8K

Toggle the prompting mode and watch the bars. All four values are the paper's reported GSM8K solve rates (Table 2 / appendix — GPT-3 = text-davinci-002). Note how the smallest model's bar shrinks with CoT.

GSM8K solve rate (%), bars scaled to 60%. 17.9 → 56.9 for PaLM-540B; 6.5 → 14.3 for LaMDA-137B (modest); 15.6 → 46.9 for GPT-3-175B; 3.2 → 1.6 for LaMDA-8B — CoT actively hurts below ~10B parameters.
Chapter 04

Results — Three Reasoning Domains

Arithmetic, commonsense, symbolic. Chain-of-thought prompting with PaLM-540B beats standard prompting across the board — with one honest caveat: the gains live in the largest models.

Arithmetic · PaLM-540B, Standard vs Chain-of-Thought (solve rate %)
BenchmarkStandardWith CoTWhat it is
GSM8K17.956.9 (+39.0)multi-step math word problems
SVAMP69.479.0 (+9.6)simpler word problems, varying structure
AQuA25.235.8 (+10.6)multiple-choice algebra
MAWPS79.293.3 (+14.2)word problems, varied templates

All values from the paper's Table 2. PaLM-540B + CoT on GSM8K reached state-of-the-art territory, surpassing even fine-tuned GPT-3 with a verifier (~55% prior best) — from prompting alone. The jump is biggest exactly where problems are hardest: multi-step GSM8K.

Commonsense Reasoning

On commonsense benchmarks (CSQA, StrategyQA, and BIG-bench evaluation sets like sports understanding), CoT helps most at the largest scale. PaLM-540B with CoT hit 75.6% on StrategyQA — outperforming the prior fine-tuned state of the art (69.4%) — and beat an unaided sports enthusiast on sports understanding (95.4% vs 84%). Being honest: the gain on CSQA was minimal, and for GPT-3 some commonsense tasks did not improve at all.

Symbolic Reasoning

On toy symbolic tasks — coin flips (state tracking) and last-letter concatenation — PaLM-540B with CoT reaches near-100% solve rates, including on out-of-distribution lengths longer than any exemplar. Standard prompting fails outright on last-letter concatenation (near 0%). One honest wrinkle: list reversal defeated two of the paper's co-authors' prompts — only a carefully engineered chain from a third solved it. Prompts are not all equal.

GSM8K · PaLM-540B + CoT
56.9%
solve rate — vs 17.9 standard, ~55 prior fine-tuned best
SVAMP · PaLM-540B + CoT
79.0%
solve rate — easier task, smaller CoT jump
DOMAINS
3
arithmetic · commonsense · symbolic
TRAINING RUNS NEEDED
0
off-the-shelf models, prompting only
The Honest Caveat

Every big number above belongs to a 540B-parameter model. The same chain-of-thought prompt applied to an 8B model does nothing or actively hurts — the paper found CoT "actually hurts performance for most models smaller than 10B parameters". The ability to follow and produce a correct reasoning chain emerges with scale. If your model is small, CoT is not your lever (Chapter 03); if it is huge, it is nearly free performance.

Chapter 05

Ablations — Why the Words Matter

Maybe the gains come merely from "having computation in the prompt"? The paper strips the chain down to bare equations, pure computation, and reasoning-after-the-fact — every stripped-down version fails. Natural-language reasoning itself is the active ingredient.

🧮 Equation only
Prompt the model to emit just a mathematical equation, then the answer. On LaMDA-137B GSM8K this scores 5.4 vs 14.3 for full CoT — the semantics of GSM8K questions are too hard to translate into the right equation without natural-language steps. A real failure from the paper: the model wrote "(4 + 20 × 0.25) = 6" for a ping-pong points problem — right shape, wrong reasoning, wrong answer.
⠿ Variable compute only
Give the model extra intermediate tokens with no content at all — a sequence of dots (…) as long as the equation would be. It performs about the same as standard prompting (6.4 vs 14.3 CoT). Variable computation alone is not the secret: the words carry the gain.
↩️ Reasoning after answer
Keep the full chain but put the answer before it. Score: 6.1 vs 14.3. The reasoning has to precede the answer to be useful — explaining after committing changes nothing.
🔀 Prompt order matters
Like all few-shot methods, CoT is sensitive to its prompt. Variance across exemplar orders was small in most cases, but large for coin flips; performance also varied by annotator (99.6% vs 71.4%). Prompt engineering still matters.
Interactive Demo — Strip the Chain, Lose the Gain

Click a prompting variant to see the exemplar format the model would copy, and the resulting PaLM-540B GSM8K solve rate. The 56.9 is exact; ~ values are illustrative of the direction reported in the paper's ablation study (Figure 5 / Tables 6–7, where ablations land at or below standard prompting).

Legacy

Impact — Reasoning Becomes an Axis

The paper that made "reasoning" an LLM benchmark axis. Everything that followed — sampling chains, distilling them, training on them with RL — builds on this one change to the prompt.

🗳️ Self-consistency (2022)
Generate many chains, majority-vote the final answers — stacks a second large gain on top of CoT. Self-Consistency Guide → — the natural next paper.
✨ Zero-shot CoT (2022)
Kojima et al.: one magic sentence — "Let's think step by step" — unlocks reasoning with no exemplars at all.
🎓 CoT distillation
Fine-tune smaller models on model-generated chains of thought to distill the ability down below the emergence scale.
🤖 Reasoning models: o1, DeepSeek-R1
Modern systems trained with reinforcement learning to produce long chains of thought before answering — thinking as the product. DeepSeek-R1 Guide →
📊 The benchmark axis
Post-CoT, "can it reason?" became a headline question for every frontier model — GSM8K and friends became standard reporting.
🔬 Emergence research
CoT became the canonical example of an ability that appears suddenly with scale — fueling the whole emergence debate.
Test Yourself

Quick Quiz

Check your understanding of the key concepts from the Chain-of-Thought prompting paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ CoT = few-shot exemplars whose answers include intermediate reasoning steps, ending in "The answer is X."
✅ No training, no new architecture — the same frozen model, prompted differently with ~8 exemplars.
✅ Emergence: gains appear around ~100B parameters; below ~10B, CoT often hurts (fluent but wrong chains).
✅ PaLM-540B on GSM8K: 17.9 → 56.9 solve rate (+39 points) — beating fine-tuned GPT-3 with a verifier, from prompting alone.
✅ The words matter: equation-only, variable-compute-only, and reasoning-after-answer ablations all fail — natural-language reasoning is the active ingredient.
✅ Legacy: self-consistency, zero-shot CoT, distillation, and RL-trained reasoning models (o1, DeepSeek-R1) all build on this paper.