History Problem Core Idea Afterlife Results Impact Quiz Takeaways
Interactive Paper Explainer

Two Words:
"Step by Step"
Zero-Shot Reasoners

No examples, no fine-tuning — one appended sentence. 'Let's think step by step' turns zero-shot answers into zero-shot reasoning, with large accuracy jumps on arithmetic, symbolic, and commonsense tasks.

Start Learning Read the Paper ↗
1
Prompt sentence
10+
Tasks swept
Up to ~36pt
Single-task jump (paper figure)
2022
Kojima et al.
History

Examples Were Everything

The 2022 assumption zero-shot prompting quietly broke.

2020-21
Few-shot is the GPT-3 story
In-context exemplars define capability: 'few-shot learners' (entry #6). Zero-shot is the weak sibling.
Jan 2022
CoT needs examples
Chain-of-thought prompting (entry #48) elicits reasoning — but only with hand-written worked examples in the prompt.
May 2022
🚀 Zero-shot-CoT
Kojima et al.: the same model, no exemplars — append "Let's think step by step" and extract the answer afterward. MultiArith +36pp, GSM8K +17pp class gains (paper figures).
2022-23
The prompting wave
Self-consistency (entry #49) stacks on it; the trick becomes folklore, then default practice, then built-in (chat models think before answering by training).
2024-25
Reasoning as a product mode
Think-before-answer graduates from a prompt suffix to a trained behavior — o1/R1-style models institutionalize what two words began.
Reasoning Was There All Along

The paper's quiet thesis: models had latent multi-step ability that few-shot exemplars were merely eliciting, not creating. If elicitation needed examples, it was a formatting problem — and formatting problems have cheap fixes. The fix here is almost embarrassing: a single imperative sentence triggers the step-by-step register; a second prompt ("…so the answer is") then harvests it. The pattern generalizes: zero-shot capability is often a prompt-shape discovery away.

Chapter 01

Examples or Nothing

The 2022 constraint zero-shot CoT removed.

📚
The Exemplar Tax
  • CoT prompting (entry #48) requires hand-crafted worked examples per task family
  • Example choice dominates results — a new hyperparameter surface, sensitive and manual
  • Zero-shot prompting collapses on multi-step tasks: models blurt answers before thinking
  • Nobody had isolated the question: do models reason zero-shot at ALL?
✨
The Two-Word Answer
  • Prompt: {question} + "Let's think step by step" — no exemplars, one template for all tasks
  • The model generates a rationale, then a second pass ("Therefore, the answer is") extracts it
  • Large zero-shot gains: MultiArith +36pp, GSM8K +17pp-class (paper figures) — same weights
  • Result: CoT-style reasoning without example engineering — elicitation, not teaching
Analogy — The Counting Question

Ask a child "17 minus 8?" out of nowhere and they blurt a guess. Ask the same child "17 minus 8 — take your time, work it out" and they murmur "8 plus what makes 17… 9." The ability was always there; the register you invite determines whether it shows up. The paper found the magic invitation for LLMs — and it was two words long.

Chapter 02

The Two-Step Elicitation

The full recipe — trigger, then harvest.

1️⃣ Trigger
"Q: {question}\nA: Let's think step by step." — the model continues with a free-form rationale, working the problem aloud.
2️⃣ Harvest
Append "Therefore, the answer is" (or A: …) — a second short completion that extracts the final answer from the rationale above.
🧪 Task sweep
Arithmetic (MultiArith, GSM8K, AQUA-RAT, SVAMP), symbolic (last-letter, coin flips), commonsense, and other benchmark suites — one template for everything.
📊 Scale interplay
Largest models gain the most — consistent with reasoning as an emergent-ability phenomenon (weak models do not suddenly think because asked nicely).
Reported effect sizes
  • MultiArith: ~+36 percentage points over standard zero-shot (paper figure)
  • GSM8K: ~+17pp-class jump (InstructGPT-class models)
  • Symbolic tasks (last letter, coin): large gains where the task is pure procedure
  • Consistent across model families — the sentence works broadly, not on one model
What the paper checked (and you should too)
  • Control prompts tested: other imperative sentences ("solve this") — weaker; the step-by-step register is special
  • Sensitivity: prompt phrasing matters at the margin — prompt engineering remains a real variable
  • Answer extraction is its own failure mode — the harvest step needs care on free-form outputs
Interactive Demo — The Two-Step Recipe, Live

Watch one question go from blurting to reasoning — trigger, rationale, harvest.

Chapter 03

From Suffix to Training

The trick's afterlife: folklore, default, then weight-level behavior.

The Institutionalization Path

The sentence went through three careers: (1) folklore — every prompt-engineering guide taught it; (2) default — instruction-tuned models absorbed so much step-by-step training data that reasoning became their natural register (the suffix became redundant); (3) trained behavior — o1/R1-class models (entries #47, #58) generate long reasoning chains by construction, spending tokens on thinking as a policy. The conceptual arc the paper opened: capability and elicitation are different variables — and elicitation can be moved both by prompts (cheap, brittle) and by training (durable, expensive).

Interactive Demo — What the Prompt Actually Does

Tab through hypotheses for why two words move performance — the paper's own analysis lens.

Chapter 05

One Sentence, Ten Benchmarks

The sweep's pattern: big jumps on procedure-heavy tasks, model-scale dependent.

MULTIARITH
+36pp
zero-shot baseline → zero-shot CoT (paper figure)
GSM8K
+17pp-class
same weights, one template
SYMBOLIC TASKS
large ↑
last-letter concatenation, coin-flip tracking
TEMPLATE
1
single prompt for every task — no per-task examples
Interactive Demo — From Suffix to Trained Behavior

The two-word trick's career arc. Press reveal to trace it from folklore to o1-style reasoning models.

Legacy

Legacy — The Cheapest Capability Unlock

Zero-shot CoT is the highest ROI sentence in LLM history.

✂️ The examples axe
It deleted CoT's example-engineering cost: reasoning on a new task with one shared template — the accessibility moment for chain-of-thought.
🧪 The elicitation frame
Capability vs. elicitation became a standard analysis lens — later applied to tool use, self-consistency (entry #49), and the 'emergent abilities' debate.
🏗 Prompting's professional era
Zero-shot CoT legitimized systematic prompt research (controls, phrasing sensitivity, answer extraction) — the paper's methodology, not just its sentence, was copied.
🚀 The reasoning-model prequel
The eventual product-mode answer (o1/R1-class models that think by construction) is this paper's question answered at the weight level — elicitation institutionalized.
⚠️ What it did NOT solve
Small models gain little (emergence's double edge); phrasing sensitivity persists; answer extraction from free-form rationales is its own failure surface; and hallucinated steps look identical to valid ones.
🛤 Read next
Test Yourself

Quick Quiz

Check your understanding of the key concepts from Zero-Shot Reasoners.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ "Let's think step by step" + answer-harvest prompt = zero-shot chain-of-thought, no exemplars.
✅ Reported jumps: ~+36pp MultiArith, ~+17pp-class GSM8K — one template, all tasks, unchanged weights.
✅ Effect concentrates in large models — reasoning as emergent, elicited ability.
✅ Control prompts show the register (not just any imperative) matters — elicitation is real but fragile.
✅ The suffix's career: folklore → default → built-in (chat models) → trained behavior (o1/R1).
✅ Read it as the proof that prompting was capability testing in disguise.