History Problem Core Idea Dataset Results Impact Quiz Takeaways
Interactive Paper Explainer

Grade the Steps,
Not the Answer
Let's Verify Step by Step

Outcome supervision checks the final answer; process supervision checks every reasoning step. On MATH, the difference is decisive — and the paper releases 800K step-level labels to prove it.

Start Learning Read the Paper ↗
78%
MATH subset solved
800K
Step-level labels (PRM800K)
3×
Better active learning (paper claim)
2023
Lightman et al.
History

Where Did the Error Begin?

The question supervision design had to answer: when a model gets a math problem wrong, what exactly do you train against?

2021-22
Chain-of-thought arrives
CoT prompting (entry #48) makes reasoning visible — but the training signal is still all-or-nothing on the final answer.
2022
The ORM era
Outcome reward models + rejection sampling: sample many solutions, keep the ones that land the right answer. Errors mid-chain survive whenever the ending is lucky.
2022 · Usher/Snigdha
Early process work
Usher et al. and peers experiment with step-level verifiers — promising, small-scale, under-validated.
May 2023
🚀 Let's Verify Step by Step
OpenAI's careful comparison at scale: 800K human-verified step labels, process vs outcome vs no supervision, same base model. Process wins decisively; PRM800K released.
2024-25
Process supervision everywhere
Math RL (DeepSeekMath/R1 lineage), test-time search (entry #58), and verifiable-reward pipelines inherit the step-verifier pattern.
The False-Positive Machine

Outcome-based training has a subtle poison: a chain with three logic errors that stumbles into the right final answer gets reinforced — every mistake in it included. Process supervision breaks that coupling: each step is labeled correct / incorrect / neutral as it is generated, so reinforcement reaches exactly the steps that earned it. The paper's finding: on MATH, this trains significantly more reliable models than outcome supervision — 78% of a representative test subset solved by the process-supervised model.

Chapter 01

Right Answer, Wrong Reasoning

The false-positive problem at the heart of outcome supervision.

🎯
Outcome Supervision's Blind Spot
  • Final-answer checking cannot see WHERE reasoning went wrong — only that it did (or got lucky)
  • Lucky-wrong chains get reinforced: false-positive training signal baked into the weights
  • Rejection sampling amplifies the issue: any sampled path ending in the right answer becomes supervision
  • On multi-step problems, reliability — not peak accuracy — collapses first
🪜
The Process Answer
  • PRM: a reward model that scores EACH step (correct / incorrect / neutral) as the solution unfolds
  • PRM800K: ~800K step-level human labels on GPT-4-generated MATH solutions, released publicly
  • Active learning: label the steps the current PRM is most unsure about — 3× the data efficiency of naive sampling (paper's claim)
  • Result: 78% of a representative MATH test subset solved — with the reliability profile outcome models can't match
Analogy — The Math Teacher's Red Pen

Outcome supervision is a teacher who only grades the answer key — 'final answer 42: correct, A+'. The student who botched three steps and lucked into 42 is rewarded identically to the flawless proof. Process supervision is the teacher who reads every line with a red pen: errors get marked WHERE they happen, partial credit is honest, and lucky endings can't launder broken logic.

Chapter 02

The Three-Way Comparison

Same base model, same budget framing — three supervision designs.

The competitors
  • ORM (outcome RM): score the full solution from its final answer correctness
  • PRM (process RM): score each step — correct / incorrect / neutral — trained on PRM800K labels
  • No supervision: the finetuned-on-solutions baseline
  • Best-of-N sampling: generate N solutions, pick by RM score — the decoupling test of verifier quality
The result pattern
  • Process supervision significantly outperforms outcome supervision for training models to solve MATH problems
  • The process-supervised model solves 78% of a representative 500-problem MATH test subset
  • PRM wins even at matched label budget — the signal quality, not volume, dominates
  • Active learning (uncertainty-directed labeling) improves data efficiency ~3× over naive random labeling
Interactive Demo — One Solution, Two Graders

Tab through a worked solution with a mid-chain error — watch ORM bless it and PRM catch it.

Chapter 03

How PRM800K Was Built

The labeling protocol — and why active learning was the multiplier.

Dataset Construction

The neutral label is a quiet design win: forcing every step into binary correctness would have manufactured noise — ambiguous steps exist, and giving them a home keeps the signal clean.

Interactive Demo — The Active Learning Loop

How 800K labels punched above their weight — label where the verifier is unsure, retrain, repeat.

Chapter 05

78% — and You Can See Why

The score and the transparency arrived together.

MATH SUBSET
78%
representative 500-problem test set, process-supervised
vs OUTCOME SUP.
significant
process wins at matched budget
DATA EFFICIENCY
~3×
from active learning on step labels
RELEASE
PRM800K
800K step-level labels, public
Interactive Demo — Why 'Neutral' Exists

Binary step labels would look cleaner. Press reveal for why the third class keeps the signal honest.

DimensionOutcome RM (ORM)Process RM (PRM)
Signals per solution1 (final answer)every step (correct/incorrect/neutral)
False positivescommon — lucky chains reinforcedmarked at the step where they occur
MATH subset solvedlower78%
Label costcheaper per examplehigher per example, ~3× more efficient with active learning
Failure diagnosisopaquelocalizable — which step failed

Summary of the paper's comparison; exact per-benchmark numbers are in its tables and figures.

Legacy

Legacy — The Verifier Layer

Step-level verification became the backbone of the reasoning era.

🪜 Math RL's foundation
Verifiable per-step rewards feed DeepSeekMath/R1-style RL (entries #46, #47) and test-time search (entry #58) — the reasoning-model stack is built on this comparison's verdict.
📦 PRM800K as infrastructure
The released 800K step labels became a standard resource for verifier research and training — a community contribution that outlived the specific model.
🔍 Honest reliability
The paper's framing — 'reliable' as the target, not 'accurate' — reoriented how the field scores reasoning: false-positive accounting entered the conversation.
⚗️ Methodological template
Matched-budget three-way comparison + released data + active-learning ablation: the design standard for supervision research since.
⚠️ What it did NOT solve
Label cost remains high (humans grading steps is slow); PRMs can be fooled by plausible-looking steps; and the MATH-domain focus left open how far process labels transfer to open-ended domains.
🛤 Read next
Test Yourself

Quick Quiz

Check your understanding of the key concepts from Let's Verify Step by Step.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Process supervision grades every step; outcome supervision grades only the ending.
✅ False positives are outcome training's poison: lucky-broken chains get reinforced whole.
✅ The process-supervised model solves 78% of a representative MATH test subset.
✅ PRM800K: ~800K released step-level labels — verifier research's standard resource.
✅ Active learning (uncertainty-directed labeling) makes step supervision ~3× more efficient.
✅ Read it as the reliability turn: the reasoning era's foundation is per-step honesty.