Outcome supervision checks the final answer; process supervision checks every reasoning step. On MATH, the difference is decisive — and the paper releases 800K step-level labels to prove it.
The question supervision design had to answer: when a model gets a math problem wrong, what exactly do you train against?
Outcome-based training has a subtle poison: a chain with three logic errors that stumbles into the right final answer gets reinforced — every mistake in it included. Process supervision breaks that coupling: each step is labeled correct / incorrect / neutral as it is generated, so reinforcement reaches exactly the steps that earned it. The paper's finding: on MATH, this trains significantly more reliable models than outcome supervision — 78% of a representative test subset solved by the process-supervised model.
The false-positive problem at the heart of outcome supervision.
Outcome supervision is a teacher who only grades the answer key — 'final answer 42: correct, A+'. The student who botched three steps and lucked into 42 is rewarded identically to the flawless proof. Process supervision is the teacher who reads every line with a red pen: errors get marked WHERE they happen, partial credit is honest, and lucky endings can't launder broken logic.
Same base model, same budget framing — three supervision designs.
The labeling protocol — and why active learning was the multiplier.
The neutral label is a quiet design win: forcing every step into binary correctness would have manufactured noise — ambiguous steps exist, and giving them a home keeps the signal clean.
The score and the transparency arrived together.
| Dimension | Outcome RM (ORM) | Process RM (PRM) |
|---|---|---|
| Signals per solution | 1 (final answer) | every step (correct/incorrect/neutral) |
| False positives | common — lucky chains reinforced | marked at the step where they occur |
| MATH subset solved | lower | 78% |
| Label cost | cheaper per example | higher per example, ~3× more efficient with active learning |
| Failure diagnosis | opaque | localizable — which step failed |
Summary of the paper's comparison; exact per-benchmark numbers are in its tables and figures.
Step-level verification became the backbone of the reasoning era.
Check your understanding of the key concepts from Let's Verify Step by Step.
Everything you need to remember about this paper.