Can a model learn to reason without a single human-written reasoning example? DeepSeek-R1 said yes — pure reinforcement learning produced emergent "aha moments", and the open weights went on to match OpenAI's o1.
Reasoning spent three years as a prompting trick, then a closed product. With R1 it became an open recipe anyone can run.
Reasoning doesn't need to be demonstrated — only rewarded. If a task's answers can be checked, reinforcement learning can find the thinking on its own. That was DeepSeek's bet against "you must fine-tune on human reasoning data first" — and R1-Zero's emergent "aha moments" proved it.
By late 2024, reasoning models existed. All of them were closed — and the field's default recipe assumed a step nobody had questioned: imitate human reasoning first.
You can train a chess student two ways. Hand them 10,000 annotated grandmaster games to memorize (SFT on human CoT) — they learn the openings, but their ceiling is the games they copied. Or let them play thousands of matches and tell them just one bit of information each time: won or lost (RL with a verifiable reward). DeepSeek bet on the second — and the student started inventing habits nobody taught, like double-checking a move before committing. That habit is the aha moment.
One strong open base model, and one RL algorithm that deleted the most expensive part of PPO.
A Mixture-of-Experts LLM: 671B total parameters, but only ~37B active per token — each token routes to a small subset of experts. Already strong on math, code, and general knowledge, it was a good enough starting point to attempt RL directly, with no reasoning fine-tune in between.
PPO — the workhorse of RLHF — estimates advantages with a critic (value) network typically as large as the policy itself: double the memory, double the compute, more things to go wrong. GRPO (Group Relative Policy Optimization, introduced with DeepSeekMath) deletes the critic: for each prompt, sample a group of G responses, and let the group's own average play the baseline.
The pure experiment: take V3, run GRPO with two rule-based rewards, and add zero — zero! — supervised reasoning data. What appeared was never prompted, never taught.
Is the final answer correct? On math, answers can be checked exactly with rules — no learned reward model, nothing to reward-hack. Just: right or wrong.
Put the reasoning inside <think> … </think> tags, with the answer after. That is the only structural demand — how to think is completely free.
Emergent reasoning came with emergent mess: readability problems (rambling, repetitive reasoning) and language mixing (English and Chinese interleaved mid-thought). The reasoning was there; the packaging wasn't. Fixing that without losing the magic is Chapter 04's job.
R1-Zero proved reasoning can emerge. R1 makes it shippable: readable, language-consistent, and aligned — without forgetting how to reason.
Because pure-RL reasoning is a mess to read and mixes languages mid-thought. The full recipe interleaves small amounts of supervised data with rounds of RL — each stage fixing what the previous one broke, keeping the emergent reasoning while restoring usability. Four stages:
On the benchmarks that defined the reasoning era, R1 lands at o1-1217's level — with weights anyone can download.
| Benchmark | DeepSeek-R1 | OpenAI-o1-1217 |
|---|---|---|
| AIME 2024 (pass@1) | 79.8 | 79.2 |
| MATH-500 | 97.3 | 96.4 |
| MMLU | 90.8 | ~91.8 |
| Codeforces (percentile) | 96.3 | — (not directly comparable) |
AIME and MATH-500 are the paper's exact numbers; o1's MMLU is approximate (~), and its Codeforces rating isn't directly comparable. Beyond benchmarks, R1 stays strong on non-reasoning tasks too — writing, factual QA, summarization.
The team fine-tuned six open small models — Qwen and Llama checkpoints from 1.5B to 70B — on ~800k R1 samples. No RL for the small models: just imitation of R1's verified reasoning traces.
The team also trained the same small models with RL directly. Result: the distilled models beat their RL-trained counterparts across benchmarks. For small models, reasoning transfers better by distilling it from a big RL model than by re-learning it from scratch — the big model's expensive exploration gets exported for free.
One January release reframed what frontier AI costs to build — and who gets to build it.
Five questions on the paper that taught models to think for themselves.
Everything you need to remember about DeepSeek-R1.