History Problem Base R1-Zero Pipeline Results Impact Quiz
Interactive Paper Explainer

Reasoning via Reinforcement
DeepSeek-R1

Can a model learn to reason without a single human-written reasoning example? DeepSeek-R1 said yes — pure reinforcement learning produced emergent "aha moments", and the open weights went on to match OpenAI's o1.

Start Learning Read the Paper ↗
671B
Total Params (37B Active)
79.8%
AIME 2024 pass@1
6
Distilled Models (1.5B→70B)
2025
Year Published
History

From Prompted Thinking to Trained Thinking

Reasoning spent three years as a prompting trick, then a closed product. With R1 it became an open recipe anyone can run.

2022 · Jan
Chain-of-Thought prompting (Wei et al.)
A few worked examples in the prompt make models reason step by step. Reasoning by imitation — of human-written traces. Guide →
2023
Self-consistency & the reasoning-eval era
Sample many chains and vote (self-consistency); GSM8K and MATH become the yardsticks. Reasoning is still something you prompt, not train.
2024
o1 (OpenAI)
Test-time reasoning arrives: think longer at inference, score dramatically higher on math and code. Proprietary recipe, closed weights.
2025 · Jan
🚀 DeepSeek-R1 (DeepSeek-AI)
RL alone elicits reasoning — the famous "aha moment" emerges untrained — and the MIT-licensed weights go on to match o1 on math benchmarks.
2025 →
The open reasoning wave
R1's distilled models (1.5B–70B) and a wave of open reasoning models carry test-time reasoning to every scale.
Key Insight

Reasoning doesn't need to be demonstrated — only rewarded. If a task's answers can be checked, reinforcement learning can find the thinking on its own. That was DeepSeek's bet against "you must fine-tune on human reasoning data first" — and R1-Zero's emergent "aha moments" proved it.

R1-ZERO'S ENTIRE SUPERVISION
answer correct? → reward ✓
reasoning inside <think>…</think>? → reward ✓
human-written reasoning → none. zero.
Two rule-based rewards. That is the whole supervision.
Chapter 01

The Problem — Reasoning Behind Locked Doors

By late 2024, reasoning models existed. All of them were closed — and the field's default recipe assumed a step nobody had questioned: imitate human reasoning first.

🔒
The Closed Era
  • OpenAI's o1 proved test-time reasoning works — but the training recipe stayed secret and the weights closed
  • The default alternative — SFT on human-written reasoning traces — is expensive to collect at scale
  • Human traces cap exploration: the model can only imitate reasoning it has seen
  • Imitation ≠ reasoning — copying a grandmaster's moves is not playing chess
🔓
DeepSeek's Bet
  • Question the assumption: skip human reasoning data — start RL straight from the base model
  • Reward only what's checkable: is the answer right, and is the reasoning wrapped in <think> tags?
  • Let the model discover its own reasoning strategies — including ones no human would write
  • Ship the weights, MIT-licensed, for everyone
The Chess-Coach Analogy

You can train a chess student two ways. Hand them 10,000 annotated grandmaster games to memorize (SFT on human CoT) — they learn the openings, but their ceiling is the games they copied. Or let them play thousands of matches and tell them just one bit of information each time: won or lost (RL with a verifiable reward). DeepSeek bet on the second — and the student started inventing habits nobody taught, like double-checking a move before committing. That habit is the aha moment.

Chapter 02

The Ingredients — V3 + GRPO

One strong open base model, and one RL algorithm that deleted the most expensive part of PPO.

The Foundation — DeepSeek-V3

A Mixture-of-Experts LLM: 671B total parameters, but only ~37B active per token — each token routes to a small subset of experts. Already strong on math, code, and general knowledge, it was a good enough starting point to attempt RL directly, with no reasoning fine-tune in between.

671B total · 37B active · ≈5.5% of weights fire per token
The Algorithm — GRPO

PPO — the workhorse of RLHF — estimates advantages with a critic (value) network typically as large as the policy itself: double the memory, double the compute, more things to go wrong. GRPO (Group Relative Policy Optimization, introduced with DeepSeekMath) deletes the critic: for each prompt, sample a group of G responses, and let the group's own average play the baseline.

no critic / value network cheaper than PPO
Ai = ( ri − mean(r) ) / std(r)
r₁ … r_G = rewards of the G responses sampled for one prompt
Ai
Relative Advantage
How much better response i is than its group. Positive → reinforce; negative → suppress.
ri
Verifiable Reward
Rule-based: answer correctness + format adherence. No learned reward model to game.
mean(r)
The Group Baseline
The average of the sampled rewards replaces the critic network — the group IS the value estimate.
std(r)
Normalization
Dividing by the group's spread keeps gradient scale stable across easy and hard prompts.
Interactive Demo — GRPO: The Group Is the Critic

One prompt, G = 5 sampled responses. Rewards are compared to the group's own mean — that dashed line is the baseline PPO would need an entire network to learn. Press Re-roll group to sample a new batch.

PROMPT · A jacket costs $40 after a 20% discount. What was the original price?
bar = reward (1 correct / 0 wrong) · - - - dashed line = group mean · A = group-relative advantage
Chapter 03

R1-Zero — Reasoning From Nothing

The pure experiment: take V3, run GRPO with two rule-based rewards, and add zero — zero! — supervised reasoning data. What appeared was never prompted, never taught.

Reward 1 — Accuracy 🎯

Is the final answer correct? On math, answers can be checked exactly with rules — no learned reward model, nothing to reward-hack. Just: right or wrong.

Reward 2 — Format 📋

Put the reasoning inside <think> … </think> tags, with the answer after. That is the only structural demand — how to think is completely free.

🔍 Self-Verification
The model starts writing "Wait, let me check that…" and re-checks its own arithmetic — unprompted.
🪞 Reflection
Mid-solution, it reconsiders its assumptions: "Hmm, let me reconsider whether this setup is right."
🧭 Alternative Paths
When one approach stalls, it backtracks and tries another strategy — behavior nobody demonstrated to it.
💡 The "Aha Moment"
The paper's highlight: mid-derivation, the model realizes it was wrong and visibly course-corrects (see the demo below).
Interactive Demo — The Aha-Moment Player

A stylized R1-Zero-style trace (shortened — real traces run for thousands of tokens). Press Next step and watch the model think… then catch itself. None of what follows was prompted or rewarded.

PROMPT · A train covers 240 km. The first 120 km at 60 km/h, the second 120 km at 40 km/h. What is the average speed for the whole trip?
The Catch — Why R1-Zero Didn't Ship

Emergent reasoning came with emergent mess: readability problems (rambling, repetitive reasoning) and language mixing (English and Chinese interleaved mid-thought). The reasoning was there; the packaging wasn't. Fixing that without losing the magic is Chapter 04's job.

Chapter 04

DeepSeek-R1 — The Full Recipe

R1-Zero proved reasoning can emerge. R1 makes it shippable: readable, language-consistent, and aligned — without forgetting how to reason.

Why not just ship R1-Zero?

Because pure-RL reasoning is a mess to read and mixes languages mid-thought. The full recipe interleaves small amounts of supervised data with rounds of RL — each stage fixing what the previous one broke, keeping the emergent reasoning while restoring usability. Four stages:

Interactive Demo — The Four-Stage Pipeline

Press Play stage to walk the pipeline. Each stage lights up with its two-line job description; progress resets with Reset.

Chapter 05

Results — Matching o1, Openly

On the benchmarks that defined the reasoning era, R1 lands at o1-1217's level — with weights anyone can download.

DeepSeek-R1 vs OpenAI-o1-1217 (pass@1 / percentile)
BenchmarkDeepSeek-R1OpenAI-o1-1217
AIME 2024 (pass@1)79.879.2
MATH-50097.396.4
MMLU90.8~91.8
Codeforces (percentile)96.3— (not directly comparable)

AIME and MATH-500 are the paper's exact numbers; o1's MMLU is approximate (~), and its Codeforces rating isn't directly comparable. Beyond benchmarks, R1 stays strong on non-reasoning tasks too — writing, factual QA, summarization.

AIME 2024 · PASS@1
79.8
vs 79.2 for OpenAI-o1-1217
MODEL SIZE (MoE)
671B
total params · only 37B active per token
≈5.5% of weights fire per token
DISTILLATION DATA
800k
R1 samples used to train 6 open models
LICENSE
MIT
weights open for research & commerce
Distillation — Reasoning for Everyone

The team fine-tuned six open small models — Qwen and Llama checkpoints from 1.5B to 70B — on ~800k R1 samples. No RL for the small models: just imitation of R1's verified reasoning traces.

R1-Distill-Qwen-1.5B
R1-Distill-Qwen-7B
R1-Distill-Qwen-14B
R1-Distill-Qwen-32B
R1-Distill-Llama-8B
R1-Distill-Llama-70B
The Surprise — Distillation Beats Re-Learning

The team also trained the same small models with RL directly. Result: the distilled models beat their RL-trained counterparts across benchmarks. For small models, reasoning transfers better by distilling it from a big RL model than by re-learning it from scratch — the big model's expensive exploration gets exported for free.

distill > re-learn (at small scale) 6 models · 1.5B → 70B
Legacy

Impact — The Open Reasoning Wave

One January release reframed what frontier AI costs to build — and who gets to build it.

🧱 The Moat Broke
Frontier-level reasoning with MIT-licensed weights: the closed-reasoning moat around o1-style training evaporated in a week.
📉 Market Shock
R1's release rippled far beyond AI — its arrival in January 2025 fed a tech-stock selloff narrative as investors repriced the cost of frontier models.
🧬 Distilled Everywhere
Six distilled models (down to 1.5B) put test-time reasoning on a laptop — reasoning stopped being a frontier-only luxury.
🧪 Paradigm Validated
RL on verifiable rewards — no human CoT — became a legitimate training paradigm, with R1-Zero's aha moment as exhibit A.
🌊 Everyone Followed
An open reasoning wave followed: R1-style training recipes and open test-time thinkers from lab after lab.
🎓 Textbook Emergence
The aha moment is now the go-to example of emergent behavior — self-correction that was never prompted, rewarded, or taught.
Test Yourself

Quick Quiz

Five questions on the paper that taught models to think for themselves.

Reference

Key Takeaways

Everything you need to remember about DeepSeek-R1.

✅ R1-Zero: pure RL on verifiable rewards (accuracy + format) elicits reasoning — no human CoT data at all.
✅ GRPO: advantages come from a group of sampled responses vs. their own mean — no critic/value network.
✅ Emergent behaviors — self-verification, reflection, backtracking, the "aha moment" — were never prompted or rewarded.
✅ Full R1 recipe: cold-start SFT → reasoning RL → rejection sampling + SFT (~600k + ~200k) → RL for all scenarios.
✅ Results: AIME 2024 79.8 and MATH-500 97.3 (pass@1), from a 671B-total / 37B-active MoE.
✅ MIT-licensed weights + 6 distilled models (1.5B→70B) kicked off the open reasoning wave.