History Problem Data GRPO Results Impact Quiz Takeaways
Interactive Paper Explainer

Math from the Crawl,
RL without the Critic
DeepSeekMath

Two moves: mine 120B math tokens out of raw Common Crawl with a iterative data-selection loop, and train with GRPO — PPO with the value network deleted. A 7B open model reaches 51.7% on competition-level MATH.

Start Learning Read the Paper ↗
120B
Math tokens mined
51.7%
MATH (no tools, no voting)
60.9%
MATH w/ 64 SC
7B
Open model
History

The Two Bottlenecks of Math

Data scarcity and RL machinery — DeepSeekMath attacked both in one paper.

2023
Math data is scarce
High-quality math corpora (GSM8K, MATH training sets) are tiny; models starve on problem diversity.
2023
Minerva/MetaMath path
Curated and augmented math sets push open models upward — but the augmentation ceiling arrives fast.
2023
RL is expensive machinery
PPO for LLMs (entries #38, #40) needs value networks, reference models, reward models — four models in memory.
Feb 2024
🚀 DeepSeekMath
Shao et al.: 120B math tokens from Common Crawl via an iterative classifier + GRPO, a PPO variant using group-relative advantages with no critic — 51.7% MATH in a 7B.
2025
The GRPO era
DeepSeek-R1 (entry #47) trains its reasoning with GRPO; the critic-free variant becomes the open world's default RL algorithm.
Two Ideas, One Thesis

The paper's claim: math capability in open models is bottlenecked by data engineering and RL plumbing, not by modeling genius. On data: a fasttext-style classifier, iteratively retrained on positive/negative math pages, extracts 120B math tokens from raw Common Crawl — a 5.9× growth over the Minerva-era corpus (per the paper). On training: GRPO drops PPO's value network and computes advantages from a GROUP of sampled solutions to the same problem — the mean of the group is the baseline, the reward comes from answer correctness. One model to train, not four.

Chapter 01

Starved and Over-Engineered

Math models lacked data; math RL lacked simplicity.

📉
The Data Famine
  • Open math training sets are small — orders of magnitude below the web's actual math content
  • Common Crawl contains vast math (forums, lecture notes, solutions) but drowning in noise
  • Manual curation doesn't scale; keyword filters (e.g. latex-density) plateau quickly
  • Result: open models trail closed ones on competition math by uncomfortable margins
🏗
The RL Overhead
  • PPO-for-LLMs runs policy + value network + reward model + reference policy simultaneously
  • The value (critic) network is the memory hog — and its advantage estimates add tuning burden
  • Simpler RL fixes (RLOO, rejection-sampling SFT) were around but under-validated at scale
  • Deep reasoning needed an RL recipe a small lab could actually run
Analogy — The Prospector and the Lean Crew

Data selection is a prospecting loop: show the sieve a few known gold nuggets, let it learn the look of ore, re-sift the mountain, keep the best dust, retrain the sieve, repeat. GRPO is the lean mining crew: instead of paying a surveyor (critic) to estimate each seam's worth, dig a small cluster of shafts on the same site and grade them against EACH OTHER — the group's average is the benchmark, no surveyor on payroll.

Chapter 02

The Data Engine

The iterative selection loop that turned the crawl into a math corpus.

1️⃣ Seed the classifier
Positive: pools of math content (known corpora pages). Negative: general crawl and non-math pages. A fast classifier learns the surface signature of math.
2️⃣ Score the crawl
Rate billions of Common Crawl pages; keep the top-scoring math candidates.
3️⃣ Filter within pages
Language and dedup rules clean the selected pages; non-math sentences inside them are tolerated (context helps).
4️⃣ Iterate
Newly found high-quality math becomes positive training data for the NEXT classifier round — the loop compounds yield.
The Yield (from the paper)
Interactive Demo — GRPO vs PPO, Step by Step

Follow one training iteration under both algorithms — and watch the critic's job get absorbed by the group.

Chapter 03

GRPO — the Critic-Free Policy Gradient

PPO's machinery, minus the value network — the algorithm R1 later rode to fame.

PPO (before)
  • Advantage  = GAE(reward, value network V(s))
  • Value net ≈ the policy's size — memory doubled
  • Trained jointly; a second optimization to tune
GRPO (this paper)
  • Sample a GROUP of G solutions per problem with the current policy
  • Baseline = group mean reward: Â_i = r_i − mean(r_1..G) (normalized by group spread)
  • Same clipped objective as PPO — stability kept, critic deleted
  • Memory footprint of one policy; straightforward math-domain reward (answer correctness)
The Two-Stage Recipe and Results

Stage 1: math SFT on problem-solution pairs (chain-of-thought). Stage 2: GRPO RL with answer-checking rewards on GSM8K and MATH. Final: 51.7% on MATH without external toolkits and voting — approaching the performance level of Gemini-Ultra and GPT-4 at the time — and 60.9% with 64-sample self-consistency. The GRPO ablations also preview the alignment surprise the follow-up would confirm: the paper observes RL eliciting reasoning behaviors (verification, reflection) beyond the SFT seed — the thread DeepSeek-R1 (entry #47) pulled until the "aha moment" emerged.

Interactive Demo — The Data Sieve

Tab through the iterative corpus-mining loop — the part that turned 'there's no math data' into 'there are 120B math tokens'.

Chapter 05

7B, Open, Competition-Grade

The result that reset expectations for open math models.

GRPO: Â_i = (r_i − mean(r)) / std(r)  ·  clipped PPO objective, no V(s)
G samples
The group
One problem, G policy-sampled solutions — the comparison set that replaces the critic.
mean(r)
The baseline
Group-average reward: a solution is 'advantaged' only by beating its siblings.
std(r)
The scale
Group spread normalizes the advantage — same clipping schedule as PPO applies afterward.
no V(s)
The savings
The value network and its GAE machinery are deleted: one model in memory, one optimizer.
MATH (no tools/voting)
51.7%
approaching Gemini-Ultra and GPT-4 at the time
MATH · 64-sample SC
60.9%
self-consistency voting
CORPUS
120B tokens
math from Common Crawl — ~7× prior open standard
RL MEMORY
1 policy
GRPO: no value network to train or host
Interactive Demo — Open Models on MATH

Press run for the competitive position the paper established — a 7B open model, with and without voting, against the era's references.

Legacy

Legacy — The Open Reasoning Recipe

Data engine + GRPO: the two pieces the reasoning era standardized on.

🚀 The R1 direct ancestor
DeepSeek-R1 (entry #47) is this recipe scaled with emergent-reward surprises — GRPO and the math-data engine are the literal predecessors.
⚙️ GRPO as open default
Critic-free RL became the open community's standard for verifiable-reward training — memory-lean, simple, and now running in countless fine-tunes.
🏗 Data-engineering doctrine
The iterative crawl-classification loop joined the Llama 3 data engine lineage as canonical 'there IS more data, build a sieve' methodology.
🔗 Verifiable-reward alignment
Answer-checking as reward bypassed reward models entirely — the pattern behind the entire reasoning-RL wave.
⚠️ What it did NOT solve
GRPO can reward-hack when 'group beats group' becomes the target; math-domain rewards don't transfer to open-ended tasks; and the data loop inherits crawl-language bias (the paper itself audits an English/Chinese imbalance).
🛤 Read next
Test Yourself

Quick Quiz

Check your understanding of the key concepts from DeepSeekMath.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ 120B math tokens mined from Common Crawl via an iterative classifier loop — the data half of the recipe.
✅ GRPO: PPO's clipped objective with group-relative advantages — the critic network deleted.
✅ 51.7% on MATH without tools or voting; 60.9% with 64-sample self-consistency — approaching GPT-4-era level in a 7B.
✅ Base: DeepSeek-Coder-1.5 7B continued on math + code + natural language.
✅ Verifiable answer-checking rewards replace reward models in RL — the reasoning-RL pattern.
✅ Read it as DeepSeek-R1's direct ancestor: same data doctrine, same algorithm, earlier scale.