Two moves: mine 120B math tokens out of raw Common Crawl with a iterative data-selection loop, and train with GRPO — PPO with the value network deleted. A 7B open model reaches 51.7% on competition-level MATH.
Data scarcity and RL machinery — DeepSeekMath attacked both in one paper.
The paper's claim: math capability in open models is bottlenecked by data engineering and RL plumbing, not by modeling genius. On data: a fasttext-style classifier, iteratively retrained on positive/negative math pages, extracts 120B math tokens from raw Common Crawl — a 5.9× growth over the Minerva-era corpus (per the paper). On training: GRPO drops PPO's value network and computes advantages from a GROUP of sampled solutions to the same problem — the mean of the group is the baseline, the reward comes from answer correctness. One model to train, not four.
Math models lacked data; math RL lacked simplicity.
Data selection is a prospecting loop: show the sieve a few known gold nuggets, let it learn the look of ore, re-sift the mountain, keep the best dust, retrain the sieve, repeat. GRPO is the lean mining crew: instead of paying a surveyor (critic) to estimate each seam's worth, dig a small cluster of shafts on the same site and grade them against EACH OTHER — the group's average is the benchmark, no surveyor on payroll.
The iterative selection loop that turned the crawl into a math corpus.
PPO's machinery, minus the value network — the algorithm R1 later rode to fame.
Stage 1: math SFT on problem-solution pairs (chain-of-thought). Stage 2: GRPO RL with answer-checking rewards on GSM8K and MATH. Final: 51.7% on MATH without external toolkits and voting — approaching the performance level of Gemini-Ultra and GPT-4 at the time — and 60.9% with 64-sample self-consistency. The GRPO ablations also preview the alignment surprise the follow-up would confirm: the paper observes RL eliciting reasoning behaviors (verification, reflection) beyond the SFT seed — the thread DeepSeek-R1 (entry #47) pulled until the "aha moment" emerged.
The result that reset expectations for open math models.
Data engine + GRPO: the two pieces the reasoning era standardized on.
Check your understanding of the key concepts from DeepSeekMath.
Everything you need to remember about this paper.