An interactive guide to "Training Compute-Optimal Large Language Models" — the DeepMind paper that showed the GPT-3 era was training models too big for their data. Same compute as Gopher, but 4× smaller and trained on 4× more data: Chinchilla rewrote the scaling rulebook and beat models 4× its size.
The Kaplan scaling laws told labs to spend new compute mostly on parameters. Two years of ever-larger models later, DeepMind re-derived the optimum — and found it is far more balanced. This page continues the story of the Scaling Laws guide.
Kaplan's recipe splits each 10× of new compute 73 : 27 toward parameters. Chinchilla's four estimation methods all landed on the same answer: the split should be 50 : 50. Scale parameters and tokens together.
The Kaplan-era "bigger is better" doctrine made frontier models too large and undertrained. GPT-3 and Gopher used far fewer tokens per parameter than optimal — wasting training compute and inflating serving costs.
GPT-3 and Gopher were race cars built around an enormous engine with a tiny fuel tank. The engine (parameters) promised speed, but the fuel (tokens) ran out long before it hit its stride — and you pay for engine size on every single lap (inference). Chinchilla's question: given the same money for the whole car, what is the best engine-to-fuel split? The answer was not a bigger engine. It was a smaller one with 4× the fuel.
Chinchilla did not just train a better model — it re-derived the compute-optimal scaling law. Three independent estimation methods, plus a combined analysis, all landed on the same answer.
Kaplan et al. (2020) fitted loss scaling in a way that implied N* ∝ C0.73 — spend most new compute on parameters. Chinchilla re-derived the optimum with a joint fit over both N and D, across ~400 runs, and double-checked it with two independent methods. Revised answer: N* ∝ C0.50 — parameters and tokens grow in lockstep. If you studied the Scaling Laws guide, this is the plot twist.
A single parametric fit can be fooled by its own functional form. IsoFLOP profiles and training-curve minima attack the same question without assuming L(N,D) at all. When three independent approaches — and the combined analysis — all land on exponent ≈ 0.50, the conclusion becomes very hard to doubt.
The paper's most quotable result: at the compute-optimal point, every parameter gets about 20 training tokens. It is the quick sanity check every lab now runs before a training run.
Before training: divide your planned token budget by 20 — that is the parameter count the rule endorses. Or multiply your planned parameter count by 20 to budget data collection. It is a heuristic, not a law — but it is the heuristic that ended the parameter race.
Chinchilla used the same compute as Gopher — but outperforms it on the large majority of evaluations, and beats models more than twice its size.
| Model | Params | Training Tokens | MMLU (5-shot) |
|---|---|---|---|
| GPT-3 | 175B | 300B | 43.9* |
| Gopher | 280B | 300B | 60.0 |
| Chinchilla | 70B | 1.4T | 67.6 (+7.6) |
*GPT-3 measured 5-shot. Same compute as Gopher, a quarter of the parameters, 4× the tokens — and a 7.6-point MMLU jump.
Chinchilla minimizes training compute. But a deployed model pays for inference every day after training — and that changed how the rule gets used.
The 20-tokens-per-parameter rule is optimal for a model you train once and throw away. Real models serve millions of requests after training, so lifetime cost is dominated by inference — which scales with N. Training "too long" on a smaller model trades cheap one-time compute for expensive forever-compute.
LLaMA-style models deliberately exceed the Chinchilla ratio — LLaMA-7B trained on ~200 tokens per parameter — paying extra training compute to shrink N and make serving cheap. "Chinchilla-optimal" and "inference-optimal" are now both part of the vocabulary. LLaMA guide →
Check your understanding of the key concepts from the Chinchilla paper.
Everything you need to remember about this paper.