History Problem Methods Rule of Thumb Results Impact Quiz Takeaways
Interactive Paper Explainer

Compute-Optimal Scaling
Chinchilla

An interactive guide to "Training Compute-Optimal Large Language Models" — the DeepMind paper that showed the GPT-3 era was training models too big for their data. Same compute as Gopher, but 4× smaller and trained on 4× more data: Chinchilla rewrote the scaling rulebook and beat models 4× its size.

Start Learning Read the Paper ↗
70B
Parameters
1.4T
Training Tokens
4
Estimation Methods
2022
Year Published
History

From Bigger Is Better to Better Balanced

The Kaplan scaling laws told labs to spend new compute mostly on parameters. Two years of ever-larger models later, DeepMind re-derived the optimum — and found it is far more balanced. This page continues the story of the Scaling Laws guide.

2020 · Jan
Kaplan et al. — Scaling Laws
Power laws for loss vs N, D, C — and a compute-optimal recipe: N ∝ C0.73. Parameters should grow faster than data. Scaling Laws guide →
2020 · May
GPT-3 (175B)
175B parameters trained on 300B tokens — planned on the Kaplan-era curves. Roughly 1.7 tokens per parameter.
2021
Gopher (280B) & Jurassic-1 (178B)
The "bigger is better" doctrine peaks: Gopher adds parameters but keeps the ~300B-token data budget.
2022 · Mar
🚀 Chinchilla (70B / 1.4T)
Same compute as Gopher, 4× fewer parameters, 4× more tokens — and better performance on the large majority of evals.
2023 →
LLaMA & the data-centric era
Training data and token budgets become headline numbers. LLaMA guide →
Key Insight

Kaplan's recipe splits each 10× of new compute 73 : 27 toward parameters. Chinchilla's four estimation methods all landed on the same answer: the split should be 50 : 50. Scale parameters and tokens together.

HOW EACH RECIPE SPENDS NEW COMPUTE
Kaplan 2020 — N ∝ C0.73
73% params
27%
Chinchilla 2022 — N ∝ C0.50
50% params
50% data
Blue = share of new compute toward parameters · Teal = toward data (log-space split, illustrative).
Chapter 01

The Problem — Big Models, Starved of Data

The Kaplan-era "bigger is better" doctrine made frontier models too large and undertrained. GPT-3 and Gopher used far fewer tokens per parameter than optimal — wasting training compute and inflating serving costs.

🐘
Overparameterized & Data-Starved
  • GPT-3: ~1.7 tokens per parameter; Gopher: ~1.1 — far below the optimum
  • Most parameters arrive undertrained — capacity goes to waste
  • Serving cost scales with N: every user query pays for the excess size
  • The inefficiency is hidden — the models still look impressive
  • The same compute, spent differently, would buy better performance
⚖️
The Compute-Optimal Fix
  • Right-size the model: 4× smaller than Gopher at the same compute
  • Feed it 4× more tokens — 1.4T instead of 300B
  • Same training budget, better loss and better benchmarks
  • Smaller N means cheaper inference for the model's whole life
  • A rule you can check before you train: ~20 tokens per parameter
Analogy — The Race Engine with a Tiny Fuel Tank

GPT-3 and Gopher were race cars built around an enormous engine with a tiny fuel tank. The engine (parameters) promised speed, but the fuel (tokens) ran out long before it hit its stride — and you pay for engine size on every single lap (inference). Chinchilla's question: given the same money for the whole car, what is the best engine-to-fuel split? The answer was not a bigger engine. It was a smaller one with 4× the fuel.

Chapter 02

The Core Idea — Four Ways to the Same Optimum

Chinchilla did not just train a better model — it re-derived the compute-optimal scaling law. Three independent estimation methods, plus a combined analysis, all landed on the same answer.

The Four Approaches
📐
1 · Parametric Loss Fit
Fit L(N,D) = E + A/Nα + B/Dβ to ~400 training runs, then minimize loss subject to a compute budget.
🎢
2 · IsoFLOP Profiles
Fix the compute budget, vary model size N, fit a parabola through the losses — its minimum marks the optimal N.
📉
3 · Training-Curve Minima
For each model size N, find the token count D at which further training stops helping.
🤝
4 · All Methods Agree
N ∝ C0.50, D ∝ C0.50 — parameters and tokens should scale equally.
L(N, D) = E + A / Nα + B / Dβ  ·  α ≈ 0.34, β ≈ 0.28
E
Irreducible loss
The floor no model can beat — the entropy of natural text itself.
A / Nα
Capacity term
Loss from the model being too small. Falls as parameters N grow; α ≈ 0.34.
B / Dβ
Data term
Loss from seeing too few tokens. Falls as data D grows; β ≈ 0.28.
~400
Training runs in the fit
The joint fit spans a wide range of model sizes and token counts — both N and D fitted together.
What Changed vs Kaplan

Kaplan et al. (2020) fitted loss scaling in a way that implied N* ∝ C0.73 — spend most new compute on parameters. Chinchilla re-derived the optimum with a joint fit over both N and D, across ~400 runs, and double-checked it with two independent methods. Revised answer: N* ∝ C0.50 — parameters and tokens grow in lockstep. If you studied the Scaling Laws guide, this is the plot twist.

Why Four Methods?

A single parametric fit can be fooled by its own functional form. IsoFLOP profiles and training-curve minima attack the same question without assuming L(N,D) at all. When three independent approaches — and the combined analysis — all land on exponent ≈ 0.50, the conclusion becomes very hard to doubt.

Interactive Demo — Compute Budget Allocator (Chinchilla)

Pick a compute budget and compare the two recipes. The Chinchilla line follows N* ∝ C0.50, calibrated so 1023 FLOPs → N* ≈ 70B params on D* ≈ 1.4T tokens. The Kaplan-style line follows N* ∝ C0.73, anchored so the same budget buys a Gopher-sized model on Gopher-sized data. Illustrative.

Interactive Demo — Kaplan vs Chinchilla: The Exponent Gap

Both recipes are power laws, so they nearly agree at small budgets — and diverge as compute grows. Toggle what the bars compare. Illustrative.

Chapter 03

The Rule of Thumb — ~20 Tokens per Parameter

The paper's most quotable result: at the compute-optimal point, every parameter gets about 20 training tokens. It is the quick sanity check every lab now runs before a training run.

D* ≈ 20 · N  ·  "20 tokens per parameter"
N
Parameters
Model size — the thing Kaplan-era scaling told you to grow fastest.
D*
Training tokens
The data budget the optimum demands — 1.4T tokens for a 70B model.
≈ 20
The ratio
Read straight off the N ∝ C0.50, D ∝ C0.50 agreement — both resources scale equally.
catch
Training compute only
The rule minimizes training cost. Deployed models deliberately break it — see Chapter 05.
Interactive Demo — Is Your Model Undertrained?

Check famous models against the ~20 tokens-per-parameter optimum. Red = undertrained, green = at or beyond the optimum.

ALL MODELS — TOKENS PER PARAMETER (LOG SCALE, | = OPTIMUM ~20)
How to Use It

Before training: divide your planned token budget by 20 — that is the parameter count the rule endorses. Or multiply your planned parameter count by 20 to budget data collection. It is a heuristic, not a law — but it is the heuristic that ended the parameter race.

Chapter 04

Results — Smaller, Yet Stronger

Chinchilla used the same compute as Gopher — but outperforms it on the large majority of evaluations, and beats models more than twice its size.

The Head-to-Head — Mean MMLU
ModelParamsTraining TokensMMLU (5-shot)
GPT-3175B300B43.9*
Gopher280B300B60.0
Chinchilla70B1.4T67.6 (+7.6)

*GPT-3 measured 5-shot. Same compute as Gopher, a quarter of the parameters, 4× the tokens — and a 7.6-point MMLU jump.

MEAN MMLU
67.6
vs 60.0 for Gopher — +7.6 points
VS GOPHER
4×
smaller, at the same training compute
TRAINING TOKENS
1.4T
vs 300B for GPT-3 & Gopher
RATIO
20
tokens per parameter — the optimum
🏆 Same compute, better model
Chinchilla outperforms Gopher on the large majority of evaluations — using the same training compute with 4× fewer parameters.
🥊 GPT-3 & Jurassic-1 beaten
Outperforms the 175B GPT-3 and 178B Jurassic-1 despite having ~2.5× fewer parameters.
⛰️ Even vs 530B
MT-NLG 530B — more than 7× Chinchilla's size — is outperformed on many tasks by this 70B model.
📚 Broad wins
Strong performance across LAMBADA, BIG-bench, and reading-comprehension benchmarks.
Chapter 05

Aftermath — Compute-Optimal ≠ Inference-Optimal

Chinchilla minimizes training compute. But a deployed model pays for inference every day after training — and that changed how the rule gets used.

The Inference-Cost Argument

The 20-tokens-per-parameter rule is optimal for a model you train once and throw away. Real models serve millions of requests after training, so lifetime cost is dominated by inference — which scales with N. Training "too long" on a smaller model trades cheap one-time compute for expensive forever-compute.

What Labs Actually Do

LLaMA-style models deliberately exceed the Chinchilla ratio — LLaMA-7B trained on ~200 tokens per parameter — paying extra training compute to shrink N and make serving cheap. "Chinchilla-optimal" and "inference-optimal" are now both part of the vocabulary. LLaMA guide →

🦙 The LLaMA turn
Meta's LLaMA made the data-centric strategy explicit: small models, huge token budgets. LLaMA guide →
📏 Ratios now reported
Model announcements quote tokens-per-parameter alongside parameter counts — a Chinchilla invention of convenience.
✨ Data-quality era
If tokens matter as much as parameters, data curation, dedup, and filtering became first-class engineering.
🐣 Small-model renaissance
Right-sized 7B–70B models became first-class citizens — not just distillation targets.
🧩 The MoE alternative
Mixture-of-experts models grow capacity without growing per-token FLOPs — a different escape from the rule.
🗺️ Every lab re-planned
Scaling curves were re-fit and frontier runs re-priced. "Compute-optimal" entered the standard vocabulary.
Test Yourself

Quick Quiz

Check your understanding of the key concepts from the Chinchilla paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Same compute as Gopher, but 70B parameters on 1.4T tokens — 4× smaller, 4× more data.
✅ Four estimation methods, one answer: N* ∝ C^0.50, D* ∝ C^0.50 — scale parameters and tokens equally.
✅ The joint fit: L(N,D) = E + A/N^0.34 + B/D^0.28, fitted on ~400 training runs.
✅ Rule of thumb: ~20 training tokens per parameter.
✅ Mean MMLU 67.6 vs Gopher's 60.0 — and it beats GPT-3 (175B) and Jurassic-1 (178B) with 2.5× fewer parameters.
✅ Compute-optimal ≠ inference-optimal: deployed models (LLaMA) deliberately overtrain for cheap inference.