An interactive guide to "Scaling Laws for Neural Language Models" — the paper showing that performance depends strongly on scale and only weakly on shape, and that smooth power laws predict loss from model size, data, and compute across many orders of magnitude.
Before 2020, making models bigger was an art. Kaplan et al. turned it into quantitative science — and the field used it to plan the largest training runs ever attempted.
A power law means loss falls as a straight line on a log-log plot: multiply the resources, subtract a fixed amount of loss. That straightness — holding over six-plus orders of magnitude — is what turned scale from a gamble into a planning problem.
Before scaling laws, every size increase was a bet. How big should the next model be? How much data does it need? When do you stop training? There was no quantitative guidance.
Before stress calculations, builders made bridges thick and hoped. Stress math turned bridge-building into engineering: given a load, the numbers tell you the beams. Kaplan et al. did the same for compute. Given a budget C, the laws say how big to build the model (N), how much data to feed it (D), and what loss to expect when you stop — so the guesswork, and the costly over-engineering, goes away.
Test loss follows smooth power laws in each of the three resource axes. On a log-log plot, loss versus scale is a straight line — one simple formula per resource.
Plot loss against N, D, or C with both axes logarithmic, and the points land on a straight line — across six-plus orders of magnitude of compute. A straight line in log-log space is the fingerprint of a power law: multiply the resources by 10, subtract a fixed amount of loss. The fits hold from toy models up to 1.5B parameters, which is exactly why extrapolating beyond them became credible.
Each resource obeys its own law, and they interlock: compute is roughly C ≈ 6ND, tying FLOPs to model size and data. Fit the constants on cheap, small runs, then read off the expected loss of a run you could never afford to explore blindly — the extrapolation that later sized GPT-3.
Within broad ranges, performance depends strongly on total scale — and only weakly on the shape of the model. Depth versus width is nearly a wash.
Models with the same non-embedding parameter count but very different depth-to-width ratios — shallow-wide versus deep-narrow Transformers — land on essentially the same loss curve. Shape knobs matter within broad ranges, but they are second-order effects compared to total scale: a kind of "universality" across architectures.
If shape barely matters, you don't need to search architectures before scaling. Pick a workable shape and spend the budget on size and data — the ceiling is set by N, D, and C, which you can plan, not by architecture roulette.
Given a fixed compute budget C, don't train a small model to convergence. Train a larger model and stop early — the best allocation follows N ∝ C0.73.
For every budget C there is a best (N, D) pair — the efficient frontier. Depart from it and you pay: the same FLOPs land you at a worse loss. The frontier is what lets a lab say "with our budget, build exactly this model, on exactly this much data, and stop exactly here."
Because loss curves are power laws in training time too, a big model that has not converged still beats a smaller one that has — at equal compute. So spend the budget on size, not on squeezing out the last few percent of training. Most compute should go to bigger models on modest data.
GPT-3 — 175B parameters trained on 300B tokens — followed this paper's compute-optimal prescription within its framework: the first 100× scale-up planned on curves instead of hope.
| Model | Parameters | Training Tokens | Year |
|---|---|---|---|
| GPT-2 | 1.5B | ~10B (WebText) | 2019 |
| GPT-3 | 175B | 300B | 2020 |
| Chinchilla | 70B | 1.4T | 2022 |
Honest footnote: Chinchilla (2022) revised the compute-optimal exponents — optimal models need more data per parameter than Kaplan's laws suggested. Read the two papers together → Chinchilla guide.
Scaling laws turned "bigger is better" into a quantitative research program that guided the largest training runs in history.
Check your understanding of the key concepts from the scaling-laws paper.
Everything you need to remember about this paper.