History Problem Power Laws Findings Compute Impact Quiz Takeaways
Interactive Paper Explainer

The Power Laws
Scaling Laws

An interactive guide to "Scaling Laws for Neural Language Models" — the paper showing that performance depends strongly on scale and only weakly on shape, and that smooth power laws predict loss from model size, data, and compute across many orders of magnitude.

Start Learning Read the Paper ↗
6+
Orders of Compute Studied
1.5B
Largest Model in Study
3
Resource Axes (N, D, C)
2020
Year Published
History

From Craft to a Science of Scale

Before 2020, making models bigger was an art. Kaplan et al. turned it into quantitative science — and the field used it to plan the largest training runs ever attempted.

2012–2017
Hand-crafted architectures
Deep learning gains came from clever architectures — CNNs, LSTMs, attention. Bigger helped, but nobody could say by how much, or predict it ahead of time.
2018
GPT-1 & BERT — pre-training arrives
Pre-train on unlabeled text, fine-tune on tasks. Scale starts to matter — but each size jump is still a costly experiment, not a prediction.
2019
GPT-2 (1.5B)
A 1.5-billion-parameter model shows surprisingly broad language abilities. The hunch forms: maybe raw scale is the secret ingredient.
2020 · Jan
🚀 Kaplan et al. — Scaling Laws
Empirical laws: test loss follows smooth power laws in parameters N, data D, and compute C — over six orders of magnitude. Scaling becomes predictable.
2020 · May
GPT-3 (175B)
A 100× jump over GPT-2, designed using the scaling-law framework — planned on curves, not guesswork. Few-shot learning emerges at scale.
2022
Chinchilla — the correction
DeepMind shows the compute-optimal balance leans more toward data (N ∝ C^0.5, not C^0.73). Study it together with the Chinchilla guide.
Key Insight

A power law means loss falls as a straight line on a log-log plot: multiply the resources, subtract a fixed amount of loss. That straightness — holding over six-plus orders of magnitude — is what turned scale from a gamble into a planning problem.

LOSS vs RESOURCES — LOG-LOG (ILLUSTRATIVE)
1×
10×
10²×
10³×
10⁴×
10⁵×
10⁶×
Equal steps down in loss for equal 10× steps in resources — a straight line in log-log space.
Chapter 01

The Problem — Scaling by Guesswork

Before scaling laws, every size increase was a bet. How big should the next model be? How much data does it need? When do you stop training? There was no quantitative guidance.

🎲
Scaling Without Laws
  • How big should the next model be? Nobody could predict
  • How much data? Folklore and expensive trial-and-error
  • When to stop training? Watch the loss curve and hope
  • Wasted compute on runs that under-delivered
  • Unpredictable performance made planning impossible
📐
Scaling With Laws
  • Smooth, predictable power laws for loss vs scale
  • Compute-optimal recipes: the best loss per FLOP
  • Plan and price a run before spending the compute
  • Know the loss floor a model size can reach
  • A science of scale, not folklore
Analogy — Building Without Stress Math

Before stress calculations, builders made bridges thick and hoped. Stress math turned bridge-building into engineering: given a load, the numbers tell you the beams. Kaplan et al. did the same for compute. Given a budget C, the laws say how big to build the model (N), how much data to feed it (D), and what loss to expect when you stop — so the guesswork, and the costly over-engineering, goes away.

Chapter 02

The Core Idea — Power Laws

Test loss follows smooth power laws in each of the three resource axes. On a log-log plot, loss versus scale is a straight line — one simple formula per resource.

L(N) = (Nc / N)αN  ·  αN ≈ 0.076
L(D) = (Dc / D)αD  ·  αD ≈ 0.095
L(C) = (Cc / C)αC  ·  αC ≈ 0.050
N
Non-embedding parameters
Model size, counting only transformer-body parameters. The strongest single lever: L(N) = (N_c/N)^0.076.
D
Dataset size (tokens)
Number of training tokens. Data obeys its own power law: L(D) = (D_c/D)^0.095.
C
Training compute
Estimated non-embedding training FLOPs (the paper reports PF-days). L(C) = (C_c/C)^0.050.
L
Cross-entropy loss
Test loss on held-out text — the single number all three laws predict.
Nc, Dc, Cc
Critical constants
Scale thresholds from the fit — the resources must pass these before the power law takes hold.
α ≈ 0.05–0.10
Small exponents
Each 10× in resources buys a modest but very predictable slice of loss.
Read It Off a Log-Log Plot

Plot loss against N, D, or C with both axes logarithmic, and the points land on a straight line — across six-plus orders of magnitude of compute. A straight line in log-log space is the fingerprint of a power law: multiply the resources by 10, subtract a fixed amount of loss. The fits hold from toy models up to 1.5B parameters, which is exactly why extrapolating beyond them became credible.

Three Levers, One Recipe

Each resource obeys its own law, and they interlock: compute is roughly C ≈ 6ND, tying FLOPs to model size and data. Fit the constants on cheap, small runs, then read off the expected loss of a run you could never afford to explore blindly — the extrapolation that later sized GPT-3.

Interactive Demo — Power-Law Loss Explorer

Pick a model size and watch the loss predicted by L(N) = (Nc/N)αN with αN ≈ 0.076 and Nc = 8.8×1013. Every 10× in parameters buys roughly the same cut in loss — the signature of a power law. Loss values are illustrative.

Chapter 03

Finding 1 — Shape Barely Matters

Within broad ranges, performance depends strongly on total scale — and only weakly on the shape of the model. Depth versus width is nearly a wash.

Depth ⇄ Width — Almost a Wash

Models with the same non-embedding parameter count but very different depth-to-width ratios — shallow-wide versus deep-narrow Transformers — land on essentially the same loss curve. Shape knobs matter within broad ranges, but they are second-order effects compared to total scale: a kind of "universality" across architectures.

Why That's Liberating

If shape barely matters, you don't need to search architectures before scaling. Pick a workable shape and spend the budget on size and data — the ceiling is set by N, D, and C, which you can plan, not by architecture roulette.

⚖️ Universality of Overfitting
N and D must scale together — roughly D ∝ N0.74 — or the loss floor rises: more parameters on the same data eventually stop helping.
📈 Universality of Training
Late-time loss is predictable from the early training curve — the same learning-curve shape shows up across scales.
🔀 Transfer
Models trained on one text distribution still do well on others after fine-tuning — scale, not your corpus choice, sets the ceiling.
⚡ Sample Efficiency
Large models reach any target loss with fewer training samples — big models learn more per example.
Interactive Demo — The N vs D Balance

You've built a 100B-parameter model. Slide the dataset size: starve it of data and the loss floor rises (overfitting); drown it in data and the model can't absorb it all (undertrained). The efficient frontier sits near D ∝ N0.74. Illustrative.

DATA BUDGET
← overfitbalanceddata-rich →
Chapter 04

Finding 2 — Compute-Optimal Training

Given a fixed compute budget C, don't train a small model to convergence. Train a larger model and stop early — the best allocation follows N ∝ C0.73.

N* ∝ C0.73  ·  D* ∝ C0.27  →  train larger, stop early
C
Fixed compute budget
The FLOPs you can afford — the constraint every training decision hangs on.
N* ∝ C^0.73
Optimal model size
The best N grows fast with budget — most new compute buys parameters.
D* ∝ C^0.27
Optimal dataset size
The best D grows slowly — bigger models run on relatively modest data.
stop early
Never converge
The optimum sits before convergence — an unconverged big model beats a converged small one at equal compute.
The Efficient Frontier

For every budget C there is a best (N, D) pair — the efficient frontier. Depart from it and you pay: the same FLOPs land you at a worse loss. The frontier is what lets a lab say "with our budget, build exactly this model, on exactly this much data, and stop exactly here."

Why Stopping Early Wins

Because loss curves are power laws in training time too, a big model that has not converged still beats a smaller one that has — at equal compute. So spend the budget on size, not on squeezing out the last few percent of training. Most compute should go to bigger models on modest data.

Interactive Demo — Compute Budget Allocator

Choose a total compute budget. The bars show the Kaplan-optimal model size N* (following N* ∝ C0.73) and dataset size D* (from C ≈ 6ND), anchored so the ~1023 FLOP point matches GPT-3's 175B run. Watch the tokens-per-parameter bar fall as the budget grows — the model grows faster than the data. Illustrative.

Chapter 05 — GPT-3: The Proof at Scale

GPT-3 — 175B parameters trained on 300B tokens — followed this paper's compute-optimal prescription within its framework: the first 100× scale-up planned on curves instead of hope.

ModelParametersTraining TokensYear
GPT-21.5B~10B (WebText)2019
GPT-3175B300B2020
Chinchilla70B1.4T2022

Honest footnote: Chinchilla (2022) revised the compute-optimal exponents — optimal models need more data per parameter than Kaplan's laws suggested. Read the two papers together → Chinchilla guide.

Legacy

Impact — The Scaling Hypothesis

Scaling laws turned "bigger is better" into a quantitative research program that guided the largest training runs in history.

🧭 A Research Program
The "scaling hypothesis" — that capabilities keep improving as scale grows — became a deliberate program, not a hope.
📐 Compute Planning
Compute-optimal curves guided GPT-3's 175B run and much that followed — scale-ups sized on paper before hardware.
🔮 Loss Forecasting
Teams now fit the laws on small runs and predict the loss of frontier runs before committing the compute.
⚖️ Frontier Thinking
"Train larger, stop early" reshaped training practice — convergence stopped being the target at every serious lab.
📚 Chinchilla (2022)
Data scaling laws rebalanced the optimum toward more tokens per parameter — read it alongside this guide.
🚀 The Scale-Up Era
From GPT-3 onward, frontier models were bets placed on curves — see the GPT-3 guide.
Test Yourself

Quick Quiz

Check your understanding of the key concepts from the scaling-laws paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Smooth power laws: L ≈ (N_c/N)^0.076, (D_c/D)^0.095, (C_c/C)^0.050 — straight lines on log-log plots.
✅ Scale ≫ shape: loss depends strongly on N, D, C and only weakly on depth/width within broad ranges.
✅ The fits hold over six-plus orders of magnitude of compute — from toy models to 1.5B parameters.
✅ Compute-optimal: at fixed C, N ∝ C^0.73 — spend on bigger models on modest data, stopped early.
✅ N and D must grow together (≈ D ∝ N^0.74) or overfitting raises the loss floor.
✅ GPT-3 (175B, 300B tokens) was planned on these curves — Chinchilla (2022) later rebalanced toward more data.