A 2023 counter-argument in model form: skip proprietary data, train smaller models on more public tokens — and LLaMA-13B outperforms GPT-3 (175B) on most benchmarks.
Chinchilla moved the optimum; LLaMA cashed it in with public data and a research-license release.
LLaMA's two wagers: (1) public data suffices — carefully filtered Common Crawl, C4, GitHub, Wikipedia, books, and arXiv beat proprietary corpora when curated; (2) train past Chinchilla-optimal — the models spend more tokens per parameter than the compute-optimal point because inference cost, not training cost, dominates a deployed model's life. Smaller-but-longer-trained models are cheaper forever.
The 2022 status quo: the best models were closed, their data recipes secret, and smaller players priced out.
Closed 2022 models were restaurant dishes — delicious, recipe secret. LLaMA published the cookbook: here is the shopping list (public data), here is the oven schedule (token budget), here is how each portion size (7B-65B) turns out. Anyone with a kitchen can now cook — and millions did.
Everything public, everything filtered, everything documented — the mixture is the paper's most-copied table.
The compute-optimal point optimizes one thing; deployment optimizes another.
Chinchilla's law minimizes training compute for a target loss. LLaMA's 7B and 13B consume ~20× the compute-optimal token count — because after deployment, a model is inferred billions of times and trained once. A smaller model at slightly worse loss is dramatically cheaper to serve, distill, and fine-tune. The paper reports the 65B model trained in 21 days on 2,048 A100-80GB GPUs — everything reproducible in principle from public parts.
The ladder of four sizes turned one release into an entire ecosystem's substrate.
| LLaMA | 7B | 13B | 33B | 65B |
|---|---|---|---|---|
| Training tokens | 1.0T | 1.0T | 1.4T | 1.4T |
| A100 GPU days (paper) | ~82k | ~135k | ~530k | ~860k |
| Beats GPT-3 175B | — | yes — most benchmarks | yes | yes — vs Chinchilla-70B / PaLM-540B class |
Token budgets and device counts from the paper's training table. The 13B 'beats GPT-3' claim is the abstract's headline: state-of-the-art per size class using public data only.
A research release with a non-commercial license still restructured the entire industry.
Check your understanding of the key concepts from LLaMA.
Everything you need to remember about this paper.