Smaller models, more tokens, open weights. LLaMA-65B rivals the best closed systems of its day, and LLaMA-13B beats GPT-3 (175B) on most benchmarks — trained on 1.4T tokens of entirely public data. This is the paper that detonated the open-model ecosystem.
LLaMA landed at the peak of the "bigger is better — and closed" era. Here is the road that led to it, and the explosion that followed.
Chinchilla asked: given a training budget, what model is best? LLaMA asks the deployment question: given a model you will serve millions of times, what should you have trained? Once inference dominates the cost, the optimum moves to smaller models trained longer.
By 2022, the best language models were black boxes. LLaMA's bet: you don't need 175B parameters — or a private dataset — to reach that class of performance.
Can we match GPT-3-class performance with smaller models trained longer on more public data? Think of closed frontier models as restaurants: you can order from the menu (an API), but you can never walk into the kitchen. BLOOM and OPT handed out photos of a huge kitchen that cooked worse. LLaMA hands researchers the keys to a compact kitchen that plates nearly the same dishes — and lets you renovate it.
Chinchilla's rule (~20 tokens/param) minimizes training compute. LLaMA borrows the spirit and breaks the letter at the small end: overtraining buys cheaper inference, forever.
| Model | Params | Tokens | Tokens / param | Strategy |
|---|---|---|---|---|
| LLaMA-7B | 7B | 1.4T | 200 | 10× beyond optimal — built for cheap serving |
| LLaMA-13B | 13B | 1.4T | 108 | The GPT-3 beater |
| LLaMA-33B | 33B | 1.4T | 42 | ~2× optimal — quality vs serving trade-off |
| LLaMA-65B | 65B | 1.4T | 22 | ≈ Chinchilla-optimal — maximum quality |
Every size sees the identical 1.4T-token corpus — the family doubles as a clean study of "more tokens, fewer params."
Chinchilla minimizes training compute — a cost you pay once. A deployed model pays inference on every request, forever, and inference cost scales with parameter count. A 7B model overtrained to 200 tokens/param is wasteful to train once and cheap to serve a billion times.
At 1.4T tokens, even the smallest model (7B) sees a huge, deduplicated, diverse corpus. One corpus, four training budgets: the smaller sizes simply trade peak quality for serving cost — and all four publish as a clean scaling family.
No private scrapes, no licensed archives: every one of the 1.4 trillion tokens is publicly available, so anyone can audit — and in principle rebuild — the corpus.
The mix mirrors what a general-purpose model needs: web prose for breadth, code for programming, Wikipedia for facts, books for long-horizon reasoning, ArXiv for math, StackExchange for dialogue-shaped Q&A. Web text dominates by volume; the curated sources dominate by signal per token.
Because every source is public, the recipe is inspectable and reproducible in principle — a deliberate contrast with closed frontier training sets. It also kept the release story simple: no proprietary data to license around, weights to share with researchers.
No exotic machinery: a standard GPT-style decoder-only Transformer. Three targeted upgrades over the GPT-3 recipe, each borrowed from the 2021–22 literature.
| Component | GPT-3 | LLaMA |
|---|---|---|
| Normalization | LayerNorm | RMSNorm, pre-norm |
| FFN activation | GELU | SwiGLU (gated) |
| Positions | Learned absolute | RoPE (rotary) |
| Tokenizer | BPE, ~50k | SentencePiece BPE, ~32k |
| Context length | 2048 | 2048 (unchanged) |
| Objective | Next-token prediction | Same — nothing new needed |
The lesson: LLaMA's edge is not architectural — it's data and training. The best 2023 decoder is barely three components away from GPT-3.
One GPU cluster, three weeks, public data — and results that made closed models 10× larger look inefficient.
| System | Params | MMLU 5-shot | Notes |
|---|---|---|---|
| GPT-3 (closed) | 175B | 43.9 | API only; ~300B training tokens |
| LLaMA-13B (open) | 13B | ~55 | Beats GPT-3 with ~13× fewer parameters |
| LLaMA-65B (open) | 65B | 63.4 | Public data only; rivals the frontier |
| Chinchilla (closed) | 70B | 67.5 | DeepMind; the compute-optimal reference |
| PaLM (closed) | 540B | ~69 | 8× larger than LLaMA-65B |
Beyond MMLU: strong zero-shot common-sense reasoning (CommonsenseQA, PIQA, HellaSwag, ARC, BoolQ), and the 65B model stays competitive with Chinchilla-70B and PaLM-540B across many evaluations.
A gated research release became the accidental ignition point of the entire open-weights ecosystem.
Check your understanding of the key concepts from the LLaMA paper.
Everything you need to remember about this paper.