History Problem Core Idea Efficiency Results Impact Quiz Takeaways
Interactive Paper Explainer

Open Weights, Trained Longer
LLaMA

A 2023 counter-argument in model form: skip proprietary data, train smaller models on more public tokens — and LLaMA-13B outperforms GPT-3 (175B) on most benchmarks.

Start Learning Read the Paper ↗
7B → 65B
Model sizes
1.0T → 1.4T
Training tokens
100%
Public data
2023
Research release
History

The Compute Rebalance

Chinchilla moved the optimum; LLaMA cashed it in with public data and a research-license release.

2020
GPT-3 and the moat
175B parameters trained on filtered web-scale corpora; access only via API — capability locked behind a frontier-lab wall.
2022
Chinchilla rebalance
Hoffmann et al. showed models were under-trained: same compute, more tokens, smaller model wins — the training recipe, not just size, was the lever.
Feb 2023
🚀 LLaMA
Meta AI releases 7B/13B/33B/65B models trained on trillions of tokens from entirely public sources — state-of-the-art at each size, weights to researchers.
Mar 2023+
The leak and the boom
Within a week of the weights leaking, the community fine-tuned them on a single GPU — Alpaca, Vicuna, and the open-weights ecosystem detonated.
2023-25
Open goes frontier
Llama 2 (chat + commercial), Llama 3, Mistral, DeepSeek — the lineage LLaMA started now competes with closed models.
The Bet that Aged Perfectly

LLaMA's two wagers: (1) public data suffices — carefully filtered Common Crawl, C4, GitHub, Wikipedia, books, and arXiv beat proprietary corpora when curated; (2) train past Chinchilla-optimal — the models spend more tokens per parameter than the compute-optimal point because inference cost, not training cost, dominates a deployed model's life. Smaller-but-longer-trained models are cheaper forever.

Chapter 01

Frontier = Locked

The 2022 status quo: the best models were closed, their data recipes secret, and smaller players priced out.

🔒
The Closed Frontier
  • GPT-3-class models were API-only: no weights, no data recipe, no scrutiny of training mix
  • Proprietary web-scale datasets were unreplicable, so open models trailed by a visible margin
  • Research on interpretability, safety, and efficiency needed weight access that licenses forbade
  • Every deployment paid frontier-API pricing for capabilities an open model might carry
🔓
The LLaMA Answer
  • 7B-65B models trained exclusively on public, documented, filterable sources
  • Token budgets stretched to 1-1.4T — past Chinchilla-optimal, buying inference efficiency
  • State-of-the-art results at every size; 13B beats GPT-3 175B on most benchmarks
  • Weights released to the research community (non-commercial license at first)
Analogy — The Published Recipe

Closed 2022 models were restaurant dishes — delicious, recipe secret. LLaMA published the cookbook: here is the shopping list (public data), here is the oven schedule (token budget), here is how each portion size (7B-65B) turns out. Anyone with a kitchen can now cook — and millions did.

Chapter 02

The Data Is the Moat

Everything public, everything filtered, everything documented — the mixture is the paper's most-copied table.

The public mixture
  • Common Crawl — 67% of the mixture; 5 pipelines of dedup + heuristics (line length, symbol ratios) + low-quality-page classifiers
  • C4 — 15%; Common Crawl cleaned of boilerplate, non-English pages filtered
  • GitHub — 4.5%; permissive licenses only, quality filters, file-level dedup
  • Wikipedia + books + ArXiv — high-density knowledge and reasoning, 11% combined
The architecture deltas
  • RoPE positional encoding instead of learned absolutes
  • SwiGLU activations replacing ReLU/GELU in the FFN
  • RMSNorm instead of LayerNorm — fewer statistics, faster training
  • Efficient attention implementation for long contexts — plus context length 2k
Interactive Demo — The Scale Ladder

Press run to see the parameter ladder LLaMA climbed — and where GPT-3 and PaLM sat. 13B outperforming 175B is the whole thesis in one chart.

Chapter 03

Chinchilla, Deliberately Ignored

The compute-optimal point optimizes one thing; deployment optimizes another.

Why Train Past the Optimum

Chinchilla's law minimizes training compute for a target loss. LLaMA's 7B and 13B consume ~20× the compute-optimal token count — because after deployment, a model is inferred billions of times and trained once. A smaller model at slightly worse loss is dramatically cheaper to serve, distill, and fine-tune. The paper reports the 65B model trained in 21 days on 2,048 A100-80GB GPUs — everything reproducible in principle from public parts.

Interactive Demo — The Public Data Mixture

The mixture behind every open model since. Tab through the components and their cleaning pipelines.

Chapter 05

State of the Art per Size Class

The ladder of four sizes turned one release into an entire ecosystem's substrate.

LLAMA-13B vs GPT-3
outperforms
on most benchmarks — at 13× fewer parameters
LLAMA-65B
competitive
with Chinchilla-70B and PaLM-540B
DATA
100% public
Common Crawl, C4, GitHub, Wikipedia, books, arXiv
RELEASE
weights
research community license — the ecosystem's big bang
Interactive Demo — Fine-Tune on One GPU

How a leaked 7B model became an ecosystem in ten days — the affordability chain LLaMA's efficiency unlocked.

LLaMA7B13B33B65B
Training tokens1.0T1.0T1.4T1.4T
A100 GPU days (paper)~82k~135k~530k~860k
Beats GPT-3 175B—yes — most benchmarksyesyes — vs Chinchilla-70B / PaLM-540B class

Token budgets and device counts from the paper's training table. The 13B 'beats GPT-3' claim is the abstract's headline: state-of-the-art per size class using public data only.

Legacy

Legacy — The Big Bang of Open Weights

A research release with a non-commercial license still restructured the entire industry.

💥 The ecosystem detonation
Leaked weights + Alpaca's $600 data + QLoRA fine-tuning produced a functional open assistant in two weeks — proving open-weights viability at ChatGPT speed.
📜 Llama 2's commercial turn
Meta's next release (entry #15) added a chat-tuned, commercially usable license — the direct descendant that made open models deployable products.
🔬 Research access at scale
Interpretability, safety, quantization, and finetuning research all accelerated once weight access was table stakes for academics.
🧪 The data-recipe publication
Documenting the mixture (CC 67% / C4 15% / code 4.5% / knowledge 11%) gave every later lab a reproducible starting point instead of guesswork.
⚠️ What it did NOT solve
A research-only license, a 2k context window, no chat alignment, no safety training — LLaMA was raw substrate, not a product; and the leak that seeded the ecosystem was, itself, a governance failure.
🛤 Read next
The lineage: Llama 2 · Llama 3 · Chinchilla
Test Yourself

Quick Quiz

Check your understanding of the key concepts from LLaMA.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ 7B-65B models, trained exclusively on public data, state-of-the-art in each size class.
✅ LLaMA-13B outperforms GPT-3 175B on most benchmarks — 13× fewer parameters.
✅ Trained past Chinchilla-optimal on purpose: inference, not training, dominates deployment cost.
✅ The published data mixture and architecture deltas (RoPE, SwiGLU, RMSNorm) became the open-LLM canon.
✅ Research-licensed weights that leaked seeded Alpaca, Vicuna, and the entire open ecosystem.
✅ Read it as the moment 'frontier' stopped meaning 'closed'.