History Problem Recipe Data Architecture Results Impact Quiz
Interactive Paper Explainer

Open and Efficient
LLaMA

Smaller models, more tokens, open weights. LLaMA-65B rivals the best closed systems of its day, and LLaMA-13B beats GPT-3 (175B) on most benchmarks — trained on 1.4T tokens of entirely public data. This is the paper that detonated the open-model ecosystem.

Start Learning Read the Paper ↗
65B
Largest Model
1.4T
Training Tokens (Public Only)
4
Model Sizes (7/13/33/65B)
2023
Year Published
History

The Road to Open Weights

LLaMA landed at the peak of the "bigger is better — and closed" era. Here is the road that led to it, and the explosion that followed.

2020
GPT-3 (Brown et al.)
175B parameters and few-shot magic — and fully closed. You could query an API, but never see the weights.
2021–22
PaLM, BLOOM, OPT
Frontier quality stays locked behind APIs (PaLM 540B). The few open 175B-class releases — BLOOM, OPT — underperform GPT-3 and ship research-only licenses.
2022
Chinchilla (Hoffmann et al.)
The compute-optimal rule: ~20 tokens per parameter. Same compute, 4× smaller model, 4× more data → better. See the Chinchilla guide.
2023 · Feb
🚀 LLaMA (Touvron et al., Meta AI)
7/13/33/65B models trained on 1.4T tokens of 100% public data. Weights released to researchers under a gated license.
2023 · Mar →
The explosion
The weights leak within a week. Alpaca and Vicuna fine-tunes appear; llama.cpp runs a quantized 7B on a laptop.
2023 →
Llama 2 & 3 — the open era
A truly open license (commercial use allowed), chat-tuned models, and open weights as the ecosystem default.
Key Insight

Chinchilla asked: given a training budget, what model is best? LLaMA asks the deployment question: given a model you will serve millions of times, what should you have trained? Once inference dominates the cost, the optimum moves to smaller models trained longer.

THE SAME 1.4T TOKENS, TWO ALLOCATIONS
70B params × 20 tokens — cheapest to train
7B params × 200 tokens — cheapest to serve
LLaMA trains both ends of this spectrum — and every point in between.
Chapter 01

The Problem — a Closed Frontier

By 2022, the best language models were black boxes. LLaMA's bet: you don't need 175B parameters — or a private dataset — to reach that class of performance.

🔒
The Closed Frontier
  • GPT-3, PaLM, Chinchilla: best quality, zero public weights
  • No weights → no science: probing, fine-tuning, and bias analysis are impossible
  • A reproducibility crisis: papers benchmark APIs, not models
  • The few open 175B-class models (BLOOM, OPT) underperformed GPT-3
  • Big models are expensive to serve — 175B-class inference needs a datacenter
🦙
LLaMA's Bet
  • Smaller models trained longer: 7–65B on 1.4T public tokens
  • Strong per-parameter efficiency — 13B beats GPT-3 on most benchmarks
  • 100% public data: no proprietary scrape, fully inspectable recipe
  • Weights available to researchers under a gated license
  • One 65B run: ~2,048 A100-80GB GPUs for ~3 weeks — commodity scale
The Question the Paper Asks

Can we match GPT-3-class performance with smaller models trained longer on more public data? Think of closed frontier models as restaurants: you can order from the menu (an API), but you can never walk into the kitchen. BLOOM and OPT handed out photos of a huge kitchen that cooked worse. LLaMA hands researchers the keys to a compact kitchen that plates nearly the same dishes — and lets you renovate it.

THE BET, IN ONE LINE
13B params + 1.4T tokens > 175B params + 300B tokens
Chapter 02

The Recipe — Train Longer Than Optimal

Chinchilla's rule (~20 tokens/param) minimizes training compute. LLaMA borrows the spirit and breaks the letter at the small end: overtraining buys cheaper inference, forever.

tokens per param = 1.4T ÷ N
7B → 200 · 13B → 108 · 33B → 42 · 65B → 22
20
Chinchilla rule
The compute-optimal ratio: ~20 tokens per parameter for the cheapest training run.
22
LLaMA-65B
Sits right on the rule — the maximum-quality end of the family.
42
LLaMA-33B
~2× beyond optimal — near-frontier quality, cheaper to serve.
200
LLaMA-7B
10× the rule — deliberately overtrained so it runs on a single GPU.
One Corpus, Four Sizes
ModelParamsTokensTokens / paramStrategy
LLaMA-7B7B1.4T20010× beyond optimal — built for cheap serving
LLaMA-13B13B1.4T108The GPT-3 beater
LLaMA-33B33B1.4T42~2× optimal — quality vs serving trade-off
LLaMA-65B65B1.4T22≈ Chinchilla-optimal — maximum quality

Every size sees the identical 1.4T-token corpus — the family doubles as a clean study of "more tokens, fewer params."

Interactive Demo — The Tokens-per-Param Slider

Fix the corpus at 1.4T tokens and slide the tokens-per-parameter ratio. Watch the implied model size shrink while the green Chinchilla line (~20/param) stays put — that gap is inference savings banked on every request.

Why break the rule?

Chinchilla minimizes training compute — a cost you pay once. A deployed model pays inference on every request, forever, and inference cost scales with parameter count. A 7B model overtrained to 200 tokens/param is wasteful to train once and cheap to serve a billion times.

The 1.4T decision

At 1.4T tokens, even the smallest model (7B) sees a huge, deduplicated, diverse corpus. One corpus, four training budgets: the smaller sizes simply trade peak quality for serving cost — and all four publish as a clean scaling family.

Chapter 03

The Data — 1.4T Public Tokens

No private scrapes, no licensed archives: every one of the 1.4 trillion tokens is publicly available, so anyone can audit — and in principle rebuild — the corpus.

Interactive Demo — The Data-Mix Explorer

Click a source row to see what it contributes. Two thirds is filtered web crawl; the rest is a carefully chosen blend of code, encyclopedic text, books, math, and Q&A.

Why this blend?

The mix mirrors what a general-purpose model needs: web prose for breadth, code for programming, Wikipedia for facts, books for long-horizon reasoning, ArXiv for math, StackExchange for dialogue-shaped Q&A. Web text dominates by volume; the curated sources dominate by signal per token.

Public = accountable

Because every source is public, the recipe is inspectable and reproducible in principle — a deliberate contrast with closed frontier training sets. It also kept the release story simple: no proprietary data to license around, weights to share with researchers.

Chapter 04

Architecture — a Vanilla Decoder, Three Swaps

No exotic machinery: a standard GPT-style decoder-only Transformer. Three targeted upgrades over the GPT-3 recipe, each borrowed from the 2021–22 literature.

Architecture Overview
🧱
Decoder-Only
Causal self-attention, next-token prediction.
📏
2048 Tokens
The context window for prompts + generations.
✂️
SentencePiece BPE
Subword tokenizer, ~32k vocabulary.
🔁
4 Sizes
7B / 13B / 33B / 65B — same recipe, scaled.
⚡ Swap 1 — RMSNorm (pre-norm)
Replaces LayerNorm: drop the mean-subtraction, keep root-mean-square scaling. Why: simpler and faster — and pre-norm keeps very deep stacks stable.
🌀 Swap 2 — SwiGLU activation
Replaces GELU in the FFN with a gated unit — two branches, one gates the other. Why: better quality per FLOP, the same upgrade PaLM used.
🧭 Swap 3 — Rotary embeddings (RoPE)
Replaces learned absolute positions: rotate query/key vectors so attention depends on relative distance. Why: no position table, and it extrapolates naturally.
LLAMA-7B
7B
parameters
32 layers · 4096 hidden · 32 heads
LLAMA-13B
13B
parameters
40 layers · 5120 hidden · 40 heads
LLAMA-33B
33B
parameters
52 layers · 6656 hidden · 52 heads
LLAMA-65B
65B
parameters
80 layers · 8192 hidden · 64 heads
GPT-3 (2020) vs LLaMA (2023)
ComponentGPT-3LLaMA
NormalizationLayerNormRMSNorm, pre-norm
FFN activationGELUSwiGLU (gated)
PositionsLearned absoluteRoPE (rotary)
TokenizerBPE, ~50kSentencePiece BPE, ~32k
Context length20482048 (unchanged)
ObjectiveNext-token predictionSame — nothing new needed

The lesson: LLaMA's edge is not architectural — it's data and training. The best 2023 decoder is barely three components away from GPT-3.

Chapter 05

Training & Results — Frontier on a Budget

One GPU cluster, three weeks, public data — and results that made closed models 10× larger look inefficient.

🖥️ 2,048 A100-80GB
The cluster behind the 65B run — a single, commodity-scale training job.
📅 ~21 days
Wall-clock time for the 65B model to consume all 1.4T tokens.
⚡ ~1 day
The same 1.4T tokens through the 7B model on the same 2,048-GPU cluster.
🔁 One recipe
Identical data, architecture, and training procedure across all four sizes.
13B vs 175B
beats
LLaMA-13B outperforms GPT-3 on most benchmarks
65B
rivals
the frontier: competitive with Chinchilla-70B & PaLM-540B
CORPUS
1.4T
tokens — 100% publicly available data
65B TRAINING RUN
21
days on 2,048 A100-80GB GPUs
MMLU 5-Shot — the Headline Table
SystemParamsMMLU 5-shotNotes
GPT-3 (closed)175B43.9API only; ~300B training tokens
LLaMA-13B (open)13B~55Beats GPT-3 with ~13× fewer parameters
LLaMA-65B (open)65B63.4Public data only; rivals the frontier
Chinchilla (closed)70B67.5DeepMind; the compute-optimal reference
PaLM (closed)540B~698× larger than LLaMA-65B

Beyond MMLU: strong zero-shot common-sense reasoning (CommonsenseQA, PIQA, HellaSwag, ARC, BoolQ), and the 65B model stays competitive with Chinchilla-70B and PaLM-540B across many evaluations.

Interactive Demo — MMLU 5-Shot Showdown

Toggle between raw scores and the gain over GPT-3 — then notice how little the blue and amber bars cost to serve.

Legacy

Impact — the Explosion

A gated research release became the accidental ignition point of the entire open-weights ecosystem.

📦 The gated release (Feb 2023)
Weights to researchers under a non-commercial license. They leaked within a week — and Meta leaned into the flood instead of fighting it.
🦙 Alpaca (Mar 2023)
LLaMA-7B + 52k instructions, fine-tuned for a few hundred dollars — a GPT-3-class assistant on a laptop-friendly base model.
🐫 Vicuna (Mar 2023)
Fine-tuned on shared ChatGPT-style conversations; reported ~90% of ChatGPT quality — and helped popularize LLM-as-a-judge evaluation.
⚙️ llama.cpp + quantization
4-bit quantized LLaMA-7B running on consumer laptops and gaming GPUs — no datacenter required to serve a foundation model.
🦁 Llama 2 (2023) & Llama 3 (2024)
A genuinely open license — commercial use allowed — plus chat-tuned variants. The open-weights era went mainstream.
🌍 The ecosystem
Thousands of fine-tunes, open eval harnesses, serving and quantization tooling — open weights became the research default.
Test Yourself

Quick Quiz

Check your understanding of the key concepts from the LLaMA paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ 1.4T training tokens, 100% public data — no proprietary scrapes, an inspectable recipe.
✅ Four sizes (7/13/33/65B), one corpus: the small ones deliberately overtrained for cheap serving.
✅ Decoder-only Transformer with three swaps vs GPT-3: RMSNorm (pre-norm), SwiGLU, RoPE.
✅ LLaMA-13B beats GPT-3 (175B) on most benchmarks; LLaMA-65B rivals Chinchilla-70B and PaLM-540B.
✅ 65B training: ~2,048 A100-80GB GPUs, ~21 days — frontier class at commodity scale.
✅ The release detonated the open ecosystem: Alpaca, Vicuna, llama.cpp, Llama 2/3.