History Problem Core Idea Post-Training Results Impact Quiz Takeaways
Interactive Paper Explainer

The Herd Report
Llama 3

The most complete open-weights engineering report of its era: multilingual, coding, reasoning, and tool-native models — topped by a 405B dense Transformer with a 128K context window, released with pre-trained and post-trained weights.

Start Learning Read the Paper ↗
405B
Flagship dense model
128K
Context window
15T+
Pre-training tokens
2024-25
The 3.x progression
History

One Paper, Many Models

Llama 3 is a herd: base models, instruct models, safety guardrails, and multimodal — one report for the whole program.

Feb 2023
LLaMA — the substrate
Public-data base models prove open viability (entry #12).
Jul 2023
Llama 2 — the assistant
RLHF-tuned chat models + commercial license (entry #15).
Apr 2024
Llama 3 8B & 70B
First wave: 15T+ tokens, strong post-training, major multilingual gains; the 70B closes on closed-source mid-tier.
Jul 2024
🚀 The herd paper
The full report: 405B flagship (dense, 128K context), quality comparable to GPT-4-class on many tasks, plus Llama Guard 3 and compositional image/video/speech experiments.
2024-25
3.1 / 3.2 / 3.3
The herd grows: 128K contexts fleet-wide, distilled small models, quantized and multimodal variants — one paper's program, extended.
The Engineering Report Genre

Llama 3's real artifact is not a model but a documented production system: a data engine that mines and deduplicates at 15T+ scale, scaling-law observations in the Chinchilla-revised regime, a long-context extension protocol (progressive RoPE rescaling on 128K data), a rejection-sampling+DPO post-training stack, and honest negative results (distillation that didn't transfer, the 'scaling down' trade-offs). Comparable quality to leading models like GPT-4 on a plethora of tasks — with the recipe attached.

Chapter 01

Open Quality at Frontier Class

The 2024 question: can a fully published pipeline produce GPT-4-class quality with weights public?

🏔
The Frontier Gap
  • Closed models led on multilinguality, long context, reasoning, and tool use simultaneously
  • Open recipes published bits and pieces — data engine, post-training, evaluation — never one whole system
  • 405B-class training with reproducible curation was unproven in public
  • Safety tooling (guard models) shipped separately from capability releases
🐑
The Herd Answer
  • 405B dense Transformer, 128K context, GQA — trained on 15T+ tokens from a fully documented data engine
  • Post-training: SFT, rejection sampling, and DPO across dialogue/safety/reasoning/code domains
  • Native multilinguality, coding, reasoning, and tool usage from the base model up
  • Llama Guard 3 + compositional multimodal experiments in the same report — the program, not just the model
Analogy — The Farm, Not the Animal

Closed 2024 releases were prize cattle at a county fair — impressive, single, and behind rails. The Llama 3 paper is a working farm: the breeding program (data engine), the feeding schedule (token budget), the training regimen (post-training), the veterinary plan (Guard 3), and the stud book (3.1/3.2/3.3 lineages) — with the animals themselves handed over.

Chapter 02

The Data Engine

Fifteen trillion tokens is a curation problem, not a download problem.

Pipeline stages
  • Heuristic filters — language ID, length, repetition, toxicity screens
  • Model-quality classifiers — fast trained scorers decide page-level keep/drop
  • Massive dedup — URL-, document-, and fuzzy-dedup across the whole corpus
  • Domain reweighting — deliberate upsampling of code, math, multilingual, and knowledge-heavy domains
Architecture & scale facts
  • 405B dense Transformer — deliberately dense (not MoE) for training simplicity and inference quality
  • 128K context via progressive long-context extension stages (RoPE rescaling on curated long documents)
  • GQA on all sizes; tokenizer with 128K vocabulary (upsampled from prior 32K)
  • Scaling-law studies across 8B→70B→405B runs to extrapolate flagship training decisions
Interactive Demo — From Web to 15T Tokens

Walk one page through the data engine — the part of the paper every later lab copied.

Chapter 03

Post-Training and the Extension Program

The recipe behind the instruct herd — and how 8K became 128K without retraining from scratch.

From Base to Herd

Post-training runs per model: supervised fine-tuning on curated dialogue/reasoning/code data, then rejection sampling (best-of-N SFT rounds) plus direct preference optimization (entry #44) — PPO largely retired from this pipeline. The long-context program grows 8K → 128K in stages, each stage mixing long documents with general data to avoid near-context forgetting. The report closes with the compositional multimodal experiments (image, video, speech added onto the frozen base) and the honest failures — including knowledge-distillation attempts that underperformed data distillation.

Interactive Demo — 8K → 128K Context Extension

Slide through the long-context growth stages — each stage re-trains on longer documents mixed with general data so nothing is forgotten.

Stage names simplified; the principle — progressive rescaling with curriculum mixing and short-context regression checks — is the paper's.
Chapter 05

Comparable to Leading Models

The paper's own claim, backed by an extensive empirical evaluation — with weights to check it against.

405B DENSE
128K ctx
the largest fully-documented open Transformer of its day
QUALITY
GPT-4-class
comparable on a plethora of tasks, per the paper's evaluation
DATA ENGINE
15T+ tokens
fully documented curation, dedup, and reweighting
SAFETY
Guard 3
input/output safety model released alongside the herd
Interactive Demo — The Herd Family Tree

What actually shipped, per release wave. Tab through the program the paper documents.

Llama 3 milestoneReleaseModelsSignature change
Llama 3Apr 20248B · 70B15T+ tokens; new tokenizer; strong multilingual base
Llama 3.1Jul 20248B · 70B · 405B128K context; the herd paper; open frontier claim
Llama 3.2Sep 20241B · 3B + multimodal 11B · 90BDistilled small models; vision added compositionally
Llama 3.3Dec 202470BFlagship-class quality at 70B via post-training alone

The progression table the community watched live — each row is the same herd paper, extended.

Legacy

Legacy — The Reference Release

For two years, 'the open frontier' meant Llama 3-class weights and this paper's recipes.

🌍 Open frontier-quality weights
A fully documented 405B/128K model comparable to leading closed models on many tasks reset the open-closed boundary — every open lab benchmarked against it.
🏭 The data-engine genre
The paper's curation pipeline became the public template for 15T-scale data work — filters, classifiers, dedup, and reweighting as first-class engineering.
🧪 Post-training without PPO
The herd's rejection-sampling + DPO stack mainstreamed PPO-free alignment at frontier scale (DPO, entry #44, going industrial).
📐 The negative results
Failed distillation branches and documented trade-offs taught the field as much as the wins — the report is honest where marketing would not be.
⚠️ What it did NOT solve
Multilinguality remains uneven across languages; long-context evaluation was young (Lost in the Middle, entry #33); tool use was trained but agent-level reliability lagged (the VIII category); and the herd paper evaluates its own models — the standard caveat applies.
🛤 Read next
The context problem it inherited: Lost in the Middle · the alignment stack: DPO · the efficiency frontier: DeepSeek-V3
Test Yourself

Quick Quiz

Check your understanding of the key concepts from Llama 3.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ A herd, not a model: 8B-405B base + instruct models, Guard 3, and multimodal experiments in one report.
✅ Flagship: 405B dense, 128K context, GQA, 15T+ curated tokens — quality comparable to leading models like GPT-4 on many tasks.
✅ The data engine (filters → quality classifiers → dedup → reweighting) is the paper's most-copied export.
✅ Long context grew 8K → 128K via staged rescaling with curriculum mixing and regression checks.
✅ Post-training went PPO-free: SFT + rejection sampling + DPO at frontier scale.
✅ The 3.1/3.2/3.3 progression turned one release into a program — distilled, quantized, multimodal, and 70B-flagship variants.