The most complete open-weights engineering report of its era: multilingual, coding, reasoning, and tool-native models — topped by a 405B dense Transformer with a 128K context window, released with pre-trained and post-trained weights.
Llama 3 is a herd: base models, instruct models, safety guardrails, and multimodal — one report for the whole program.
Llama 3's real artifact is not a model but a documented production system: a data engine that mines and deduplicates at 15T+ scale, scaling-law observations in the Chinchilla-revised regime, a long-context extension protocol (progressive RoPE rescaling on 128K data), a rejection-sampling+DPO post-training stack, and honest negative results (distillation that didn't transfer, the 'scaling down' trade-offs). Comparable quality to leading models like GPT-4 on a plethora of tasks — with the recipe attached.
The 2024 question: can a fully published pipeline produce GPT-4-class quality with weights public?
Closed 2024 releases were prize cattle at a county fair — impressive, single, and behind rails. The Llama 3 paper is a working farm: the breeding program (data engine), the feeding schedule (token budget), the training regimen (post-training), the veterinary plan (Guard 3), and the stud book (3.1/3.2/3.3 lineages) — with the animals themselves handed over.
Fifteen trillion tokens is a curation problem, not a download problem.
The recipe behind the instruct herd — and how 8K became 128K without retraining from scratch.
Post-training runs per model: supervised fine-tuning on curated dialogue/reasoning/code data, then rejection sampling (best-of-N SFT rounds) plus direct preference optimization (entry #44) — PPO largely retired from this pipeline. The long-context program grows 8K → 128K in stages, each stage mixing long documents with general data to avoid near-context forgetting. The report closes with the compositional multimodal experiments (image, video, speech added onto the frozen base) and the honest failures — including knowledge-distillation attempts that underperformed data distillation.
The paper's own claim, backed by an extensive empirical evaluation — with weights to check it against.
| Llama 3 milestone | Release | Models | Signature change |
|---|---|---|---|
| Llama 3 | Apr 2024 | 8B · 70B | 15T+ tokens; new tokenizer; strong multilingual base |
| Llama 3.1 | Jul 2024 | 8B · 70B · 405B | 128K context; the herd paper; open frontier claim |
| Llama 3.2 | Sep 2024 | 1B · 3B + multimodal 11B · 90B | Distilled small models; vision added compositionally |
| Llama 3.3 | Dec 2024 | 70B | Flagship-class quality at 70B via post-training alone |
The progression table the community watched live — each row is the same herd paper, extended.
For two years, 'the open frontier' meant Llama 3-class weights and this paper's recipes.
Check your understanding of the key concepts from Llama 3.
Everything you need to remember about this paper.