History Problem Core Idea License Results Impact Quiz Takeaways
Interactive Paper Explainer

Open Weights, Aligned
Llama 2

The model that made 'open' and 'assistant' compatible: 7B-70B pretrained models, then Llama 2-Chat — optimized for dialogue with an RLHF pipeline published in enough detail to reproduce.

Start Learning Read the Paper ↗
7B → 70B
Model sizes
2T
Pre-training tokens
RLHF
Chat pipeline
2023
Commercial license
History

From Weights to Assistant

Llama 1 proved open weights; Llama 2 proved open assistants with a license to deploy.

Feb 2023
LLaMA — raw substrate
Strong base models, research-only license, no chat tuning: the community supplied alignment itself (Alpaca, Vicuna).
Mar 2023
ChatGPT defines the bar
Helpfulness, multi-turn instruction following, refusal behavior — closed models set the assistant standard.
Jul 2023
🚀 Llama 2 + Llama 2-Chat
Meta releases pretrained and RLHF-tuned chat models, 7B-70B, with a commercially usable license and a paper documenting the alignment recipe.
2023-24
The default deployment base
Thousands of products fine-tune Llama 2; the 70B with GQA becomes the reference open model of the era.
2024+
The herd grows
Llama 3 (entry #18) scales the recipe to 405B and 128K context — the lineage this paper stabilized.
What Changed vs Llama 1

Three upgrades: more data (2T tokens vs 1-1.4T, with a heavier data-quality mix), GQA in the 70B (entry #14's trick, native this time), and — the real contribution — a published, reproducible RLHF pipeline: supervised fine-tuning, then iterative rejection-sampling + PPO with separate helpfulness and safety reward models. Llama 2-Chat outperformed open chat models on most tested benchmarks and, on Meta's human evaluations of helpfulness and safety, plausibly substituted for closed-source models.

Chapter 01

Open but Not Aligned

The gap Llama 2 attacked was not capability — it was the assistant layer, and the license to use it.

🧱
The 2023 Open-Model Gap
  • Open base models were strong but raw: no dialogue tuning, no refusals, no multi-turn discipline
  • Community fine-tunes (Alpaca, Vicuna) approximated chat behavior with cheap SFT — inconsistent safety
  • Closed assistants held the helpfulness + safety frontier behind API pricing
  • Research-licensed predecessors (LLaMA 1) could not legally power products
🦙
The Llama 2 Answer
  • 7B/13B/70B pretrained on 2T tokens — more data, better mix, 4K context
  • GQA in the 70B for production-grade inference economics
  • A five-stage alignment pipeline: SFT → helpfulness RM → safety RM → rejection sampling → PPO
  • Commercial-friendly license + safety-mechanism documentation (Llama Guard lineage)
Analogy — The Restaurant Opens

LLaMA 1 was a commissary kitchen — great ingredients, no service, invite-only. Llama 2 is the restaurant: same kitchen craft, but a trained waitstaff (RLHF), a menu people can order from (chat), health inspections (safety RMs + red teaming), and a public door (commercial license).

Chapter 02

The Alignment Pipeline

Five stages, iterated — the recipe that turned a base model into a deployable assistant.

1️⃣ Supervised fine-tuning
Publicly collected instruction data teaches the base model the shape of helpful dialogue — quality over quantity, since RLHF follows.
2️⃣ Reward models ×2
Separate helpfulness and safety RMs trained on human preference rankings — safety needs its own critic, not a helpfulness checkbox.
3️⃣ Rejection sampling
Sample many responses per prompt; keep the RM-preferred best; fine-tune on winners — a cheap, stable stepping stone to PPO.
4️⃣ PPO iterations
Proximal Policy Optimization (entry #38) on RM scores, alternating with fresh rejection sampling rounds — capability and alignment advance together.
Ghost Attention (GA)
  • Multiturn problem: "answer in German" is respected in turn 1, forgotten by turn 5
  • GA fix: fine-tune with the instruction synthesized into every turn of the dialogue
  • Result: constraints persist across the whole conversation — cheap, purely data-side
Training-scale facts
  • 2T tokens for all sizes; context length 4K
  • 70B uses GQA; 7B/13B keep standard MHA
  • Meta's own human preference comparisons score Llama 2-Chat above open peers on helpfulness and safety
  • Safety techniques: safety-centric SFT data, red teaming, safety RLHF, system prompts
Interactive Demo — The RLHF Assembly Line

Follow one model through the full alignment pipeline — from base weights to a chat model that knows when to refuse.

Chapter 03

The License Gift

The technical recipe was reproducible; the legal one was the unlock.

Usable by Companies

The release's quiet masterstroke was commercial usability (with a carve-out for the very largest platforms). Startups that could never license a frontier model could now ship products on a 7B they control — which is exactly what thousands did. The paper pairs the license with unusual recipe transparency: reward-model design choices, annotation guidelines, red-team methodology, and scaling observations — the community's first complete map of an industrial RLHF pipeline.

Interactive Demo — Helpfulness vs Safety — Two Critics

Tab through the same prompts under the two reward models — and see why one RM is never enough.

Chapter 05

The Open-Assistant Benchmark Moment

Where Llama 2-Chat landed: clearly ahead of open peers, plausibly substituting for closed models on Meta's human evals.

OPEN CHAT MODELS
outperforms
on most benchmarks tested
CLOSED MODELS
substitute?
per Meta human evals of helpfulness and safety
TRAINING TOKENS
2T
for every size; 4K context
INFERENCE (70B)
GQA
grouped-query attention native
Interactive Demo — Ghost Attention, Before and After

A three-turn conversation with a system-style constraint. Without GA the rule evaporates; press reveal to see the trained behavior.

Legacy

Legacy — Open Assistants at Product Scale

Llama 2 made 'we host our own aligned model' a normal engineering decision.

🏢 The deployment base
For two years, '7B/13B/70B fine-tuned on our data' was the default architecture of AI products — a thousand startups built on the 70B with GQA.
📘 The RLHF textbook
SFT → rejection sampling → PPO with separate safety/helpfulness RMs became the canonical published pipeline — reproduced, critiqued, and improved (DPO era, entry #44) by the whole field.
🛡 Safety as publication
Red-teaming methodology, over-refusal tracking, and Llama Guard made 'document your safety stack' part of the open-model genre.
⚖️ The license precedent
Commercially usable weights with a mega-platform carve-out became the template every open lab copied (Llama 3, Mistral variants, Gemma).
⚠️ What it did NOT solve
4K context, no tool-use training, English-centric data, RLHF's reward-hacking exposure — and human evals run by the model's own authors deserve the standard skepticism.
🛤 Read next
The lineage and the toolbox: Llama 3 · PPO · GQA
Test Yourself

Quick Quiz

Check your understanding of the key concepts from Llama 2.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ 7B-70B pretrained (2T tokens) + Llama 2-Chat RLHF-tuned — outperforms open chat models on most benchmarks tested.
✅ Alignment pipeline: SFT → helpfulness RM + safety RM → iterated rejection sampling + PPO.
✅ Separate safety and helpfulness reward models make the trade-off explicit instead of accidental.
✅ Ghost Attention: re-inject constraints each turn in training data — multi-turn persistence learned, not hoped for.
✅ 70B ships GQA; commercial license made open assistants deployable at product scale.
✅ Read it as the RLHF recipe book the field then spent two years simplifying (DPO) and scaling (Llama 3).