The model that made 'open' and 'assistant' compatible: 7B-70B pretrained models, then Llama 2-Chat — optimized for dialogue with an RLHF pipeline published in enough detail to reproduce.
Llama 1 proved open weights; Llama 2 proved open assistants with a license to deploy.
Three upgrades: more data (2T tokens vs 1-1.4T, with a heavier data-quality mix), GQA in the 70B (entry #14's trick, native this time), and — the real contribution — a published, reproducible RLHF pipeline: supervised fine-tuning, then iterative rejection-sampling + PPO with separate helpfulness and safety reward models. Llama 2-Chat outperformed open chat models on most tested benchmarks and, on Meta's human evaluations of helpfulness and safety, plausibly substituted for closed-source models.
The gap Llama 2 attacked was not capability — it was the assistant layer, and the license to use it.
LLaMA 1 was a commissary kitchen — great ingredients, no service, invite-only. Llama 2 is the restaurant: same kitchen craft, but a trained waitstaff (RLHF), a menu people can order from (chat), health inspections (safety RMs + red teaming), and a public door (commercial license).
Five stages, iterated — the recipe that turned a base model into a deployable assistant.
The technical recipe was reproducible; the legal one was the unlock.
The release's quiet masterstroke was commercial usability (with a carve-out for the very largest platforms). Startups that could never license a frontier model could now ship products on a 7B they control — which is exactly what thousands did. The paper pairs the license with unusual recipe transparency: reward-model design choices, annotation guidelines, red-team methodology, and scaling observations — the community's first complete map of an industrial RLHF pipeline.
Where Llama 2-Chat landed: clearly ahead of open peers, plausibly substituting for closed models on Meta's human evals.
Llama 2 made 'we host our own aligned model' a normal engineering decision.
Check your understanding of the key concepts from Llama 2.
Everything you need to remember about this paper.