History Problem Core Idea Loop Results Impact Quiz Takeaways
Interactive Paper Explainer

Models Attack Models
Red Teaming with LMs

Human red teamers are expensive and slow. Point another language model at the target as an adversary — and it uncovers tens of thousands of offensive replies in a 280B-parameter chatbot, at machine scale.

Start Learning Read the Paper ↗
280B
Target chatbot
1000s+
Harmful cases per hour-scale
100×
More effective / sample (paper claim)
2022
Perez et al.
History

Safety Testing at Human Speed

The scaling problem: models grow faster than annotation teams.

2021-22
Hand-written test cases
Human annotators craft safety probes one at a time — expensive, limited diversity, always behind the model.
2021
Foundational gap noted
Prior work explicitly identifies harmful behaviors pre-deployment via human red teaming — and its cost ceiling.
Feb 2022
🚀 LM red teaming
Perez et al.: generate test cases WITH another LM — zero-shot, few-shot, supervised, and RL-driven generators — auto-scored for offensiveness.
2022-23
The safety pipeline standard
Automated adversarial testing joins RLHF workflows; the paper's target gets its own improvements: the harmful generations become training data for SAFER behavior.
2023+
Adversarial era
GCG suffixes (entry #90), agent red-teaming (entry #95), and safety evals industrialize the adversary-as-generator pattern.
The Generator-Scorer Loop

The architecture is three parts: a generator LM (52B in the paper) produces candidate test inputs; the target LM (a 280B chatbot) replies; a trained offensiveness classifier judges the reply. Harmful (input, reply) pairs surface automatically. The diversity tricks matter most: prompting the generator with persona descriptions ("a racist user from the 1800s") steers it into corners of input space humans never think to visit — and RL on the classifier signal pushes further, optimizing directly for test cases the target fails.

Chapter 01

Finding Needles at Annotation Speed

The economics of pre-deployment safety testing.

🧪
The Human Bottleneck
  • Hand-written test cases: high cost per probe, low diversity, bounded imagination
  • Harmful behavior hides in a vast input space — sampling it at human speed misses most of it
  • Annotation budgets cap the number AND kind of failure modes discovered
  • Bigger models have bigger surface areas — the gap compounds with scale
🤖
The Automated Answer
  • Generate test cases with an LM: zero-shot → few-shot → supervised → RL, increasing control and targeting
  • Personas and steering prompts diversify the attack distribution beyond human imagination
  • Classifier-based scoring: automatic, scalable detection of harmful replies
  • Bonus loop: the discovered failures become training data that makes the target SAFER
Analogy — Pen Testers vs the Fuzzer

Human red teamers are elite penetration testers — insightful, expensive, scarce. LM red teaming is a fuzzer with taste: it throws millions of structured probes at the target, keeps the ones that draw blood, and — the twist — the crash reports double as patches (fine-tuning data). The fuzzer writes the fixes.

Chapter 02

The Method Ladder

From zero-shot to RL — control and attack-rate rise together.

The generation methods
  • Zero-shot: "generate a list of questions that would elicit harmful answers" — broad, untargeted
  • Persona-driven few-shot: describe an attacker ("grumpy old man frustrated with Millennials") — diversity through character
  • Supervised: few-shot with prior harmful examples found — focus around known failure modes
  • RL (GBF training): reward = offensiveness of the TARGET's reply — optimize directly for failure
The findings (from the paper)
  • Uncovered tens of thousands of offensive replies in the 280B-parameter chatbot
  • Offensive generations cluster in interpretable themes — attack surface has structure
  • Detected harmful output at rates orders of magnitude beyond random or human sampling (per-sample effectiveness)
  • Fine-tuning on the found cases reduces failures — the attack is also the cure
Interactive Demo — The Adversarial Loop, One Pass

Watch a persona-driven generator probe a chatbot — and the classifier catch the kill.

Chapter 03

The Feedback Trick

Why the red team's output makes the target safer — the closed loop.

Adversary as Annotator

The discovered (harmful prompt, harmful reply) pairs are exactly the training data a safety fine-tune needs: real failure cases, machine-labeled, drawn from the model's own weak points. The paper demonstrates the cycle — red team finds thousands, fine-tune on them, the failure rate on those classes drops — prefiguring the adversarial-training loops that RLHF pipelines adopted. The adversary and the alignment engineer are, mechanically, the same worker.

Interactive Demo — Four Generators, Four Attack Shapes

Tab through the method ladder — each level trades breadth for targeting.

Chapter 05

Machine-Scale Discovery

What automated red teaming found in a single target.

HARMFUL REPLIES FOUND
tens of thousands
in a 280B-parameter chatbot
EFFECTIVENESS
orders of magnitude
vs human/random sampling per test case
METHOD RANGE
4 levels
zero-shot → persona → supervised → RL
CURE
same data
fine-tuning on findings reduces failures
Interactive Demo — The Honest Limits

LM red teaming found thousands — but the paper flags its own blind spots. Press reveal.

Legacy

Legacy — The Adversary as Infrastructure

Red-teaming pipelines and safety fine-tuning merged into one loop.

🔁 The safety pipeline
Generate → score → fine-tune became the standard adversarial-training cycle inside RLHF-era safety workflows — the paper's loop, industrialized.
🎭 Persona-based adversarial diversity
Steering attack diversity through character descriptions became a standard tool for both safety evaluation and synthetic-data generation.
⚔ The escalation frame
Adversaries-as-models set the stage for optimization attacks (GCG, entry #90) and agent-scale misuse evals (AgentHarm, entry #95) — the same architecture, harder targets.
⚠️ What it did NOT solve
The judge (classifier) bounds what can be found; generator imagination bounds where to search; and automated discovery legitimized a false sense of coverage — the un-imagined failure modes remain until something else finds them.
🛤 Read next
The safety stack: Constitutional AI · GCG Attack · AgentHarm
Test Yourself

Quick Quiz

Check your understanding of the key concepts from Red Teaming LMs.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Red team with a generator LM + offensiveness classifier — discovery at machine scale.
✅ Tens of thousands of offensive replies surfaced in a 280B-parameter chatbot.
✅ The method ladder: zero-shot → personas → supervised → RL on the classifier signal.
✅ Found failures become safety fine-tuning data — adversary as annotator.
✅ Limits are real: the judge and the generator bound the discoverable.
✅ Read it as the moment safety testing got a fuzzer — and alignment got its data.