History Problem Core Idea Self-Critique RLAIF Results Impact Deep Dive Quiz
Interactive Paper Explainer

Harmlessness from AI Feedback
Constitutional AI

A visual, step-by-step guide to the paper that replaced tens of thousands of human harm-preference labels with a written constitution and the model's own critiques — RLAIF, the alignment recipe behind Claude.

Start Learning Read the Paper ↗
52B
Model Scale Tested
16
Constitution Principles
2
Phases (SL + RL)
2022
Year Published
History

From Labelers to Principles

RLHF works — but it outsources your values to whoever you hired. CAI writes the values down.

2017 →
RLHF lineage (Christiano, Stiennon, Ouyang)
Human preference labels + reinforcement learning align models with what people rate as better.
2022 · Apr
Helpful & Harmless assistants (Bai et al.)
Anthropic's RLHF assistant line documents the harmlessness–helpfulness tension quantitatively.
2022 · Dec
🚀 Constitutional AI (Bai et al.)
A written list of principles + model-generated critiques and revisions + AI preference labels: RLAIF.
2023 →
Claude, RLAIF everywhere
The technique scales alignment: labeled data for harmfulness becomes optional, and transparency becomes a design axis.
The Problem With Preference Labels

RLHF's harm labels come from crowdworkers with their own (varying) judgments. That is expensive, inconsistent, slow to iterate on, and — most importantly — opaque: no one can read why a model learned where its boundaries are. Constitutional AI makes the boundary-setting step a public document, then lets models generate the training signal from it.

🧭 Lineage
Compare with InstructGPT — same RL machinery, human-sourced labels instead.
Chapter 01

The Scaling Problem of RLHF

RLHF needs tens of thousands of human preference labels. As models grow, that human pipeline becomes the bottleneck — and the hidden values problem becomes worse.

🏷
Human Labels Don't Scale or Explain
  • Every boundary judgment needs a fresh batch of human comparisons
  • Labelers disagree with each other — and with the eventual users
  • Values are implicit in the data; nobody can audit why the model refuses what it refuses
  • RLHF tends to produce evasive, wishy-washy refusals to avoid disagreement
📜
A Constitution Scales With the Model
  • Write the principles down once — 16 of them in this paper
  • The model itself generates critiques, revisions, and preference labels from those principles
  • Human labor shrinks to writing and iterating the constitution
  • Boundaries become auditable text, not statistical folklore
Analogy — Case Law vs Written Law

RLHF is case law: behavior emerges from thousands of individual verdicts nobody wrote down. Constitutional AI is a code of written law: principles first, judgments derived from them, and any citizen (or auditor) can read the statute book.

Chapter 02

Two Phases, One Document

Supervised learning from self-critique, then reinforcement learning from AI feedback — both driven by the same constitution.

Phase 1 · SL
Sample responses to harmful prompts → model critiques its own output against a principle → rewrites it → finetune on the revisions.
Phase 2 · RL
From the finetuned model, sample response pairs; an AI labels preferences by asking which response better follows a (random) principle; train a preference model; run RL against it.
📜 The constitution
16 principles spanning harmlessness, helpfulness within safety, honesty, and process values like impartiality — each with an operational "how to choose" instruction.
🤖 52B scale
Experiments through 52B parameters; harmlessness improves without collapsing helpfulness.
Chapter 03

The Self-Critique Loop

Phase 1 in slow motion: a harmful response gets critiqued, then rewritten — all by the model, prompted with one principle.

Interactive Demo — Critique & Revise Stepper

Step through the supervised phase on one prompt. Each click advances: initial response → constitutional critique → revision.

Why Critique-Then-Revise Works

Asking a model to rewrite a bad answer directly produces shallow edits. Asking it to critique first — "what is wrong with this response, per principle X?" — surfaces the failure mode explicitly, and the revision then has something concrete to fix. The paper found critique-with-a-specific-principle beats generic "improve this."

Non-Evasive Refusals

A headline behavioral change: CAI models refuse less evasively. Instead of "I can't help with that," the trained model explains the objection and offers safe alternatives — because the constitution's principles ask it to be helpful within harmlessness, and to explain its objections.

Chapter 04

RL From AI Feedback

Phase 2 replaces the human preference labeler with a model consulting the constitution.

reward ← PM( response pair + randomly sampled principle )
Pair sampling
Two responses
From the phase-1 model, sample two responses to the same prompt.
Principle
Random principle
A random constitution principle + its instruction is inserted into the preference prompt.
AI label
Model judgment
The model says which response better follows that principle.
PM → RL
Preference model
Train a PM on these AI labels, then optimize the policy against it.
Interactive Demo — AI Preference Labeler

Given a harmful-ish prompt and two candidate replies, pick a principle and see which the AI judge prefers under it.

🎲 Why random principles
Sampling a different principle per comparison prevents the PM from over-fitting one axis — an implicit ensemble over the constitution.
🔗 Decoupled preference labeling
The paper also shows the AI labeler can be decoupled from the policy — a general RLAIF pattern later papers reused.
🔍 Transparency audit
When the label disagrees with intuition, you can trace it to a principle — debuggable alignment data.
Chapter 05

Does It Actually Work?

The paper's evaluation: crowdworkers comparing CAI against RLHF baselines on helpfulness and harmlessness, plus a 438-question HHH test.

Harmlessness
  • RL-CAI (constitutional + RL) substantially improves harmlessness over the RLHF baseline at the 52B scale
  • Notably: non-evasive refusals — models explain objections and offer alternatives
  • Feedback quality: AI preference labels, prompt-by-principle, produce PMs close enough to human-labeled ones to train from
The Honest Trade-offs
  • Helpfulness dips somewhat versus pure helpfulness-RLHF — the classic HHH trade-off, made visible
  • Constitution quality is the new human-in-the-loop: bad principles → confidently misaligned behavior
  • SL-only CAI improves harmlessness less than the full SL+RL pipeline — RL phase matters
Interactive Demo — Harmlessness vs Helpfulness Frontier

Where each recipe lands on the two axes. Click a recipe to move the frontier dot.

Legacy

Impact — Alignment at Machine Speed

RLAIF made preference data cheap enough to iterate on values daily instead of quarterly.

🧬 The Claude lineage
Anthropic's assistant line grew directly from this recipe — constitutions evolved in public for later models.
💸 Preference-label economics
If AI labels ≈ human labels for many comparisons, alignment iteration cost collapses — an industry-level shift.
🧭 DPO and label-free alignment
The same pressure (less RL machinery, cheaper signals) drives Direct Preference Optimization.
👁 Scalable oversight research
CAI is a founding example of weaker-supervisors-stronger-models: the assistant critiques itself toward principles humans can read.
⚖️ Constitutional criticism
Who writes the constitution? Whose values? The paper opens governance questions it cannot close — deliberately.
⚠️ What it did NOT solve
Sycophancy, sandbagging, and values-drift-under-RL remain open; a constitution constrains legible behavior, not latent capabilities.
Deep Dive

The Constitution Is the Real Product

The training technique is transferable. The values in the document are where every downstream difference actually originates.

🧠
The Mechanism Illusion
  • "RLAIF" sounds like the algorithm is the contribution
  • But any constitution trains a model — the output behavior tracks the principles
  • Two RLAIF models with different constitutions are different products, same math
  • Principles written informally can conflict, and the PM silently resolves the conflicts
📖
The Honest Reading
  • The paper's 16 principles were "ad hoc," research-purpose — stated openly
  • That honesty is the point: constitution-writing is now visibly a first-class engineering act
  • Iteration moved from data collection to document editing
  • Public principles = public accountability surface, whatever their flaws
Interactive Demo — Constitution Sandbox

Same borderline request, three different principle weightings — watch the model's answer change with the document, not the algorithm.

Verdict

Constitutional AI's real invention is a separation of powers: the legislature (principles) is now a readable document, the judiciary (preference labels) is delegated to machines, and the executive (RL) is unchanged machinery. Whether that legislature is accountable — who votes, who amends — is the alignment debate this paper started and no training run can finish.

Test Yourself

Quick Quiz

Check your understanding of the key concepts from the Constitutional AI paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Two phases: supervised self-critique (critique → revise → finetune) then RL from AI feedback.
✅ The constitution: 16 principles, deliberately ad hoc and research-purpose — the auditable values layer.
✅ RLAIF replaces human harm-preference labels with model judgments prompted by random principles.
✅ Harmlessness improves at 52B scale without collapsing helpfulness — and refusals become non-evasive.
✅ Critique-first prompting beats direct revision: surface the failure before fixing it.
✅ The mechanism is transferable; the constitution is the product — values live in the document.