A visual, step-by-step guide to the paper that replaced tens of thousands of human harm-preference labels with a written constitution and the model's own critiques — RLAIF, the alignment recipe behind Claude.
RLHF works — but it outsources your values to whoever you hired. CAI writes the values down.
RLHF's harm labels come from crowdworkers with their own (varying) judgments. That is expensive, inconsistent, slow to iterate on, and — most importantly — opaque: no one can read why a model learned where its boundaries are. Constitutional AI makes the boundary-setting step a public document, then lets models generate the training signal from it.
RLHF needs tens of thousands of human preference labels. As models grow, that human pipeline becomes the bottleneck — and the hidden values problem becomes worse.
RLHF is case law: behavior emerges from thousands of individual verdicts nobody wrote down. Constitutional AI is a code of written law: principles first, judgments derived from them, and any citizen (or auditor) can read the statute book.
Supervised learning from self-critique, then reinforcement learning from AI feedback — both driven by the same constitution.
Phase 1 in slow motion: a harmful response gets critiqued, then rewritten — all by the model, prompted with one principle.
Asking a model to rewrite a bad answer directly produces shallow edits. Asking it to critique first — "what is wrong with this response, per principle X?" — surfaces the failure mode explicitly, and the revision then has something concrete to fix. The paper found critique-with-a-specific-principle beats generic "improve this."
A headline behavioral change: CAI models refuse less evasively. Instead of "I can't help with that," the trained model explains the objection and offers safe alternatives — because the constitution's principles ask it to be helpful within harmlessness, and to explain its objections.
Phase 2 replaces the human preference labeler with a model consulting the constitution.
The paper's evaluation: crowdworkers comparing CAI against RLHF baselines on helpfulness and harmlessness, plus a 438-question HHH test.
RLAIF made preference data cheap enough to iterate on values daily instead of quarterly.
The training technique is transferable. The values in the document are where every downstream difference actually originates.
Constitutional AI's real invention is a separation of powers: the legislature (principles) is now a readable document, the judiciary (preference labels) is delegated to machines, and the executive (RL) is unchanged machinery. Whether that legislature is accountable — who votes, who amends — is the alignment debate this paper started and no training run can finish.
Check your understanding of the key concepts from the Constitutional AI paper.
Everything you need to remember about this paper.