History Problem Core Idea Limits Results Impact Quiz Takeaways
Interactive Paper Explainer

Humans Rate, Agents Learn
RLHF, 2017

The original RLHF paper: instead of reward functions, show humans pairs of short trajectory clips and ask which is better — a learned reward model solves Atari and robot locomotion with feedback on under 1% of interactions.

Start Learning Read the Paper ↗
<1%
Of interactions labeled
2
Feedback types tested
~700
Human comparisons (lab study)
2017
Christiano et al.
History

The Reward Problem

Reinforcement learning worked when you could write the objective; RLHF was born for everything else.

1950s-2016
Hand-crafted rewards
RL progresses on objectives engineers can specify — game scores, simulated physics. Real goals resist specification.
2016
Specification gaming documented
Reward hacking folklore accumulates: boat racers looping for bonuses (the classic CoastRunners clip), agents exploiting simulators. The 'reward function is wrong' era.
2016-17
Preference ideas brewing
Earlier work (including Akrour et al. and Wirth et al.) learns from ordinal preferences — but not scalable deep RL on modern tasks.
Jun 2017
🚀 Christiano et al.
Deep RL from human preferences: asynchronous reward-model training on human clip comparisons, coupled to deep RL agents — Atari, simulated locomotion, no reward function access.
2022
The LLM inheritance
InstructGPT (entry #40) ports the loop to language: preferences → reward model → policy optimization. The 2017 structure runs inside every aligned chat model.
Preferences Scale Functions Don't

A reward function is one expert's total specification, written before training and debugged after. A preference comparison is one non-expert's local judgment — 'this robot walk looks better' — cheap, natural, and robust: humans agree on qualities they could never formalize. The paper's loop learns a reward model from a few hundred to a few thousand such comparisons and lets RL do the rest. The result: complex behaviors specified with feedback on less than 1% of agent interactions.

Chapter 01

Objectives That Resist Writing

The 2016 problem set that made reward specification the bottleneck of applied RL.

✍
The Specification Wall
  • Real goals (a natural robot gait, a helpful answer) have no clean reward formula
  • Proxy rewards get gamed — agents exploit the gap between score and intent
  • Expert demonstration (imitation) is expensive and caps out at demonstrator skill
  • Watching every agent action and correcting it is a human-time impossibility
👍
The Preference Answer
  • Humans compare pairs of short trajectory segments — seconds of judgment, no rubric
  • A reward model learns to predict human preference on arbitrary segments
  • RL agent optimizes the LEARNED reward — deep RL machinery unchanged
  • Asynchronous loop: reward model keeps improving from fresh feedback mid-training
Analogy — The Talent Scout, Not the Coach

Writing a reward function is being a coach who must fully describe a perfect tennis swing before anyone plays. Preference RLHF is a talent scout: watch two 3-second clips, point at the better one, repeat. The scout never writes a manual — yet after a few hundred judgments, the training program they implicitly define produces better swings than any manual could.

Chapter 02

The Loop

Three components, one asynchronous cycle — the architecture InstructGPT inherited wholesale.

🤖 RL agent
Standard deep RL (A3C-era actors) interacting with the environment — it only ever sees the reward model's scores, never a true reward.
📊 Reward model
A network mapping (trajectory segment → predicted human preference score), trained on the growing comparison dataset.
👤 Human raters
Shown pairs of 1-2 second clips, asked which is better (or tie). Non-experts; feedback arrives asynchronously.
🔄 The cycle
Agent explores → clips sampled → humans compare → reward model retrains → agent's reward signal improves → repeat. No synchronization required.
What the humans saw
  • Simulated robotics: Hopper / Walker / Swimmer locomotion gaits
  • Atari clips from an agent trained on the learned reward (pong, boxing, space invaders)
  • Pairs with optional "equal" rating — judgment under seconds
  • Lab study: ~700 comparisons from non-expert humans sufficed for locomotion behaviors
Design details that mattered
  • Segments sampled for informativeness — not random frames, but agent states where preference data reduces uncertainty
  • Reward model ensembles / uncertainty handling for stability
  • Preference loss: Bradley-Terry style logistic model of pair outcomes
  • Feedback budget: under 1% of interactions across experiments
Interactive Demo — One Round of the Preference Loop

Walk a full cycle: agent explores, clips get sampled, a human judges, the reward model updates, the policy follows.

Chapter 03

The Honest Limits

The 2017 paper's own caveats — the same caveats that still shape RLHF safety debates.

What It Did Not Claim

The paper is careful: learned rewards can be exploited by the agent — the reward model is an approximation of preferences, and a policy optimizing it hard will find the approximation's seams (the ancestor of today's reward hacking and overoptimization literature). Human raters are also the objective: rater bias, inconsistency, and scalable oversight limits are inherited directly. These aren't footnotes — they are the research agenda of alignment's next decade, visible from day one.

Interactive Demo — Reward Specification: Three Interfaces

Tab through how you can tell an agent what you want — and the cost profile of each.

Chapter 05

Behaviors from Seconds of Judgment

What non-expert humans achieved without ever writing a reward.

L(r) = −E[(σ, τ1, τ2)~human] log σ( r(τ1) − r(τ2) )
σ ∈ {τ1, τ2, tie}
A human judgment
Which of the two shown segments is better — the entire supervision signal.
r(τ)
Reward model
Predicted human-approval score of a trajectory segment — the learned objective.
Bradley-Terry
The loss shape
A logistic model of pairwise outcomes: preference probability grows with the score difference.
<1%
The budget
Feedback on less than one percent of agent-environment interactions across the paper's experiments.
LOCOMOTION
natural gaits
hopper/walker behaviors from ~700 comparisons
ATARI
playing behavior
game behavior learned from clip preferences
FEEDBACK COST
<1%
of interactions labeled by humans
LOOP
async
reward model improves while the agent trains
Interactive Demo — From 2017 Robot to 2022 Chatbot

The same diagram, relabeled. Press reveal to watch the loop become InstructGPT.

Legacy

Legacy — The Alignment Blueprint

Every aligned chat model runs a descendant of this loop.

🔁 InstructGPT and the RLHF industry
The 2017 loop — preference data → reward model → policy optimization — is the alignment mechanism of modern assistants (entry #40), unchanged in structure five years later.
🎯 Preference interfaces everywhere
Constitutional AI (entry #42) replaces the human rater with a constitution; DPO (entry #44) collapses the loop's math — both are edits to THIS paper's architecture.
📉 The <1% budget insight
Preference supervision is information-dense: hundreds of comparisons steer behaviors millions of interactions couldn't — the economic case for RLHF over demonstration.
⚠️ What it did NOT solve
The paper itself flags the seams: reward-model exploitation (overoptimization), rater disagreement and bias, scaling to judgments humans can't make consistently — the open problems of alignment, visible in 2017.
🛤 Read next
The alignment stack it seeded: PPO · InstructGPT · DPO
Test Yourself

Quick Quiz

Check your understanding of the key concepts from RLHF Origins.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ 2017 RLHF: human clip preferences → learned reward model → deep RL policy — no reward function needed.
✅ Feedback on under 1% of interactions: preference supervision is information-dense.
✅ Asynchronous loop: reward model retrains mid-training while the agent explores.
✅ Bradley-Terry logistic loss turns pairwise judgments into a scalar objective.
✅ The paper itself flags reward-model exploitation and rater limits — alignment's open problems, visible from day one.
✅ Read it as the blueprint InstructGPT ported to language five years later.