The original RLHF paper: instead of reward functions, show humans pairs of short trajectory clips and ask which is better — a learned reward model solves Atari and robot locomotion with feedback on under 1% of interactions.
Reinforcement learning worked when you could write the objective; RLHF was born for everything else.
A reward function is one expert's total specification, written before training and debugged after. A preference comparison is one non-expert's local judgment — 'this robot walk looks better' — cheap, natural, and robust: humans agree on qualities they could never formalize. The paper's loop learns a reward model from a few hundred to a few thousand such comparisons and lets RL do the rest. The result: complex behaviors specified with feedback on less than 1% of agent interactions.
The 2016 problem set that made reward specification the bottleneck of applied RL.
Writing a reward function is being a coach who must fully describe a perfect tennis swing before anyone plays. Preference RLHF is a talent scout: watch two 3-second clips, point at the better one, repeat. The scout never writes a manual — yet after a few hundred judgments, the training program they implicitly define produces better swings than any manual could.
Three components, one asynchronous cycle — the architecture InstructGPT inherited wholesale.
The 2017 paper's own caveats — the same caveats that still shape RLHF safety debates.
The paper is careful: learned rewards can be exploited by the agent — the reward model is an approximation of preferences, and a policy optimizing it hard will find the approximation's seams (the ancestor of today's reward hacking and overoptimization literature). Human raters are also the objective: rater bias, inconsistency, and scalable oversight limits are inherited directly. These aren't footnotes — they are the research agenda of alignment's next decade, visible from day one.
What non-expert humans achieved without ever writing a reward.
Every aligned chat model runs a descendant of this loop.
Check your understanding of the key concepts from RLHF Origins.
Everything you need to remember about this paper.