Your language model is secretly a reward model. DPO aligns models with human preferences using one simple classification-style loss — no reinforcement learning, no separate reward model, no PPO.
DPO arrived after three years in which alignment got powerful — and complicated. Here is the road that led to the paper.
The reward model that RLHF painstakingly trains is already hiding inside the policy. The reward that solves "maximize human preference while staying close to your reference model" is just a log-ratio of probabilities — so a separate reward model, and the PPO loop built around it, are both unnecessary detours.
By 2023 the standard alignment pipeline was a three-leg relay: supervised fine-tuning, reward-model training, and PPO reinforcement learning. Every leg adds cost, instability, and hyperparameters — and a dropped baton anywhere ruins the run.
InstructGPT-style RLHF is a relay race. The SFT runner hands the baton to the reward-model runner, who hands it to PPO. Each handoff can drop the baton: the reward model misgeneralizes to new outputs, PPO exploits its errors (reward hacking), and every runner has a pace to tune — learning rates, KL coefficients, rollout batch sizes. DPO is one runner holding the shortcut map: the closed form says exactly where the finish line is, so the run goes straight there.
DPO starts from the exact same mathematical goal as RLHF, then does algebra until the reward model disappears. Four moves, one closed form.
Humans prefer y over y′ with probability equal to the sigmoid of the reward difference. This is the model RLHF uses to train its reward model.
The policy that maximizes reward while staying close to πref satisfies this identity: reward equals a log-ratio, up to a prompt-only constant.
Solve for the policy instead: it is the reference re-weighted by exponentiated reward — a softmax over responses. The policy can be recovered from the reward.
Plug Step 2 into Step 1. The constants cancel, the reward model disappears — and what remains is a loss over the policy's own log-probabilities, on pairs alone.
The identity r(x,y) = β log π(y|x)/πref(y|x) + const is the exact solution of the KL-constrained reward maximization — not an approximation. Substituting it into the Bradley-Terry preference model yields a maximum-likelihood objective on pairs alone: the prompt-only constants cancel in the difference r(x,yw) − r(x,yl), and the reward model vanishes.
Read the loss as binary cross-entropy: the model "predicts" σ(β · margin) for the label chosen wins. That is why DPO trains like any supervised model — forward pass, backward pass, step — with two models in memory instead of a zoo of rollouts, rewards, and value estimates.
One gradient, two pressures: push the chosen response up and push the rejected response down — both measured relative to where the frozen reference model had them.
The loss shrinks as the margin between the chosen and rejected log-ratios grows. Every step makes the chosen response more likely and the rejected less likely, relative to πref.
The gradient is scaled by 1 − σ(β·margin). Pairs the model currently ranks backwards get near-full updates; pairs already ranked well barely register.
Both terms are ratios to πref, so updates are measured from the starting model — β is the throttle on how far the new policy may drift.
where h(x,y) = log π(y|x) / πref(y|x) is the (log) implicit reward. This sigmoid gate is why DPO does not waste capacity on pairs it already ranks correctly — samples that are already right contribute little.
Three tasks borrowed from the RLHF literature: sentiment steering, summarization, and dialogue. DPO matches or beats PPO while training like plain supervised learning.
| Method | Reward Model? | RL Sampling? | Training Stages | Stability | Alignment Quality |
|---|---|---|---|---|---|
| SFT (baseline) | No | No | 1 | very stable | baseline — follows format, not preferences |
| RLHF (PPO) | Yes — separate model | Yes — rollouts every step | 3 | notoriously twitchy — many knobs | strong, when carefully tuned |
| DPO | No — implicit in the policy | No | 1 | stable as supervised learning | matches or beats PPO |
Across the three tasks, DPO is never dramatically worse than the best RLHF baseline — and it is consistently more stable and cheaper to tune than PPO.
Within a year, DPO went from a Stanford preprint to the default alignment stage of open models — and to a whole family of descendants.
DPO is the direct successor of the RLHF pipeline from InstructGPT (Training Language Models to Follow Instructions with Human Feedback) — read that guide first to see exactly what DPO removed. DeepSeek-R1 shows where large-scale RL came back into the picture for reasoning.
Check your understanding of the key ideas from the DPO paper.
Everything you need to remember about this paper.