TRPO's stability without its second-order machinery: PPO clips the policy-update ratio so each optimization step stays near the old policy — first-order simplicity that became RL's default algorithm.
The 2015-17 stability problem, and the two-line idea that solved it cheaply.
The policy ratio r_t(θ) = π_θ(a|s) / π_old(a|s) measures how far the new policy has drifted on a given action. TRPO constrained its KL divergence globally — expensive math. PPO's move: clip the ratio to [1−ε, 1+ε] inside the objective. Once the ratio leaves the band, the surrogate stops rewarding further drift — the incentive to move too far simply switches off. One hyperparameter (ε ≈ 0.1-0.2), first-order gradients, no constraint solvers.
The instability problem every policy-gradient practitioner met in 2016.
TRPO is a steel cage: updates are mathematically forbidden from exceeding the trust region — safety via constraint solver. PPO is a banister on the stairs: you're free to lean, but past a small angle the railing just stops helping you lean further. Nothing is forbidden; the incentive structure quietly refuses to reward overreach.
The clipped surrogate — the whole algorithm in one expression.
Why the LLM era adopted a 2017 robot-control algorithm verbatim.
RLHF's optimization problem (entry #40) is nasty: a huge policy (the LLM), a learned reward (the RM), and catastrophic-forgetting risk on every update. PPO's properties map exactly: the clip keeps the LM near its reference distribution (same role as TRPO's trust region, but through the ratio in token space); multi-epoch minibatches make expensive generations reusable; and the value head gives per-token credit without bespoke machinery. Hence InstructGPT: PPO + a KL penalty against the reference policy — the recipe Llama 2 ran, and the baseline GRPO (entry #46) and DPO (entry #44) were designed against.
PPO's results were 'good enough everywhere' — the most valuable property an algorithm can have.
Ten thousand baselines later, PPO is still the algorithm you reach for first.
Check your understanding of the key concepts from PPO.
Everything you need to remember about this paper.