The Reward Signal

Direct Preference Optimization vs PPO for Instruction Tuning

DPO trains faster and simpler, but PPO explores better on hard tasks.

Senior Writer · · 13 min read
Cover illustration for “Direct Preference Optimization vs PPO for Instruction Tuning”
RLHF Foundations · September 5, 2026 · 13 min read · 2,897 words

Language models learn to sound fluent by predicting the next word over and over on huge piles of text; nothing in that process teaches them to be helpful, honest, or safe when a person actually asks them for something. Closing that gap is the entire job of instruction tuning, and the two most consequential techniques for doing it, PPO-based reinforcement learning from human feedback and Direct Preference Optimization, take almost opposite approaches to the same problem. This piece walks through both methods at the mechanical level: where each one wins, where the clean benchmark numbers stop telling the whole story, and how the industry actually stitches them together once the paper's been published and the deadline's real.

Pretraining gives a model something like general knowledge; it does not give the model a preference for one answer over another. Supervised fine-tuning on instruction-response pairs helps close part of the gap, since the model at least sees examples of the format a good answer takes. SFT teaches imitation rather than comparison, with a limited way to encode that Response A is better than Response B, only that some response exists and should be copied. The missing ingredient turned out to be comparison data itself: show human raters two model outputs for the same prompt, ask which one they prefer, and record the pair. That preference signal, chosen response versus rejected response, is the raw material both PPO and DPO are built to use. They just process it in almost opposite ways.

How PPO-based RLHF works and where it struggles

The classic RLHF pipeline runs in two stages, and both stages need their own infrastructure, their own tuning, and their own failure modes.

Stage one trains a reward model: a separate neural network learns to score responses by fitting to the human preference pairs, usually through a maximum-likelihood objective on the ranking signal. Stage two takes that reward model and uses it as the training signal for Proximal Policy Optimization, the algorithm Schulman and colleagues introduced in 2017, to update the base language model so it produces higher-scoring outputs. Critically, this update happens under a KL-divergence constraint that keeps the new policy from drifting too far from the original model's distribution. Drop that constraint and the policy will find the fastest path to a high reward score, which is often gibberish that happens to exploit some quirk in what the reward model considers "good." The constraint functions as the boundary between a coherent chatbot and a reward-hacking word-soup generator.

What makes PPO genuinely powerful, at least on paper, is that it explores. It generates new candidate responses and gets scored on them, going beyond imitating responses already sitting in a dataset, which means it can discover behaviors the human annotators never explicitly labeled. That's a real advantage, and it is why systems like ChatGPT and Claude were trained using reward-based PPO pipelines rather than pure imitation learning. This reflects how some of the most-used AI products in the world got tuned, well beyond any workshop paper's scope.

The costs are equally real. PPO is notoriously sensitive to hyperparameters, and the reward model, being a proxy rather than the actual judgment of a human, can get exploited by the policy in ways that are hard to catch before they show up in production. Running the pipeline means keeping four models in memory at once: the policy being trained, a frozen reference copy, the reward model, and a value or critic network used to estimate advantages. That is a lot of GPU memory and a lot of orchestration for a training loop that is already finicky. A 2024 regulatory filing from China Construction Bank noted, in almost deadpan language, that traditional PPO-based RLHF "involves numerous hyperparameters and can lead to training instability and collapse." When a bank's risk disclosure is citing your optimizer's instability, that's a fairly strong signal the pain is not confined to research labs. PPO remains effective, though expensive and occasionally unpredictable, and those two properties are exactly what sent researchers looking for an alternative.

What DPO actually does differently at the algorithmic level

Direct Preference Optimization, introduced by Rafailov and colleagues in 2023 under the memorably blunt title "Your Language Model is Secretly a Reward Model," starts from a mathematical shortcut. The alignment objective, it turns out, can be derived directly from the preference data without requiring a separately trained reward network. That means the reward model, the separate network everyone in the PPO pipeline trains and then optimizes against, is not strictly necessary. The alignment objective can be derived straight from the preference data itself.

In practice, DPO trains on preference pairs using a binary cross-entropy loss: push the model's probability for the chosen response up relative to what a frozen reference model (a copy of the SFT checkpoint) assigns it, and push the probability for the rejected response down, also relative to that same reference. The reference model quietly does the job the KL constraint did in the PPO pipeline, anchoring the policy so it doesn't wander off into incoherence, except here it's baked into the loss function rather than enforced as a separate penalty term during an RL rollout.

No reward model gets trained. No rollouts happen. No value function, no actor-critic bookkeeping, none of the apparatus that makes PPO a genuine reinforcement learning method. DPO runs on just the policy and reference model, instead of juggling the four networks PPO requires. Operationally, that buys training stability, since the objective behaves like ordinary supervised learning with gradients that behave the way gradients are supposed to behave, plus a real cut in compute and a much simpler codebase. There's no reward-model pipeline to build, no PPO clipping ratio to tune, no advantage estimator to debug at 2 a.m.

It's worth sitting with what DPO actually is, structurally: a supervised method that happens to operate on preference pairs. That distinction sounds pedantic, but it explains a lot of what shows up later in this piece, when DPO hits tasks that require exploring outside the training data rather than just re-weighting probabilities within it.

Where each method actually wins on benchmarks

The original DPO paper reported wins on summarization and single-turn dialogue: a controlled generation win rate around 61% for DPO versus 57% for PPO at matched sampling temperature, plus a robustness advantage on the tasks tested. On the tasks the paper tested, DPO looked competitive, and dramatically easier to implement to boot.

Later studies muddy that picture considerably. A 2024 paper titled "Unpacking DPO and PPO" found that PPO outperformed DPO across every held-out benchmark tested, with average gains of 1.3 points on reasoning, 2.9 points on coding, and 2.3 points on safety. The one consistent exception ran the other direction: PPO-trained models degraded on truthfulness by an average of 2.5 points. That's not a footnote; anyone optimizing specifically for factual accuracy should treat PPO's other benchmark wins with some caution, since the same training process that sharpens reasoning appears to loosen the model's grip on being straightforwardly correct.

A separate 2024 study, "Is DPO Superior to PPO for LLM Alignment," pushed the comparison into code generation and found something closer to a wall than a gap. On the CodeContest dataset, DPO produced zero correct solutions after a full epoch of training, a flat 0% pass rate. PPO, trained from the same base model, improved meaningfully on the same task, and a 34-billion-parameter PPO-trained model hit 22.4% on the 10@1k metric, beating AlphaCode-41B's 16.4% despite having roughly a fifth of the parameters.

Line those results up and a pattern starts to form. DPO does well on open-ended generation, the kind of task where the preference signal maps fairly directly onto surface qualities like fluency, tone, or avoiding unsafe content. PPO pulls ahead once the task demands structured reasoning, verifiable correctness, or output that requires exploring beyond whatever happened to be in the static training set, tasks like code and math where there's an actual right answer to converge on. The pattern hints that the RL loop may be developing something qualitatively different from what a supervised loss on fixed pairs can produce. Whether that difference is worth the infrastructure headache is, again, the question every team runs into eventually. There's no universal answer, which is precisely why the next two sections exist.

Diagram: PPO vs. DPO: Where Each Method Wins. Visualizes: Show a ranked comparison of benchmark deltas between PPO and DPO across five named dimensions: reasoning (+1.3 pts PPO), coding (+2.9 pts PPO), safety (+2.3 pts PPO), truthfulness (−2.5 pts…

Why data quality confounds the algorithm comparison

Here's a wrinkle that ought to make anyone pause before treating the benchmark numbers above as settled fact: several studies find that the quality of the preference data matters more than the choice of algorithm sitting on top of it. The authors of Unpacking DPO and PPO quantified the gap directly, reporting roughly an 8% improvement in performance from higher-quality preference data, compared with about a 2.5% improvement from switching algorithms entirely, from DPO to PPO. The data effect outweighs the algorithm effect by something like three to one.

That number should reframe how anyone reads a DPO-versus-PPO benchmark table. A lot of published comparisons are measuring, whether the authors intend to or not, differences in data pipelines dressed up as differences in optimizers.

There's a theoretical frame that helps explain why this happens, sometimes called the coverage argument. DPO does well when the preference dataset already covers the range of outputs the model is likely to produce; the static data is simply enough, since the model doesn't need to venture outside it to find good answers. PPO earns its keep when the task requires exploration, generating and scoring responses that fall outside whatever the existing dataset happened to capture.

That has a fairly concrete practical upshot. A team sitting on a well-curated, high-coverage preference dataset can reasonably expect DPO to match PPO's output quality, at a fraction of the infrastructure cost. A team working with sparse or narrow preference data, on the other hand, is going to watch PPO's exploration advantage compound over training, because DPO simply has nowhere else to look. This also explains, at least partly, why academic benchmarks and commercial deployments sometimes reach opposite conclusions: academic work tends to run on curated data where DPO looks great, while production teams are stuck with messier, sparser, less consistent labels where PPO's flexibility starts to matter more. Before picking an algorithm, then, it might make more sense to audit the data: its coverage, how fresh it is, how consistent the human labels actually are. That audit may decide the outcome before a single line of training code gets written.

Diagram: Data Quality Beats Algorithm Choice: The 3-to-1 Gap. Visualizes: Show a magnitude comparison between two sources of model performance improvement: switching from DPO to PPO yields ~2.5% improvement, while upgrading to higher-quality…

DPO's structural weaknesses that don't show up in clean benchmarks

DPO is, by design, an offline method. It trains on a dataset collected once, ahead of time, and it has no mechanism for generating a new response mid-training and getting feedback on it. That single fact explains most of what follows in this section.

As training proceeds, the policy's output distribution drifts away from whatever distribution produced the original preference data. The implicit reward function, which lives entirely inside the policy's own weights rather than in a separate network, was calibrated against that original distribution; as the policy moves, the calibration goes stale. One measure of this: DPO's implicit reward accuracy drops by around 7% in out-of-distribution settings, compared with an explicit reward model evaluated the same way. Variants that periodically refresh the preference data soften this problem without eliminating it; between refreshes the policy keeps moving, and the data keeps trailing slightly behind.

Then there's an asymmetry in how the optimization actually behaves. The structure of the DPO loss creates an asymmetry in how the optimization proceeds, which can result in a strange kind of alignment: the signal from rejected responses may dominate over the signal from chosen ones.

Length bias shows up for a related reason. Because DPO can't explore, it simply inherits whatever length pattern happens to sit inside the preference labels it was given. If annotators subtly favored longer answers (a well-documented tendency in human preference judgments generally), DPO has no way to correct for that on its own, and models can learn to pad responses simply because padding happened to correlate with "chosen" in the training set.

Overfitting is another risk worth naming. Overfitting is another risk: the loss function can keep improving while generation quality degrades, and unlike the PPO pipeline, DPO leaves behind no reusable reward model. And unlike the PPO pipeline, DPO leaves behind no reusable reward model once training finishes; the alignment signal is baked into the policy weights themselves, with nothing separable to reuse later for best-of-N reranking or a further round of RL without recomputing everything from scratch. Aggressive DPO training also carries some risk of eroding capabilities the SFT stage established, particularly on anything underrepresented in the preference dataset; if a skill wasn't well covered in the comparison pairs, DPO has no real way to protect it.

None of this is a hypothetical footnote. It's the mechanism behind the CodeContest result from the previous section: a 0% pass rate is exactly what you'd expect from a method that cannot explore, applied to a task that requires exploring to find a correct solution. The reasoning and coding gaps observed across the 2024 studies aren't some mysterious weakness; they're DPO running straight into its own structural ceiling.

How the DPO variant landscape tries to patch these gaps

The original DPO paper set off a fairly large wave of follow-up work, and most of it reads like a punch list against the weaknesses just described.

SimPO, out of Princeton NLP and presented at NeurIPS 2024, removes the reference model from the picture entirely. In its place, it uses a length-normalized average log-probability as the implicit reward, along with an added target reward margin. Normalizing for length is a direct fix for the verbosity bias: DPO, left unchecked, will reward long responses simply because long responses skewed "chosen" in the training data, and SimPO's normalization takes that shortcut away. According to its NeurIPS paper, SimPO beat DPO and several other DPO variants on AlpacaEval 2, MT-Bench, and Arena-Hard. The tradeoff is that SimPO-trained models tend to produce shorter, denser output, which is a feature on some tasks and a real limitation on tasks where thoroughness matters more than brevity.

ORPO, from Hong and colleagues in 2024, takes a different route: fold SFT and preference alignment into one training pass. The loss combines a standard cross-entropy term on the preferred response with an odds-ratio penalty applied to the rejected one. There's no reference model here either, and no separate SFT phase to run beforehand, which roughly halves the GPU memory needed during the alignment stage. That makes ORPO particularly attractive for teams working under a tight compute budget, or anyone who finds running SFT and alignment as two fully separate stages more operationally painful than it needs to be.

KTO, from Ethayarajh and colleagues in 2024, attacks the problem from the data-collection side rather than the loss-function side. Instead of requiring paired comparisons, chosen response versus rejected response, KTO trains on individual responses labeled simply good or bad. That's a genuinely cheaper thing to collect: judging one output in isolation is faster and asks less of an annotator than holding two outputs in mind and ranking them against each other. KTO earns its keep specifically when paired preference data is hard to come by, sparse feedback logged from a live deployment, say, where users react to single outputs rather than comparing two side by side. The tradeoff is informational: a head-to-head comparison between two responses simply encodes more signal about relative quality than two independent binary labels ever will.

Online or iterative DPO takes yet another angle, generating fresh preference pairs from the current policy partway through training, so the data stays closer to whatever the model is actually producing as it improves. That reduces the distribution-shift problem described earlier, though it doesn't erase it, and the improvement on genuinely hard tasks like code generation stays limited even under the iterative version. Step back and look at the whole variant landscape and a pattern emerges: the field has gotten quite good at patching DPO's specific, named weaknesses, the reference model, the length bias, the paired-data requirement, one at a time. What nobody has produced yet is a variant that closes the exploration gap with PPO on tasks where correctness is verifiable rather than a matter of taste.

How production deployments actually combine these methods

Treating PPO and DPO as rival candidates competing for a single trophy misreads how frontier labs actually build these systems. In practice, the choice rarely comes down to picking one pure method and running it end to end; it comes down to sequencing, budget, and which failure mode a team can least afford. A team with strong preference-data coverage and a tight compute budget has a real case for running DPO, or one of its variants, as the primary alignment stage and calling it done. A team building for code generation or multi-step reasoning, where the CodeContest result from earlier in this piece is a live risk rather than an academic footnote, has a real case for keeping PPO's exploration loop somewhere in the pipeline, expensive infrastructure and all. Increasingly, the honest answer involves some sequence of the two, plus one of the variants covered above filling in whichever gap the base method leaves open, which is a less satisfying conclusion than a clean leaderboard, but a more accurate description of what actually ships.

Sources

  1. arxiv.org
  2. arxiv.org
  3. arxiv.org
  4. arxiv.org
Filed underRLHF Foundations

More in RLHF Foundations