The Reward Signal

Implicit Reward Models in Offline Preference Learning

DPO works on paper but falters on cyclic preferences and unbounded reward scaling.

Senior Writer · · 11 min read
Cover illustration for “Implicit Reward Models in Offline Preference Learning”
RLHF Foundations · September 8, 2026 · 11 min read · 2,529 words

DPO (Direct Preference Optimization) skips the reward model that standard RLHF relies on, folding a reward function directly into the policy through a mathematical trick. That trick works cleanly on paper, but a growing body of research (Apple ML Research and a handful of university labs among them) shows exactly where it falls apart in practice, and the failure points cluster around specific, measurable behaviors rather than vague hand-wringing about "alignment." This piece walks through the mechanism, the assumption baked into it, the variants built to patch it, and what the patches quietly admit about the whole idea.

Start with the pipeline DPO was built to replace. Conventional RLHF runs in two stages: train an explicit reward model on human preference annotations, then optimize a policy against that reward while a KL penalty keeps it from wandering too far from a reference model. The catch is that the reward model only ever sees the fixed, offline data it was trained on. Once the policy starts generating text on its own during optimization, it produces responses the reward model never encountered, an out-of-distribution problem that opens the door to reward hacking and over-optimization: the policy finds outputs the reward model scores well but that a human would not. On top of that, running two models means two training loops, two sets of hyperparameters, and twice the infrastructure to babysit. That expense is what pushed researchers to ask whether the reward function could just live inside the policy instead of sitting next to it.

How DPO's reparameterization trick produces an implicit reward

Under the KL-regularized RL objective that RLHF already uses, there's a closed-form relationship between the optimal policy and the reward function that produced it. DPO's authors noticed this could be run backward: instead of learning a reward and then solving for the optimal policy, express the reward in terms of the policy's own log-probability ratio and skip the middle step entirely.

The formula looks like this: r(x,y) = β log π_θ(y|x)/π_ref(y|x) + β log Z(x). The Z(x) term, a partition function that would otherwise be a nightmare to compute, cancels out the moment you're comparing two responses to the same prompt, which is exactly what preference data gives you. That cancellation is the whole trick. It's why DPO works with plain preference pairs and never needs to touch a reward network directly.

"Implicit" here means something specific, not just a marketing gloss. The reward is not a separate module that outputs a number. It's a mathematical byproduct: training a policy with the DPO loss is provably the same as picking a reward function that explains the preference data and then solving for its optimal policy, collapsed into one step. The whole setup leans on the Bradley-Terry model, which assumes P(y_w ≻ y_l | x) = σ(r(x, y_w) − r(x, y_l)). That assumption is what lets the RL objective turn into a plain maximum-likelihood loss computed straight from preference pairs, no reward network required. For practitioners the payoff is real: one loss function, one model, and on data that looks like the training distribution, accuracy roughly on par with an explicit reward model. The catch, as the next section covers, is buried in that Bradley-Terry assumption itself.

The Bradley-Terry assumption and what it cannot represent

Bradley-Terry forces every preference judgment into a single scalar ranking. If a reward score exists for each response, then A beats B, B beats C, and C beats A simply cannot happen, because scalars sort into a total order by definition. But human judgment doesn't always cooperate with that constraint. Tversky's 1969 work on cyclic preferences showed people routinely make choices that violate transitivity, preferring A to B, B to C, yet still turning around and preferring C to A when it's the pair on the table. That's not irrationality in some clinical sense; it's just how comparative judgment works when different attributes get weighted differently depending on the pairing.

So what happens when you force a Bradley-Terry model onto data with that kind of cyclic structure? Research on datasets built specifically to contain intransitive preferences found that Bradley-Terry reward models perform close to random guessing, while General Preference Embedding models, which don't assume a scalar ranking, get near-perfect accuracy on the same data. That's a wide gap, and it's not a training bug. It means any method inheriting the Bradley-Terry backbone, vanilla DPO included, is structurally incapable of learning from preference data that doesn't reduce to a clean, total order. Worth sitting with for a second: this isn't something more data or a bigger model fixes. It's baked into the shape of the equation. Practitioners collecting preference data should at least ask whether the judgments they're gathering are the kind that can be scalarized in the first place, because if they're not, DPO is solving the wrong problem elegantly.

The variant landscape that grew up around DPO's known weaknesses

Once the cracks were visible, a small industry of DPO variants sprang up, each one aimed at a specific, named failure rather than a general "let's improve alignment" gesture.

IPO (Azar et al., 2024) targets what happens when preferences are close to deterministic: DPO's logit mapping pushes the implicit reward toward infinity, which effectively cancels out the KL constraint that was supposed to keep the policy anchored. IPO swaps the Bradley-Terry loss for a squared loss on the log-likelihood gap, which keeps the implicit reward bounded and the regularization intact.

SimPO (Meng et al., 2024) goes after a subtler mismatch: DPO's reward depends on the reference model, but at inference time nobody consults a reference model, the policy just generates text based on average log-probability per token. SimPO drops the reference model entirely and uses a length-normalized average log-probability as the reward, adding a margin term to keep winning and losing responses cleanly separated.

KTO (Ethayarajh et al., 2024) solves a data-collection headache rather than a math one. DPO needs paired preferred/rejected examples, which are expensive to collect at scale. KTO borrows from prospect theory and works directly on unpaired signals, so a thumbs-up or thumbs-down on individual responses is enough.

SLiC (Zhao et al., 2023) adds a margin parameter plus a supervised fine-tuning term to balance ranking quality against generation quality. CPO (Xu et al., 2024) folds sequence likelihood into the reward itself, training it alongside an SFT objective to improve calibration.

None of these are solving hypothetical problems. Each one exists because someone ran DPO, watched it fail in a specific and reproducible way, and wrote a paper about the exact fix.

Where DPO's implicit reward reliably breaks down

Google DeepMind's Robust Preference Optimization research documented something almost comic in its predictability: DPO's implicit reward tends to overfit and trend toward infinite magnitude as training continues, since nothing in the loss puts a ceiling on it. The consequence is a degenerate policy where even the probability assigned to the preferred response collapses toward zero. The KL term, whose entire job is to prevent runaway divergence from the reference model, simply gets overwhelmed.

Rashidinejad and Tian (2025) split this into two distinct hacking modes. Type I is over-optimization on out-of-distribution actions, driven by gaps in what the offline data actually covered. Type II is degradation of the original model's competence, which happens when the preference data offers thin coverage of the genuinely high-reward responses. Different mechanisms, same downstream symptom: a model that gets worse the more you optimize it.

The most carefully measured version of this problem comes from an Apple ML Research study presented at ICLR 2025 (Lin, Seto, et al.), which ran 35 experiments across Gemma-2B, Gemma-7B, and Mistral-7B, testing five out-of-distribution settings. On in-distribution data, DPO's implicit reward model matched an explicit reward model closely enough that the choice barely mattered. Move to the five OOD settings, though, and the implicit reward model dropped a mean of 3% in accuracy relative to the explicit one, with a worst case of 7%. That might sound modest, almost like a rounding error. But in an iterative DPO pipeline, where each round's preference labels come from the model trained in the previous round, a 3 to 7 percent labeling error compounds round over round, quietly degrading the whole loop.

Then there's length bias, the sort of thing that sounds trivial until it shows up in every single output. DPO's implicit reward correlates with response length, so the policy learns that longer is safer, regardless of whether length adds any value. LD-DPO exists specifically to decouple the length signal from the substantive preference signal, which tells you how real the problem was in the first place.

Underneath all of this sits a ceiling nobody can patch around: DPO and its variants learn only from the fixed dataset they're given. The policy can shuffle probability mass around within the space the data already covers, but it cannot discover a genuinely new kind of good response that the original preference annotators never saw and never ranked.

How implicit rewards work outside LLMs, and what those settings reveal about the concept's limits

Implicit rewards didn't start with language models, and looking at where they've been tried before is useful precisely because it strips away the LLM-specific noise.

Inverse Preference Learning (Hejna and Sadigh, NeurIPS 2023) rests on an insight from control theory: under a fixed policy, the Q-function contains all the same information as the reward function, so the two are interchangeable. IPL trains a Q-function directly against expert preferences using something called the inverse soft-Bellman operator, cutting the reward network out of the picture entirely. On continuous control and robotics benchmarks, it holds its own against baselines using far more parameters, which is a genuinely impressive result. But the paper is upfront about the tradeoff: the implicit reward and the policy are both moving targets during training, non-stationary in the jargon, and that makes the whole setup less stable than a fixed explicit reward, especially early on when the Q-function hasn't settled into anything reliable yet.

Contrastive Preference Learning (CPL) pushes even further, questioning whether human preferences track reward at all. The paper's working hypothesis is that people may actually be responding to regret, meaning how far a choice falls short of what the optimal policy under their own preferences would have done, rather than to some absolute reward value. If that's closer to true, then any method that assumes preferences directly encode reward, DPO very much included, is built on a foundational premise that might just be wrong. CPL sidesteps the issue by combining regret-based preference modeling with maximum entropy, producing a supervised objective that converges to the optimal policy without ever running RL.

Look across both of these and a pattern falls out cleanly: every implicit reward approach, whether it's optimizing chatbot responses or robot arm trajectories, trades a simpler training setup for some flavor of instability or coverage limitation. That tradeoff shows up in robotics, continuous control, and language modeling alike, which suggests it's a property of the implicit-reward idea itself rather than an LLM quirk.

Proposed fixes and what they concede about the implicit paradigm

The Robust Preference Optimization approach trains the language model so its own induced implicit reward (that same scaled log-likelihood ratio from DPO's core equation) matches an explicit reward model, using an L2 regression loss over pairwise reward differences. It optimizes against a family of reward models rather than one single one, which adds robustness against variation in how different annotators label preferences, and it's theoretically equivalent to the online RLHF objective when the offline dataset is diverse enough. Notice what's happened here, though: the fix for DPO's instability is to bring an explicit reward model back into the loop through distillation. The elegant one-model promise quietly grows a second model again.

Another iterative self-alignment approach takes a different route. It uses the implicit reward from a DPO-tuned model to rank a fresh batch of outputs, builds a new preference dataset from those rankings, and reruns DPO on it. Two add-ons make this workable, length-regularized reward shaping to counter the length bias covered earlier, and experience replay to stop the dataset from degrading in quality over successive rounds. On standard benchmarks, this iterative approach has shown meaningful gains across different base models, which is a solid result by any measure. But the gains stall out after roughly three iterations, with no further improvement observed beyond that point, and if the starting DPO model is weak, the whole bootstrapping process may degrade rather than correct itself.

Put the two together and the concession is hard to miss. The most reliable fixes for DPO's implicit reward either reintroduce an explicit reward model through the back door or demand several rounds of retraining with careful engineering to keep the loop from falling apart. Neither outcome matches the original sales pitch: one loss function, one model, done in a single pass.

What practitioners should actually do with this knowledge

None of this means DPO is broken or should be avoided. It means the implicit reward trick has a domain where it works well and a domain where it quietly falls over, and the job is knowing which one applies before training starts.

Implicit rewards are probably fine when preference data is dense and closely matches the prompts the model will actually see in deployment, when the underlying human preferences are reasonably transitive rather than cyclic, and when a bit of length bias is tolerable or something a post-processing step can catch. That covers a lot of practical fine-tuning work, honestly, especially narrow-domain assistants where the input distribution doesn't shift much after launch.

The calculation changes when deployment prompts will diverge meaningfully from the training distribution. That's exactly the scenario where Apple's documented 3 to 7 percent OOD accuracy gap starts to matter, particularly once you're running iterative DPO and that gap compounds across rounds. It also changes when the preference data has thin coverage of the actual high-reward responses, raising the odds of Type II degradation, and any team planning multiple rounds of DPO should treat the preference-labeling step as the place where an implicit reward model's weaker generalization does the most damage.

Variant choice follows fairly directly from the failure mode you're worried about. Near-deterministic preferences call for IPO over vanilla DPO, since it keeps the implicit reward from running off toward infinity. Sparse or unpaired feedback points toward KTO. If verbose, padded-out responses are a known risk, SimPO's length-normalized reward or LD-DPO's explicit length term are the natural picks.

The deeper point sits underneath all of these specific choices: whether to go implicit or explicit is really a downstream decision from how the preference data was collected in the first place. The coverage and distribution of those preference pairs sets a hard ceiling on what any implicit reward can ever learn, no matter how the loss function is dressed up. For teams building alignment or content-quality systems at any real scale, the lesson from the LLM work and the older robotics and control literature lines up neatly: making the reward implicit doesn't make the reward-specification problem go away. It just moves the problem into data curation and distribution management, where it was arguably harder to see coming in the first place.

Sources

  1. proceedings.neurips.cc
  2. On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
  3. arxiv.org
  4. Robust Preference Optimization through Reward Model Distillation
  5. towardsdatascience.com
  6. proceedings.iclr.cc
  7. machinelearning.apple.com
  8. arxiv.org
Filed underRLHF Foundations

More in RLHF Foundations