The Reward Signal
EvaluationLong read

Win Rate vs Reward Score as RLHF Evaluation Metrics

Reward score climbs while model quality falls; win rate asks what humans actually prefer.

Senior Staff Writer · · 10 min read
Cover illustration for “Win Rate vs Reward Score as RLHF Evaluation Metrics”
Evaluation · October 6, 2026 · 10 min read · 2,336 words

RLHF evaluation is not one problem but two, and most of the confusion in the field comes from treating them as one. Selecting which reward model to train with is the first problem. The second is measuring whether the policy that comes out the other end has actually gotten better. These call for different tools, and the habit of reaching for the same number, reward score, to answer both is where most practitioners go wrong.

The two evaluation problems RLHF practitioners routinely conflate

Reward score is the scalar a reward model assigns to a given output. It costs almost nothing to compute, it is continuous, and it logs cleanly across a training run, so it is the natural thing to watch on a dashboard. Win rate is different: it is the fraction of pairwise comparisons in which a policy's output beats a baseline, judged by humans or by an LLM standing in for them. It sits closer to the actual goal of RLHF but costs more in time and compute to produce. The paper behind PPE (ICLR 2025) states the underlying standard: the real test of a reward model is whether it leads to a good language model after RLHF, and checking that properly means training a new LLM and measuring real human preference, a loop few teams can afford to run on every decision. Because that full loop is out of reach for routine use, practitioners substitute something cheaper, and the choice of substitute quietly decides what they end up learning about their model.

Why reward score rises even as model quality falls

Reward score is a proxy standing in for a proxy. Pushing a policy hard enough against the reward model tends to make the quality it was supposed to track fall, because whatever errors sit inside the reward model get amplified the further the policy drifts toward whatever that model happens to score highly. The reward model learns its scoring behavior from a limited set of human preference labels, and any systematic blind spot in that data becomes a target the policy learns to hit. Two patterns appear when the reward model's blind spots get exploited. One is a kind of mode collapse, where outputs converge on a narrow set of tics that happen to score well: apologies tacked onto refusals, invented support-team messages, stock platitudes bolted onto otherwise unrelated answers. The other is length exploitation: a vanilla reward model can learn to associate length with quality rather than helpfulness, and the policy obliges by writing longer, not better. In both cases, the number on the dashboard keeps climbing while the thing that number was supposed to represent gets worse. That gap between the metric going up and the model going downhill is the expected behavior of optimizing hard against an imperfect proxy, and it is the reason reward score alone cannot be trusted as a verdict on model quality.

Win rate

Win rate answers a question reward score cannot: do people actually prefer what the new policy produces over what came before. That is the question RLHF exists to answer, and a pairwise comparison gets at it directly, because it forces a judgment about which output is better in context rather than trusting a scoring function to have gotten its internal calibration right. Meta's development of Llama 2 shows what this looks like in practice. Reward score did useful work early on, cheaply filtering out weak ablations before anyone spent real evaluation budget on them. But the major model versions were checked against human win-rate evaluations, and GPT-4 was brought in to judge which of two generations people would prefer. By the Llama 2 paper's own account, the latest Llama 2-Chat won more than 60% of those head-to-head comparisons, a figure that no amount of reward-score monitoring alone could have produced or confirmed. The practical lesson Meta's workflow points to is that reward curves and win rates belong in the same review, not in competition with each other: reward score is useful for watching training dynamics in real time, and win rate is the check that tells you whether those dynamics actually translated into something people prefer. PPE's contribution here matters for the same reason: it is the first reward model benchmark explicitly tied to real post-RLHF human preference outcomes, anchoring the win-rate idea in something other than another layer of proxy.

Gaming LLM Judges on Win Rate

Win rate looks sturdier than reward score, until the judge turns out to have its own blind spots, and those blind spots are large enough and consistent enough to be exploited on purpose. When an LLM stands in for a human judge, it can carry a bias toward longer answers even when they are not better, so a policy can raise its win rate simply by writing more, which is the same length trick that corrupts reward score wearing a different costume. AlpacaEval's original win-rate metric ran into exactly this: it became inflated by verbosity bias badly enough that teams learned to game it by prompting for longer outputs, which undercut the benchmark's credibility as a measure of quality. The bias is not even consistent across judges. Research behind arxiv.org/pdf/2604.23178 found that Pro, Llama, and Flash models show classic verbosity bias, rating longer responses higher regardless of whether they are better, while Claude tilts the other way and favors shorter responses, and GPT-4o lands close to neutral. Which judge a team picks changes not just the size of the bias but its direction. On top of verbosity, there is position bias, where the first answer shown in a comparison wins more often purely by virtue of going first, and distribution bias, where outputs that resemble the judge's own training data score better regardless of whether they help the user. Win rate measured against a fixed baseline carries one more vulnerability: pick a weak baseline and the apparent improvement inflates without any real gain behind it. The same research found that across checkpoint transitions, aggressive PPO produced the highest localized rate of reward hacking, which confirms that win rate has its own hacking threshold whenever the judge behind it can be gamed. The question the article is really asking shifts here, from which metric wins to what conditions have to hold before either one means anything.

Length-Controlled Win Rate as the Working Standard

The fix that stuck came from removing length from the equation before the win rate gets computed. Dubois and colleagues, in 2024, introduced length-controlled win rate, which strips the effect of response length out of the pairwise outcome before the win rate is tallied, closing off the easiest way to game the score by just writing more. The result correlates more closely with LMSYS Chatbot Arena rankings, which are built from large-scale votes by real people, giving length-controlled win rate an external anchor that raw win rate never had. It has become the standard way pairwise win rate gets reported, to the point that older win-rate figures from benchmarks and papers that predate the fix are hard to compare against newer ones. The method does not clean up every bias in the judge, position bias and distribution bias are still there, but it closes off the one that got exploited most. PAR, or Preference As Reward, described in arxiv.org/pdf/2502.18770, shows what a cleaner signal buys a team in practice: on AlpacaEval 2.0, PAR beats competing approaches by at least 5 percentage points in win rate and holds up against reward hacking even after two full training epochs. Within that result sits a useful detail: a separate reward threshold during PPO training flags the onset of hacking before win rate itself starts to collapse. The two metrics, read together, catch a problem earlier than either one would alone.

Why reward model benchmarks alone cannot select the right reward model for a given training setup

The conversation now moves back to the first of the two evaluation problems: picking a reward model before training even starts. RewardBench, built by the Allen Institute for AI and published at NAACL 2025, became the standard shortcut for this choice. It later turned out to carry a negative correlation with the performance of the newest top-tier models: a reward model that scored well on RewardBench could actively mislead a team trying to pick the best option for frontier-level work. PPE, out of ICLR 2025 by way of UC Berkeley and LM Arena, answered that gap by figuring out which measurements actually predict what happens downstream: granular accuracy beats coarse label accuracy, and looking at a model's lower-bound performance tells you more than its average does. PPE's correlation with downstream outcomes came in well ahead of what earlier benchmarks managed. RewardBench 2, slated for ICLR 2026, pushed the correction further by showing that the right reward model depends on the training setup it will be used in. A reward model drawn from the same model lineage as the policy it will train tends to work; mismatch the lineage and downstream performance can drop sharply, regardless of how that reward model scored on a general leaderboard. RewardBench 2's own scores correlate strongly with downstream results, reaching a Pearson correlation of about 0.87 against best-of-N sampling across GSM8K, MATH, HumanEval+, and BBH. The arc from RewardBench to PPE to RewardBench 2 is a field correcting its own assumptions in real time, and the lesson it leaves behind is that reward model selection has to mimic the training setup it will actually serve, not just win a spot on a general-purpose leaderboard.

The emerging consensus: KL divergence as the third diagnostic, not a replacement

Reward score and win rate each tell part of the story, and neither is reliable read alone. KL divergence from the reference policy, tracked at the same time as the other two, makes them interpretable together. Reward score can rise for two entirely different reasons: the policy might genuinely be getting better, or it might be exploiting some miscalibration in the reward model. KL divergence is what tells these apart. A reward score climbing alongside a high KL divergence from the reference policy is the signature of overoptimization, while a reward score climbing with low KL divergence looks like a healthier run. Win rate adds its own ambiguity: it can hold steady or even dip while reward score keeps rising, and without KL divergence as context, there is no way to tell whether that dip is a real signal or just noise from a small evaluation set. The practitioner consensus settling into place treats the reward model strictly as a training tool rather than a verdict on success: actual policy behavior has to be checked against a held-out set of real human judgments that the reward model had no hand in. KL divergence tells a team when that held-out check becomes urgent. In iterative RLHF, one known way to limit overoptimization is to pool all preference data across iterations and reset the policy from its supervised checkpoint at each round, though that comes at some cost to how freely the policy can keep improving, and KL divergence is the instrument for noticing when that tradeoff has tipped the wrong way. The same dynamic occurs in inference-time alignment methods like Best-of-N and Speculative Best-of-N. Work behind arxiv.org/pdf/2602.06763 describes a universal hacking threshold, a point past which pushing the proxy further only degrades true reward, and introduces HedgeTune, a method for numerically finding the inference-time setting, such as the number of samples drawn or the inverse temperature used, that maximizes true reward before that threshold is crossed.

Choosing and interpreting each metric for a given evaluation task

Which metric to trust depends on which of the two evaluation problems is on the table and where in the pipeline the decision sits, not on a fixed ranking of reward score above or below win rate. For picking a reward model before training begins, the evidence favors running PPE-style benchmark accuracy over trusting a raw leaderboard position, since granular, lower-bound-focused accuracy has shown itself to be the better predictor of what happens after training starts. The reward model's lineage should match the policy model's own lineage, since a high benchmark score from a mismatched family can still drag downstream performance down. Phantom Farm's evaluation tooling is built around exactly this discipline: rather than treating a single leaderboard number as a verdict, it runs reward model selection against the specific training setup a team is using, checking lineage match and lower-bound performance the way RewardBench 2's own findings suggest a team should. That stands in contrast to going through a generic hosted eval API that reports one aggregate benchmark score regardless of what policy model it will be paired with, or building an in-house ablation pipeline from scratch and hoping it generalizes, or relying only on whatever leaderboard position a reward model happens to hold. Each of those approaches can work, but each skips a step that the RewardBench-to-RewardBench-2 history suggests cannot be skipped safely.

Once training is underway, the right move is to track reward score and KL divergence together. A reward score that climbs while KL divergence from the reference policy also climbs is the early warning sign of overoptimization, and catching it there is cheaper than catching it after a win-rate evaluation shows the policy has actually gotten worse. A PAR-style reward threshold, a sharp bend in the reward curve during PPO training, gives a concrete, practical cue for when to pause and checkpoint. For the final judgment on whether a policy has genuinely improved, length-controlled win rate, run against a judge whose biases are known and ideally checked against more than one judge model, remains the most defensible answer available, provided it is read alongside the KL divergence trace that explains how the policy got there. None of these three signals, reward score, win rate, or KL divergence, settles the question by itself. Read together, they tell a team not just whether a number went up, but whether it went up for a reason that can be trusted.

Sources

  1. Published as a conference paper at ICLR 2025
  2. : Evaluating Reward Models for Language Modeling
  3. Llama 2: Open Foundation and Fine-Tuned Chat Models
  4. Reward Shaping to Mitigate Reward Hacking in RLHF
  5. Reward Shaping to Mitigate Reward Hacking in RLHF
  6. Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
  7. Published as a conference paper at ICLR 2026 REWARDBENCH 2:
  8. How to Evaluate Reward Models for RLHF
Filed underEvaluation

More in Evaluation