The Reward Signal

Preference Dataset Quality Metrics for RLHF

Small, high-quality preference datasets outperform larger, noisier ones in training reward models.

Editor at Large · · 12 min read
Cover illustration for “Preference Dataset Quality Metrics for RLHF”
RLHF Foundations · September 10, 2026 · 12 min read · 2,734 words

RLHF pipelines run on preference data: a prompt, two candidate responses, and a label saying which one a human liked better. That label trains a reward model, and the reward model steers the policy. The quality of that preference data, not the volume, is what actually decides whether the resulting model gets better or worse at understanding people. This piece walks through a framework from Judy Hanwen Shen, Archit Sharma, and Jun Qin that offers concrete metrics for comparing preference dataset quality.

Right now, most teams pick from a short list of public preference datasets, and the comparison usually stops at "how many rows does it have." That's roughly like judging a restaurant by how many tables it has instead of what comes out of the kitchen. A reward model trained on a distorted signal doesn't just underperform, it actively teaches the policy the wrong lesson, and everything downstream inherits that mistake.

The three-axis framework from Shen, Sharma, and Qin (2024)

The paper is called "Towards Data-Centric RLHF: Simple Metrics for Preference Dataset Comparison," presented at the NeurIPS 2024 Workshop on ATTRIB in October 2024, done during an internship at Apple. The authors set out three axes: scale (measured as effective sample size, not raw row count), label noise (how much annotator disagreement is baked into the labels), and information content (how much the prompts and response pairs actually teach a model). None of the three works alone. A dataset can look great on one axis and still fail a reward model in practice.

The metrics had to clear one bar before anything else: they needed to work regardless of which base model was doing the reward modeling, and apply to any dataset built from pairwise preferences, not just the handful the field already knows by name. Validation ran two ways, in-distribution performance and domain generalization, tested against a standard reward modeling benchmark, with ablations across different model sizes to check the metrics didn't just work for one scale of model and quietly stop working for another.

The headline result is the part that should make practitioners sit up: a small subset, a fraction of a full dataset, was already enough to get most of the benefit. Composition mattered more than sheer size for getting good performance. That's a data-centric argument in the same spirit as data-centric work in supervised learning, except applied specifically to pairwise preferences, where the unit of "data" is a comparison, not a single labeled example.

Scale: why more data does not reliably mean better alignment

Pre-training taught everyone that more tokens equal a better model, more or less. Preference data breaks that rule. Shen, Sharma, and Qin found that increasing dataset size doesn't reliably improve reward model performance, and past a certain point it can actively hurt performance on benchmark tasks. That's worth sitting with for a second, because it cuts against a decade of scaling intuition.

The fix in how to think about scale is to separate raw count from effective sample size, the number of pairs that actually carry non-redundant signal. Two datasets with identical row counts can carry wildly different amounts of real information, which means sizing up a dataset by row count alone is a bit like judging a book by its page count instead of, well, what's written on the pages.

Composition turns out to be a quality metric in its own right. Shen et al. found some datasets dominate specific tasks across every sample size tested: UltraFeedback wins on Chat, HH-RLHF wins on Reasoning, SafeRLHF wins on Safety. So matching the dataset to the target task beats stacking the largest pile of data available, and that's a genuinely counterintuitive result for anyone raised on the "bigger corpus, better model" mantra.

Coverage is the other half of scale that raw counting misses entirely. If the collected pairs only reflect a narrow slice of tasks, tones, or risk profiles, the model trained on them generalizes poorly the moment it steps outside that slice. And here's a practical ceiling worth naming: most open-source preference datasets are in English, with non-English and domain-specific data remaining scarce, largely because that kind of annotation is slow and expensive to collect. Breadth doesn't come cheap.

Label noise: measuring how much annotator disagreement corrupts the training signal

Annotators disagree. Sometimes it's inexperience, sometimes inattention, sometimes personal bias creeping into a judgment call, and in rarer adversarial cases, someone's just gaming the label. But a lot of the disagreement isn't anyone's fault: two responses can each have real strengths and real flaws, and forcing a single ranking label onto that comparison asks the annotator to resolve an ambiguity that might not have a clean answer.

Inter-annotator agreement (IAA) is the standard tool for measuring how much this matters. Raw percent agreement is misleading on binary comparisons, though, because two reviewers will match by chance a good chunk of the time even if neither is paying close attention. That's why the field leans on chance-corrected measures: Cohen's Kappa (κ) for two annotators, Krippendorff's Alpha when there are more annotators or missing labels. Kappa corrects for chance agreement, which raw percent agreement ignores. One caveat worth flagging: the correction can overestimate quality when one label shows up far more often than the other.

Landis and Koch's interpretation scale puts κ between 0.61 and 0.80 as "substantial agreement," and anything above 0.80 as "almost perfect." In annotation practice, a target of around κ = 0.7 gets cited often enough to count as a rough industry benchmark.

The real-world numbers here are the part that should give practitioners pause. The MultiPref dataset showed diverging preferences on roughly 39% of pairs, with a reported quadratic weighted κ of just 0.268, well below "substantial." research has found human graders disagreeing on preference labels for somewhere between 20% and 30% of the data, and that disagreement clustered on prompts that were subjective, hard, or just obscure. So a meaningful chunk of the "ground truth" that reward models train on isn't ground truth at all, it's a coin flip with extra steps.

Beyond kappa, there's active work on probabilistic graph models that jointly estimate task difficulty, latent topic, rater bias, and rater confidence to try to infer the true label underneath the noise. That's the direction research goes when pairwise kappa alone can't explain what's happening in the data.

At the dataset level, the practical metric is noise invariance: inject synthetic label noise and watch how much reward model performance degrades. A dataset that holds steady is more noise-resilient than one that falls apart, and that gives teams one comparable number per dataset instead of a pile of per-session statistics that don't generalize. On the ground, this plays out as an ongoing control, not a one-time audit: disagreement analysis, consistency checks, and targeted review of the edge cases where annotators keep splitting.

Information content: what prompt diversity and response distinguishability actually capture

Information content splits into two separate questions. First, how much range does the dataset actually cover, across tasks, domains, tone, risk level? Second, when a response pair shows up, how clearly do the two responses differ in quality? A dataset can nail one and completely whiff the other.

The core distinguishability metric is reward margin, the gap in score between the preferred and rejected response in a pair. Pairs with a small margin are more likely to contain annotation noise and give the model less to learn from, which makes intuitive sense: if two responses are nearly indistinguishable in quality, whatever label they got is close to arbitrary. But margin alone can mislead across datasets. A big raw margin in a narrow, low-diversity dataset might carry less real signal than a moderate margin in a dataset that spans a much wider range of scenarios.

Two newer metrics try to fix that blind spot. Alignment potential (arXiv:2503.01864) measures the gap between a model's current implicit reward margin and the target explicit margin, essentially estimating how much room for improvement a given slice of data actually offers that specific model. Training selected by this metric consistently improved alignment performance across different base models and optimization setups. Distance Calibrated Reward Margin, or DCRM, introduced in June 2025 (arXiv:2506.14157), divides the reward margin by two distance proxies, edit distance and probability difference, to measure how dense the useful differences are among all the differences present in a pair. Across the settings tested, higher DCRM tracked with better training outcomes.

The coverage gap shows back up here too. A dataset that looks large by sample count can still cover a narrow information space if it's overwhelmingly English and skews toward a handful of domains, and information content is genuinely the hardest of the three axes to eyeball, because it needs a look at both prompt spread and response separability, neither of which shows up in a metadata file.

How the three axes interact and why treating them separately misleads

Here's where the framework gets interesting: none of these axes tells the whole story in isolation, and treating them as independent checkboxes produces some genuinely misleading conclusions.

Scale without information content just inflates a row count. A huge pile of near-identical prompts with ambiguous, hard-to-distinguish response pairs looks impressive on paper and teaches the model almost nothing. Low noise without diversity is its own trap: tight annotator agreement on a narrow domain produces beautifully clean labels that generalize badly, and a suspiciously high IAA score can be a sign the task was just easy and homogeneous, not that the dataset was well designed. Flip it around, and high information content with sloppy labeling has the opposite problem: the signal is genuinely there, diverse prompts, clearly distinguishable responses, but poor annotation discipline buries it under noise anyway.

The composition finding from Shen et al. ties all three together nicely. Choosing a dataset matched to the target task isn't only a coverage decision, it also shifts the noise profile and the information density the model actually sees during training. Domain-level composition touches all three axes at once, which is exactly why the axes shouldn't be scored in isolation.

For teams trying to triage a dataset, the sane order is scale and composition first (is there enough coverage across the domains that matter?), then noise (can the labels be trusted?), then information content (do the pairs actually carry a learning signal?). Getting this order backwards means burning a lot of effort on noise correction for a dataset that was mis-scoped from the start, which is a bit like debugging the paint job on a car that doesn't have an engine.

Benchmarks used to validate whether dataset quality translates to reward model performance

RewardBench, released by the Allen Institute for AI in March 2024, is the most widely adopted benchmark for this, with more than 150 entries logged by a 2025 survey. It runs 2,985 evaluation samples pulled from 23 data sources, split into chat, chat-hard, safety, and reasoning categories. The evaluation method is pairwise comparison accuracy: the reward model scores a chosen and a rejected response, counts it correct if the chosen response scores higher, then computes a weighted average within each category before averaging across categories.

RewardBench has a real limitation worth naming plainly. As scores climb from around 80 into the 90s, performance on other benchmarks doesn't reliably follow, and sometimes gets worse. Improving on RewardBench doesn't guarantee broader gains. Its chosen-rejected pairs were built with semi-automatic methods plus manual validation, there's a risk of spurious correlations baked into the reasoning subsets, and there's no built-in correlation analysis tying RewardBench scores to downstream performance. RewardBench v2 (Malik et al., 2025) was built to answer some of this, adding harder data, enforcing global best-of-N evaluation, and testing extremely difficult capability distinctions, including cases where responses are nearly identical and margin requirements get tight.

Preference Proxy Evaluations, or PPE, from Frick et al. at UC Berkeley (arXiv:2410.14872), takes a different route. It draws 16,038 labeled human preference pairs from Chatbot Arena, spanning responses from 20 different top LLMs across more than 121 languages, plus a second set of 2,555 prompts each paired with 32 sampled response options, 81,760 responses total across 4 models, grounded against verifiable correctness labels. PPE evaluates reward models across 12 metrics spanning 12 domains, and its real differentiator is that it's the only reward model benchmark tied explicitly to real-world human preference outcomes after RLHF: the authors ran full RLHF experiments end to end, deployed the resulting models on Chatbot Arena, and measured downstream human preference scores directly rather than proxying for them. It also grounds evaluation in crowdsourced human preferences and programmably verifiable correctness rather than model-generated labels.

Three metrics recur across these benchmarks: pairwise preference accuracy (pick the better response in a direct comparison), best-of-N accuracy (whether the reward model can identify the best response across a set of candidates), and ranking consistency (does the model's judgment hold up across paraphrased versions of the same prompt).

Biases that corrupt preference data and distort all three quality axes

Verbosity bias might be the most well-documented failure mode in this whole area. Reward models can be gamed into favoring longer responses, "length hacking," because preference labels in the training data correlate with response length independent of actual quality. This isn't purely a model problem either: researchers have traced the correlation back into the datasets themselves, meaning it's a data artifact that models then learn and amplify. Length-biased labels can distort the apparent margin between a long preferred response and a short rejected one, making the dataset look more distinguishable than it actually is.

Reward hacking, or overoptimization, is the second big one. Reward models overfit to patterns in their training set and lose fidelity to the true preference distribution they're supposed to approximate. A policy optimized against an overfit reward model finds the cracks and exploits them, producing outputs that score well on the proxy while drifting from what people actually want. On the scale axis, this means piling on more data from the same distribution doesn't reliably fix the problem.

LLM-as-annotator setups bring their own baggage. Using an LLM to generate or label preference pairs imports both verbosity bias and self-enhancement bias (a model rating its own outputs, or outputs like its own, more favorably) straight into the dataset. That makes measuring reward model overoptimization harder, because the label noise axis stops being a clean read on human disagreement and starts reflecting whatever quirks the annotating LLM brought to the table.

One design note worth flagging: pairwise comparisons tend to be more reliable than scalar scoring, because people are more consistent at judging relative quality than at assigning an absolute number. One reviewer's 7 out of 10 is another reviewer's 5, and that's not a training problem, it's just how humans rate things on scales.

Data quality as a model-relative property, not an intrinsic dataset property

The standard assumption in most preprocessing pipelines is that quality is a fixed property of the data itself: filter once, release the cleaned dataset, reuse it across any model or training configuration going forward. Work from Zhang et al. (arXiv:2510.13212, Hong Kong Baptist University) pushes back on that assumption directly. A given preference pair might genuinely help one model and genuinely hurt another, which means quality isn't sitting inside the data waiting to be measured, it emerges from the interaction between the data and whatever state the model is already in.

The evidence for this comes from experiments across a range of alignment benchmarks and different LLM families, where better alignment performance showed up using less data, as long as the selection process was model-aware rather than one-size-fits-all. The mechanism behind this is something called a truncated influence function (TIF), which quantifies each training example's effect on a specific model rather than assuming its value is fixed across every model it might ever train.

Put next to the rest of this framework, that's a genuinely humbling conclusion. The three-axis framework gives teams a much sharper lens than "how many rows does this dataset have," but even a dataset that scores well on all three axes isn't a universal answer. It's a good answer for a particular model, at a particular stage of training, and that's probably the more honest way to think about preference data quality going forward: less like a grade stamped on a dataset, more like a conversation between the data and the model asking, does this actually help you, specifically, right now?

Sources

  1. Towards Data-Centric RLHF: Simple Metrics for Preference Dataset Comparison
  2. Towards Understanding Valuable Preference Data for Large Language Model Alignment
  3. Towards Data-Centric RLHF: Simple Metrics for Preference Dataset Comparison | OpenReview
  4. arxiv.org
  5. arxiv.org
  6. arxiv.org
  7. arxiv.org
  8. arxiv.org
Filed underRLHF Foundations

More in RLHF Foundations