Bradley-Terry Model Assumptions in Preference Learning
The model's four hidden assumptions about human judgment don't hold up in production reward models.

The Bradley-Terry model is the math underneath almost every reward model in RLHF: it takes pairwise comparisons ("response A beats response B") and turns them into a single latent score for each response, so the probability of A beating B is just a sigmoid over the gap between their scores. Bradley and Terry published this in 1952 as a probabilistic framework for learning from pairwise comparisons and inferring a latent ordering over items. It has since become the load-bearing beam under most preference learning pipelines, and like most load-bearing beams, nobody looks at it until something starts to sag. The wrong move here is treating that beam as neutral infrastructure; it encodes four specific bets about how human judgment works, and at least one of those bets is losing money right now in production reward models.
The core intuition is almost quaint in its simplicity. Every item, chess player, wine, or LLM response, carries some invisible strength number nobody observes directly. When two items face off, the outcome is a probabilistic draw weighted by the gap between those numbers. Ask an annotator "is A better than B" instead of "rate A from 1 to 10," because humans are more consistent at comparisons than at absolute judgments, and Bradley-Terry gives a clean mathematical way to turn thousands of those comparisons into a ranking. In modern RLHF pipelines, that latent strength becomes a reward model: an LLM with a scalar output head, trained on human preference labels to spit out one number per response. Same logic as an Elo rating in chess. A 1200 doesn't mean anything on its own; it only means something next to an 1100, and even then it's predicting the next game's outcome, not certifying that the higher-rated player is playing "good chess" by some independent standard.
That last point matters more than it looks like at first glance. Elo ratings, and by extension Bradley-Terry scores, are relational currency. They don't mean anything outside the comparison they were minted for. So what happens when a tool built for relative ranking gets used to make claims about absolute quality, safety, or truthfulness? Things get strange in specific, traceable ways, and the rest of this piece follows those traces. But the four assumptions need laying out first, because each one is doing more work than the tidy sigmoid equation lets on.
The four load-bearing assumptions the model cannot function without
Four assumptions hold this structure up, and none of them are optional extras. Pull one out and the other three wobble too.
The first is scalar latent reward existence: every response has exactly one real number determining how likely it is to win a comparison. Helpfulness, safety, accuracy, tone, and style all get folded into that single figure. This rules out, by construction, any judgment that's genuinely multi-dimensional or context-dependent. A response can't be "safer but less helpful" in Bradley-Terry's world; it's either a higher number or a lower one, full stop.
The second is strong stochastic transitivity, and this is the strict one. If A beats B more often than not, and B beats C more often than not, A has to beat C, by a margin at least as large as the stronger of the first two margins. That's not just "no cycles allowed." It's a hard constraint on the actual size of the gaps. Any real cycle, A beats B, B beats C, C beats A, is something the model cannot represent, not just handle badly.
Third: independence of irrelevant alternatives, or IIA, borrowed from Duncan Luce's 1959 choice axiom and widely used in economics and social choice theory. The preference probability between A and B has to depend only on A's score and B's score, never on what else sits in the comparison set. Contrast effects, decoy effects, the fact that a response looks better next to a bad one and worse next to a great one: all assumed not to exist.
Fourth, homogeneous and consistent annotators. Every comparison gets treated as an equally reliable draw from one shared underlying preference distribution, as if every human labeling data agrees on what "better" means and applies that standard identically every time. There's no slot in the model for weighting one annotator's judgment over another's, no mechanism for saying "this rater is having an off day" or "this rater holds a genuinely different value system."
None of these four operate in isolation. Break transitivity and IIA usually breaks with it, because a third response changing the outcome between two others is itself evidence of a cycle. Break annotator homogeneity and transitivity likely goes too, since different people applying different standards is a natural cycle-generating machine. They're a stack, not a checklist.
How scalar rewards flatten the actual structure of human judgment
Start with the boldest claim buried in assumption one: a single number is enough to capture how a response reads across every prompt type, every annotator, every dimension of quality a human might care about. That's a big ask, and the honest answer is that it's not enough. Not close.
Human judgment doesn't work that way, and this isn't controversial; it's obvious once stated plainly. A response can be more factually accurate and less pleasant to read. It can be thorough and unsafe. It can be short, correct, and cold, or long, warm, and slightly wrong. These tensions are real, not measurement error, and whenever they show up, a scalar reward model has no way to preserve both signals. It picks a winner and discards the rest.
What that means in practice: the reward model doesn't learn some platonic definition of "good response." It learns whatever weighted average of quality dimensions the annotator pool happened to produce that week, whatever mix of helpfulness, safety, and tone that specific group of raters leaned toward. Change the pool, get a different average, with no way to tell from the number alone which quality got traded for which.
Multi-objective reward modeling is the direct answer, and it's worth naming because it proves the tradeoff isn't theoretical. Instead of one scalar, track helpfulness, accuracy, and safety as separate scores, so tradeoffs stay visible and auditable instead of buried inside one figure. That costs more annotation work and complicates how multiple scores collapse into one training signal at the end, but it at least admits the dimensions exist. Scalar rewards persist anyway not from stubbornness: a single number is fast to train against, cheap to compute, and drops straight into standard RLHF optimization. That's a real advantage. It's just an advantage bought by pretending the flattening didn't happen.
Where transitivity breaks down in real preference data
Strong stochastic transitivity asks for something genuinely hard to find in noisy, crowdsourced comparison data: no cycles, and margins that compound in one consistent direction. That's a tall order for data collected from hundreds of different humans with different priorities, and expecting otherwise is where a lot of reward model debugging goes to die.
Research on human choice behavior going back to Amos Tversky's 1969 work on intransitive preferences, and more recent work by Agranov and Ortoleva in 2017, documents that people's preferences cycle, and not from sloppy measurement. Picture three responses. A beats B because A is more helpful. B beats C because B is safer. C beats A because C is shorter and the annotator was tired of reading by comparison two hundred. That's a cycle, and it isn't noise sitting on top of some true linear order; it's the structure of the data, because different attributes win in different matchups.
Cycles compound for reasons specific to RLHF. Different annotators weight helpfulness, safety, and brevity differently to begin with, and even one annotator can shift their own criteria mid-session as fatigue sets in. When majority labels on (A, B) and (B, C) point one way and (A, C) points the other, the reward model gets a training signal that argues with itself, three data points pulling in contradictory directions. This connects straight back to IIA: when adding a third response changes which of the first two wins, that's an IIA violation and direct evidence of intransitive structure at the same time. Two names for the same crack.
Here's the part worth sitting with. Bradley-Terry doesn't detect any of this. Fed cyclic data, it finds the scalar ranking that minimizes total inconsistency across every comparison and reports that ranking with total confidence. It doesn't flag the cycle. It doesn't say "these three responses can't be ordered." It erases the cycle and hands back a clean list, as if the contradiction never happened, which is the mathematical equivalent of smoothing over a fight by pretending nobody spoke.
What annotator heterogeneity does to a model that assumes it away
Assumption four says every annotator draws from the same latent preference distribution with the same reliability, meaning disagreement between raters is pure noise that should wash out with enough data. Ask anyone who has managed an annotation team whether that holds up, and expect a tired laugh.
Annotators differ in domain expertise, attention span, cultural background, and plain values, and some share of any labeling pool, however carefully vetted, ends up inattentive or inconsistent on a given day. Multi-annotator datasets built specifically to study this, MultiPref, HelpSteer2, HelpSteer3, and PersonalLLM among them, all document the same pattern: a substantial share of pairwise comparisons come back with split votes. Disagreement isn't the exception in these datasets. It's the baseline condition.
Standard reward modeling handles a split vote the same way every time: take the majority, discard the rest. That's an aggregation choice with real consequences, because it treats one perspective as ground truth and deletes the fact that a different, coherent perspective existed in the room. Diverging preferences aren't always someone being wrong; sometimes the prompt itself was ambiguous, or two annotators applied genuinely different, defensible frames, and majority voting doesn't resolve either situation. It just papers over both.
What gets lost specifically is the minority view, which might represent a coherent value system rather than an error, and once majority aggregation runs, that view goes invisible to everything downstream. The directional problem gets worse when unreliable or biased annotators aren't discounted relative to careful ones. If their errors were random, they'd cancel out across a big enough dataset. But annotator errors aren't always random: shared biases, a preference for longer answers, a preference for formal tone, push in the same direction across many raters at once, and that consistent push survives averaging intact. The reward model ends up encoding the preferences of whatever effective majority happened to show up that day, and a policy trained against it narrows toward that one slice of behavior instead of anything broadly representative.
Length bias and sycophancy as systematic distortions the model absorbs rather than filters
Length bias is the textbook case of a correlated annotator bias hiding in plain sight. Across a wide range of the preference datasets used to train reward models, the response annotators pick as "better" tends to run longer than the one they reject, consistently enough that it reads less like coincidence and more like a systematic proxy: raters reaching for length as a stand-in for thoroughness, even when the extra words add nothing.
Bradley-Terry has no way to filter this out, because nothing in the model asks whether a preference is "for the right reasons." It trains on majority labels and treats every signal in those labels as legitimate, so length becomes a de facto feature of the reward function whether anyone intended it or not. A policy trained against that reward doesn't get better; it gets longer, padding responses because padding is what the gradient rewards, independent of whether the padding says anything useful.
Sycophancy runs on the same track. Annotators show some tendency toward responses that agree with them, flatter their framing, or avoid pushback, and because that tendency is shared across raters rather than randomly distributed, it survives aggregation the same way length bias does. Random errors cancel out over a big enough sample. Correlated ones don't; they stack in the same direction and hand the optimizer a consistent gradient to climb, one that rewards satisfying annotator psychology over satisfying the actual ask.
Here's the frustrating part: there's no clean way to catch this after the fact. The noise model underneath Bradley-Terry assumes deviations from the "true" preference are random and independent, and there's no external ground truth sitting outside the annotation process to check that assumption against. Length bias and sycophancy don't announce themselves in the loss curve. They surface later, in a model whose outputs have quietly drifted in length and tone from whatever the task actually required.
A specific failure the model's relative structure produces: safe responses scored as unsafe
Walk through this one slowly, because it shows exactly how the abstract math turns into a real alignment problem. Bradley-Terry scores are relative by construction: a response's number depends entirely on what it's measured against, never on any property the response holds by itself.
Picture a batch of responses ranked for safety, and every single one is actually fine, no real harm anywhere in the batch. But some are marginally safer than others, one phrases a sensitive topic slightly more cautiously than the rest. Ranking pressure inside that comparison set can still push the least-cautious response, the one still perfectly safe in absolute terms, into negative reward territory. Not because it said anything harmful. Because it finished near the bottom of a group where everyone happened to be safe.
That's the mechanism, worth naming plainly: absolute safety judgment gets overwhelmed by relative ranking pressure. A response that deserves reward gets penalized for the company it kept in the comparison batch, not for anything it actually said. The sigmoid loss defining Bradley-Terry only cares about the gap between two scores; there's no floor in the math that says "below this line is actually unsafe" and "above it is fine regardless of rank." Safety, in this framing, has no absolute anchor at all, which is a strange thing to discover about a system whose whole job is supposed to be catching unsafe output.
The downstream consequence is a policy that learns to avoid behaviors that were never unsafe to begin with, tiptoeing around territory it never needed to tiptoe around. Call it over-refusal, call it distorted caution: either way, it's a misalignment that traces straight back to how the scoring is structured, not to bad annotation or bad intent anywhere in the pipeline.
How reward hacking and social choice paradoxes trace back to these assumptions
Reward hacking is what happens when a policy gets optimized hard against any imperfect proxy: the gap between what the reward measures and what quality actually requires becomes something the optimizer finds and exploits, a pattern documented in scaling studies of reward model overoptimization such as Gao and colleagues in 2023. Bradley-Terry doesn't cause reward hacking on its own, but it hands the optimizer a wider set of exploits than a cleaner signal would: no calibration across prompt types, no separation between quality dimensions, length bias already baked into the training data. The policy doesn't need some exotic loophole. It just leans harder into whatever the reward already rewards for the wrong reasons.
Then there's a finding that reads almost like a contradiction if you know social choice theory. RLHF pipelines built on Bradley-Terry have been shown to violate majority consistency, pairwise majority consistency, and Condorcet consistency, the same axioms economists use to judge whether a voting system behaves sensibly. When annotator preferences cycle, no Condorcet winner exists by definition, and Bradley-Terry's aggregation still confidently picks a winner regardless, an outcome no principled voting rule would select. It's running an election, ignoring that the election has no fair winner, and certifying a winner anyway.
And yet, RLHF systems built on exactly this foundation have worked, at scale, in deployed systems; the alignment work behind models described in Ouyang and colleagues' 2022 paper and Touvron and colleagues' 2023 Llama work both used variations on this approach with results good enough to ship. That gap, a formal violation of foundational social choice axioms sitting right next to strong practical performance, is genuinely unresolved. Nobody has explained it away cleanly, and it's worth sitting with rather than smoothing over.
Calibration problems are the more mundane symptom of the same root cause. A score gap that reliably predicts preference on coding questions might mean nothing on creative writing prompts, because Bradley-Terry has no built-in way to calibrate across wildly different comparison types. And because everything collapses into one scalar difference, diagnosing which specific assumption failure is driving a given weird reward output is close to impossible from the outside. The model doesn't leave a trail; it just leaves a number.
What the alternatives sacrifice and what they recover
Every fix on the table trades away something Bradley-Terry was good at, usually speed, simplicity, or interpretability, to buy back some expressiveness the four assumptions threw out. There's no free lunch here, and treating any one alternative as a strict upgrade misreads what's actually happening.
Pairwise preference models, sometimes called PairRM or PairPM, drop the scalar reward idea entirely and model win probability between two responses directly, given both as input at once. This represents cyclic and intransitive preferences that Bradley-Terry structurally cannot. What it gives up is the ability to score a single response on its own; without a second response to compare against, the model has nothing to say.
Nash learning from human feedback treats alignment as a two-player game, optimizing win rate in direct head-to-head matchups instead of fitting one scalar score per response. It handles non-transitive annotation data more gracefully than Bradley-Terry does. It also costs more to compute and is a good deal harder to interpret, since there's no clean per-response number left to point at.
A handful of direct preference optimization variants skip the reward-model step and the transitivity assumption altogether, estimating preference probabilities straight from comparison data. That sidesteps the scalar and transitivity failures at once, but it complicates the link between the training signal and what the policy actually ends up doing, which makes debugging a stranger exercise than it needs to be.
Multi-objective reward modeling, mentioned earlier as the answer to scalar collapse, belongs here too: split the single number into tracked dimensions, helpfulness, safety, accuracy, so tradeoffs stay visible and length bias can't sneak in disguised as quality. The cost is heavier annotation and a new open question: how to recombine those separate dimensions into one decision at the exact moment the policy has to act.
None of these alternatives kill all four assumption failures at once; each targets a subset and leaves the rest standing. The real mistake is chasing a "best" option in the abstract instead of asking which assumption a given application is most exposed to. Bradley-Terry itself stays a perfectly reasonable default when the situation actually resembles its assumptions: short response pairs, a roughly consistent annotator pool, low risk of real preference cycles. It turns into a liability exactly where one of those conditions clearly stops holding, and the job, if there is one, is noticing which condition failed before the reward model quietly bakes that failure into every response downstream.


