Reward Model Ensembling for Uncertainty Estimation
Using multiple reward models to quantify and act on training uncertainty.

A single reward model in RLHF gives you one number per response. No error bar, no flag for "trust this less." That's the fragility this piece is about: reward models are deterministic point estimators trained on thin, noisy preference data, and as the policy pushes into territory the reward model never saw during training, the model keeps handing back confident-looking scalars anyway. Ensembling is the fix that's gained the most traction, and the reason isn't elegance. It's that ensembling turns disagreement among models into a number the training loop can actually act on. Most teams reach for it as an afterthought bolted onto a finished pipeline, and that's backwards: the uncertainty signal has to shape how the RL update behaves from the start, or it doesn't do much of anything.
Start with what a reward model is asked to do. It takes a prompt and a response and outputs one scalar, supposedly standing in for how much a human would like that response. That scalar gets trained on preference data that is small and noisy compared to the pretraining corpus behind the policy itself, and the optimization that produces it (SGD over mini-batches, usually just one or two passes through the data to avoid overfitting) means two runs with identical hyperparameters can land on meaningfully different parameters. Banerjee and Gopalan (arXiv:2410.23726) show this directly: train 10 reward models on the same dataset with the same settings, feed all 10 the identical prompt-response pair about who created a well-known fictional superhero, and the scores spread out visibly. Same inputs, same recipe, different answers.
That gap matters most once RL training starts pushing the policy toward responses further from the distribution the reward model was trained on. The reward model's uncertainty grows as the policy explores, but nothing in a standard setup tells the policy that. Pan et al. (arXiv:2606.19818) trace this to two compounding failures: deterministic scores get treated as equally trustworthy no matter how out-of-distribution the input is, and group-based methods like GRPO standardize advantages uniformly, so a shaky reward estimate gets the same leverage in the update as a solid one. The policy learns to chase whatever the reward model rewards, including its blind spots, a failure mode usually called reward hacking or reward overoptimization. Tuning learning rates or adding more preference labels doesn't fix any of this. The architecture itself has no slot for "I'm not sure," so the fix has to come from somewhere else.
What epistemic uncertainty means in the reward modeling context
Two kinds of uncertainty get lumped together too often, and the difference decides what you're supposed to do about them. Epistemic uncertainty comes from limited data coverage: the model hasn't seen enough examples like this one, and in principle more or better-distributed training data would shrink it. Aleatoric uncertainty is different. It's noise baked into the labels themselves, annotators disagreeing with each other, or annotators simply disagreeing with one another in ways that more data cannot resolve.
Treating these two as the same problem is the mistake worth naming outright, because the right response to each is opposite. Epistemic uncertainty calls for caution: don't let the policy update aggressively on inputs the reward model barely understands. Aleatoric uncertainty calls for robustness instead, since the ambiguity is real and permanent, and pretending you can resolve it with more labels just wastes annotator time.
Out-of-distribution inputs, the kind an RL policy generates as it explores, mainly produce epistemic uncertainty. The reward model is being asked to score something structurally unlike what it trained on. This disagreement pattern can be turned into a practical signal: if several independently trained reward models agree on a score, that's a decent indicator the score is reliable, and when they scatter widely, that's a reason to distrust the average. Banerjee, Saha, and Gopalan (arXiv:2507.15906) back this up theoretically, showing that high reward model variance drives policy overfitting and provably raises the odds that RL training produces a worse policy than doing nothing at all. A system blind to its own confidence can't apply caution selectively. It applies none, everywhere, or all of it, nowhere. Ensembling is the practical way to break that all-or-nothing failure.
How an ensemble converts disagreement into a usable uncertainty signal
The basic recipe is simple enough to state in a sentence: train several reward models with different random seeds, reinitialized final layers, or bootstrap-sampled slices of the training data, then look at how much they disagree. The now-standard version of this approach involves copying a pretrained LLM, reinitializing just the last linear layer (the reward head), training on bootstrap-sampled data so each copy sees a slightly different view of the same dataset, and at inference time computing the variance across the ensemble's scalar outputs.
Banerjee and Gopalan (arXiv:2410.23726) scaled this up with 10 independently trained reward models built on Gemma-2B-it. Each model runs about 9.34 GB on disk with its scalar reward head included, so the full set of 10 comes out to roughly 90 GB of storage. That's a real number, and it shows the cost isn't abstract even at a comparatively modest 2B parameter scale. The variance across those 10 models, computed per prompt-response pair, becomes the empirical stand-in for "how much should I trust this reward."
Get precise about what that variance is and isn't, because people conflate the two constantly. It measures how much the model family disagrees, not how close the ensemble's average is to the truth. High variance correlates with distance from the training distribution, which is useful, but low variance doesn't guarantee the mean score is correct. Models can agree and still be wrong together.
Once you have the signal, it does two jobs. It can guide active learning, sending human annotators toward the prompt-response pairs where the ensemble disagrees most, since that's where labeling effort pays off. And it can regularize the RL update itself, penalizing or down-weighting policy changes driven by uncertain reward estimates. Without it, a reward of 4.2 from a confident model and a 4.2 from an ensemble in total disagreement look identical to the policy. The ensemble is what makes that difference visible in the first place.
Three aggregation strategies and the tradeoffs each one makes
Mean aggregation is the obvious starting point: average the scalar outputs across ensemble members. It smooths out idiosyncratic errors from any one model and reduces variance in the final estimate relative to a single reward model. But it doesn't penalize disagreement at all. A prompt where all 10 models cluster tightly and a prompt where they're scattered across a wide range get treated identically once you've averaged. Mean aggregation reduces noise. It doesn't manage risk, and treating it as a safety mechanism is a mistake teams make more often than they should.
Lower Confidence Bound optimization, sometimes called worst-case optimization, actually does manage risk. The formula is mean minus beta times standard deviation, where beta is a hyperparameter controlling how risk-averse the policy training becomes. Instead of optimizing against the average score, the policy optimizes against this deliberately pessimistic lower bound, so high-variance predictions get explicitly punished rather than smoothed over. Coste et al. (2024) report that worst-case optimization and a related method called uncertainty-weighted optimization eliminate or sharply cut down overoptimization in both best-of-n sampling and PPO, with gains up to 70% over single-model optimization in best-of-n, holding whether or not the training labels carried extra noise. The tradeoff is real, though: crank beta too high and the policy starts under-rewarding responses that are genuinely good but just happen to be novel enough that the ensemble hasn't converged on them yet.
Uncertainty-weighted optimization comes in two flavors: an additive version that works like the LCB penalty, and a multiplicative "energy factor" version that scales the reward down as uncertainty rises rather than subtracting a fixed amount. The energy factor approach produces smoother gradients and does noticeably better on ambiguous prompts, because the update is modulated rather than eliminated.
All three strategies assume you already have a variance estimate worth trusting. That assumption isn't free, and it's exactly why building the ensemble efficiently, without wrecking the quality of the disagreement signal, turns out to be the real engineering problem.
LoRA-based ensembles and the efficiency-fidelity tradeoff
Ten full copies of a reward model at 90 GB is already a lot to carry for a 2B-parameter base. Scale that pattern up to the model sizes used in production RLHF pipelines, and training and serving a full ensemble stops being a matter of inconvenience. It becomes a matter of whether it fits on the hardware at all, since both compute and memory scale linearly with ensemble size.
LoRA-based ensembles sidestep this by keeping one frozen backbone shared across the whole ensemble and training multiple sets of low-rank adapter weights on top of it, each adapter diverging from the others just enough to generate a useful disagreement signal. Zhang et al. (2024), a collaboration spanning the MIT-IBM Watson AI Lab, Tsinghua, CMU, and UMass Amherst, tested this against full independent ensembles and against a cheaper alternative that only ensembles the last linear layer. The LoRA ensembles were evaluated on AlpacaEval and MT-Bench across both best-of-n and PPO settings, offering a more storage-efficient path to ensemble-based alignment benefits. The last-layer-only ensembles didn't hold up nearly as well, yielding only limited improvements and a weaker disagreement signal.
That diversity is the whole point, and it's also the failure mode to watch for. If the LoRA adapters converge toward similar solutions during training, the ensemble quietly collapses into something that behaves like a single model, and the variance estimate stops meaning anything even though it's still being computed. Zhai et al. (2023) addressed this directly with nuclear norm maximization on the LoRA parameter matrices, an explicit regularizer that pushes the adapters to span different directions in parameter space rather than drifting toward the same solution. Combined with KL divergence regularization and an uncertainty term, this produced better alignment and gold reward metrics under both best-of-n and PPO.
A related approach, WARM (Weight Averaged Reward Models, Ramé et al., ICML 2024), takes a different angle entirely: instead of ensembling predictions at inference time, it averages the weights of multiple trained reward models into one set of weights. That cuts storage costs sharply while keeping some of the diversity benefit, though it trades away the option to compute per-example variance directly, since you end up with one merged model rather than several to compare. Choosing between LoRA ensembles, last-layer ensembles, and weight averaging isn't just an infrastructure decision. It determines whether the variance signal feeding into worst-case or uncertainty-weighted optimization is trustworthy enough to build a training strategy on. Last-layer ensembles are the wrong default here, no matter how tempting the storage savings look on paper.
What ensembles still cannot do: the limits of variance as a safety signal
Ensembles help, and they don't solve the underlying problem, and that second half matters more than people give it credit for. Every member typically shares the same pretrained backbone and trains on the same underlying preference distribution, so any systematic blind spot in that distribution shows up in every member at once, not just one. This is a recognized failure mode: reward hacking can overoptimize the entire ensemble simultaneously, since a policy searching hard enough can find responses that fool every member at once rather than just slipping past one weak link.
That points to a sharper limitation. In genuinely out-of-distribution regions, without extra information about what those regions even look like, ensembles reduce reward hacking but don't close it off. Variance can stay low precisely in the cases where the whole ensemble is collectively wrong, because low variance measures agreement among the models, not agreement with the actual human preference the models are supposed to be approximating. A tight, confident-looking ensemble and a correct one are not the same claim, and conflating them is where a lot of RLHF pipelines quietly go wrong.
There's also the aleatoric floor. Adding more ensemble members shrinks variance toward zero as the ensemble grows, regardless of whether the underlying preference labels are genuinely, irreducibly noisy. An ensemble can look more and more certain about a question that human annotators would never agree on no matter how many were asked. And the compute overhead doesn't disappear just because LoRA made the storage manageable: online RLHF still needs a forward pass through every ensemble member for every candidate response, a cost that compounds over the length of an RL run rather than showing up once during training.
Variance-based uncertainty is necessary and not sufficient, plainly. It needs to sit alongside calibration methods, out-of-distribution detection, or some other measure of distributional coverage to close the gap it leaves open on its own.
How recent work extends uncertainty estimation beyond variance-based ensembles
The frontier here isn't abandoning ensembles. It's finding ways to get calibrated uncertainty without paying the full ensemble tax, or catching failure modes ensembles miss outright.
Pan et al. (arXiv:2606.19818), from Zhejiang University and Xiaohongshu, propose Uncertainty-Aware Reward Modeling (UARM), which equips a single reward model with calibrated uncertainty through quantile-based conformal prediction, a distribution-free method that produces coverage guarantees adapted to give a sample-specific reliability signal. UARM then reweights GRPO's advantage calculation using heteroscedastic variance decomposition, going straight at the amplification problem where group-based methods hand equal weight to reliable and unreliable reward estimates alike. Tested across HelpSteer, UltraFeedback, and PKU-SafeRLHF, it beat standard GRPO and uncertainty-agnostic baselines on calibration and reward hacking. No ensemble required. Just one model architecture with the uncertainty built in.
Bayesian reward models take a more principled route to the same destination. Treat the reward parameters as random variables, approximate the posterior over the fine-tuned weights (often restricted to LoRA adapters to keep it tractable) as a Gaussian centered on the MAP solution, and reward prediction at inference becomes a distribution instead of a single number. Yang et al. (2024) show this distribution can be used to penalize high-variance responses directly in both best-of-n and RL settings. It's a cleaner theoretical foundation than ensemble variance, but it only scales if the Bayesian update is restricted to a small enough parameter set to keep the posterior approximation tractable.
Energy-Based Reward Models, from Lochab et al. in April 2025, take existing reward models and retrofit them with a tractable distribution over possible rewards rather than a single point estimate, capturing uncertainty and cutting down reward hacking without needing ensemble infrastructure at all.
The uncertainty problem has also spread past text. InternLM-XComposer2.5-Reward (Zang et al., January 2025) and Skywork-VL Reward (Wang et al., May 2025) apply ensemble-based reward fusion to image and video-language benchmarks, supporting test-time selection and data cleaning in multimodal alignment. The same underlying problem keeps showing up wherever a model has to score something as varied and noisy as human preference.
And for reasoning models specifically, process reward models are starting to estimate uncertainty over individual steps of a chain of thought rather than just the final answer. CoT Entropy aggregates the entropy of the rationales generated at each verification step, turning uncertainty into something granular enough to catch a shaky step buried inside an otherwise confident-looking answer.
What connects all of it: uncertainty in reward modeling has gone from an afterthought nobody measured to a quantity built directly into the architecture, whether through ensembles, conformal prediction, Bayesian posteriors, or energy-based distributions. The through-line is the same one that opened this piece. A reward model that can't say how sure it is will eventually get exploited by the very policy it's meant to guide.


