Iterative Online RLHF With Continuously Updated Preference Data
Continuously updating preference data beats training reward models once on static data.

Iterative online RLHF keeps updating its own preference data as the model improves, instead of training once on a fixed dataset and calling it done. The distinction sounds small. It isn't: it changes what a reward model can actually learn, and it changes the cost structure of the entire alignment pipeline.
What the iterative online loop does
Traditional RLHF runs in a straight line: collect a fixed set of human preference judgments, train a reward model on that set, fine-tune the policy against it, and stop. A reward model trained once on static data has no way to generalize to prompts or response styles it never saw, and the trouble appears when the policy improves through fine-tuning and starts producing those unseen styles. The reward model ends up scoring outputs from a distribution it was never trained on. Researchers have stated that reward models trained on fixed data "struggle to generalize to out-of-distribution samples, constraining the effectiveness of the aligned models." The gap between what the reward model knows and what the policy now does doesn't hold steady, either. It widens with every fine-tuning step, because the policy keeps moving and the reward model stays put.
The iterative loop replaces that straight line with a circle. Deploy the current policy, collect fresh preference data from its actual generations, update the reward model on that new data, re-optimize the policy, repeat. The difference from offline RLHF is that data collection happens on-policy: the model being trained is the same model producing the outputs that get judged. The preference data reflects wherever the model's capability frontier sits right now, not wherever it sat months ago when the original dataset got built. Anyone defending the offline approach on cost grounds alone is arguing for a method that's structurally blind to its own model's progress.
That's more than intuition. Song et al. (arXiv:2406.01462) showed online RL methods only need a weaker partial coverage condition to perform well, while offline methods need global coverage of the response space in advance. Online methods don't need to have seen everything ahead of time; they just need to keep pace with what the policy is doing right now. Empirical work across multiple research groups backs this up: online iterative variants of direct preference learning beat their offline counterparts consistently, not occasionally.
How fresh preference data changes reward model learning
A reward model takes a prompt and two candidate responses and outputs a scalar saying which one wins. Every downstream training step depends on that number being right. If the reward model is guessing in regions it's uncertain about, the whole optimization process inherits that uncertainty, and no amount of downstream tuning fixes a bad scoring signal at the source.
On-policy data fixes this by feeding the reward model examples drawn from exactly the distribution the policy is generating right now. Accurate scoring matters most there, for the next gradient update. Tan et al. Preference data split into three subsets and used sequentially across training iterations produced LC win-rate gains of 13.74%, 3.22%, and 3.44% across the three rounds, as Tan et al. at Notre Dame and Amazon (arXiv:2506.04463) showed. The pattern says something specific. Early iterations catch the biggest distributional gaps, the widest space between what the old reward model knew and what the improving policy was actually doing. Later rounds work at the margins, refining rather than correcting. That's a built-in feature of the method, not a flaw. It's a useful signal, since the diminishing returns tell a team exactly when another round stops paying for itself.
Active versus passive data collection during online training
Not all on-policy data collection looks the same. Either the policy wanders on its own, or someone points it somewhere on purpose; if the policy wanders on its own, the reward model gets scored against off-target data and the resulting gradient updates drift from what the policy is actually doing, while deliberate targeting keeps the reward signal aligned with the policy's current outputs.
Passive exploration lets the natural randomness in an LLM's outputs generate variety on its own. It's the easier path: no added machinery, no extra objective to tune. But it carries a structural weakness. A policy gravitates toward its own high-probability outputs, the responses it's already confident about, so the ambiguous cases, the comparisons where the reward model is genuinely unsure, often never get sampled. Passive collection systematically under-explores exactly the region where more data would help the most.
Active exploration corrects for this by building an exploration bonus directly into the data-collection objective, steering queries toward high-uncertainty regions or areas where the policy has the most room to improve. Most active methods bolt this bonus onto a DPO-style objective. The tradeoff is complexity for coverage: active methods take more engineering to set up, but they stop the loop from just reinforcing what the model already believes about itself. For any team running more than a couple of iterations, passive collection is the wrong default. It looks cheaper on day one and gets expensive later, when the reward model turns out to be blind precisely where it needed to see clearly.
The computational cost that grows with every iteration, and a proposed fix
Standard practice adds each round's new data to a growing historical pile, then retrains the reward model from scratch over the entire accumulated dataset, every iteration. That's the loop's biggest practical problem, and it's a self-inflicted one.
The cost scales linearly, no better. Ten rounds of training cost roughly ten times what one round costs, in compute and in storage both. The method that's statistically correct to run fully turns out to be the one that punishes a team financially for running it fully. That's why most practitioners cap their pipelines at one or two iterations. The budget runs out before the benefit does, and calling that a principled stopping point gives it more credit than it deserves.
Li, Qian, Zhao, and Zhou's method offers a way out. It formalizes RLHF as a contextual preference bandit problem and swaps standard maximum likelihood estimation for online mirror descent with a tailored local norm. The practical upshot: no need to store historical data at all, and reward model updates run in constant time, O(1), no matter how many rounds have already happened. That's a real departure from prior approaches needing O(t) storage and compute time as rounds pile up, and the paper reports this holds without giving up the statistical guarantees older methods offered. The method was validated on Llama-3-8B-Instruct and Qwen2.5-7B-Instruct against the Ultrafeedback and Mixture2 datasets, and it covers passive collection, active collection, and deployment-time adaptation. If this kind of approach becomes standard, the case for stopping at one or two iterations mostly disappears.
Algorithm choice, PPO, DPO, GRPO, and its interaction with the iterative loop
The optimization algorithm shapes how the loop behaves, and none of the three major options substitute cleanly for one another.
PPO keeps an explicit reward model in the loop, so each iteration updates the reward model and the policy separately. That's more moving parts, but the payoff is a reward signal a team can actually inspect, rather than a black box it has to trust. A 2024 ICML study asking whether DPO beats PPO for LLM alignment found PPO still wins by roughly 2.5% on math benchmarks and roughly 1.2% on general benchmarks when data quality is held constant. That's the real reason PPO with a KL penalty remains the default for teams running production iterative pipelines, whatever DPO has going for it on simplicity.
DPO skips the explicit reward model and trains directly on preference pairs, which is cheaper to implement, but it was built for the offline setting from the start. Online variants exist, and research from the University of Washington (arXiv:2604.17207) found they can hit O(1) cumulative regret under a temperature-zero, decision-centric regret criterion, which explains why online DPO performs well in practice even where other regret criteria paint a bleaker picture. Even so, DPO carries real weaknesses once it's inside an iterative loop. Research has documented an over-optimization problem where DPO's performance degrades over extended training rather than holding steady. It's also sensitive to distribution shift between the base model's original outputs and the preference data it trains on, which leaves it prone to favoring out-of-distribution responses it has no business trusting.
GRPO drops the reference model altogether, which sidesteps the KL-divergence constraint both PPO and DPO have to manage. Some practitioners like it for exactly that reason. But adopting it means rewriting the training loop, not swapping a hyperparameter, and that switching cost is why most teams still default to PPO with a KL penalty instead.
A middle path exists too. Bose et al. proposed Hybrid Preference Optimization, combining a DPO loss trained on offline data with an exploration-aware regularizer applied to online data. The paper proves this hybrid converges faster than pure offline DPO or pure online exploration running alone. The choice between algorithms doesn't have to be all-or-nothing.
Where human preference data comes from
Human preference data starts with prompts going out to annotators, who compare pairs of model responses and record which one they prefer. Everything downstream depends on how consistent those annotators are, how well-written their instructions are, and how well the annotator pool actually covers the domain the model needs to operate in.
None of that comes cheap. Getting real value out of human preference data means running iterative training cycles, spending sums that run into the hundreds of thousands or millions of dollars, writing detailed annotation instructions, and either working through a data foundry business or hiring annotators directly in real numbers. It's a genuine operational undertaking.
There's an opacity problem stacked on top of the cost. As of 2026, no open model ships with fully open human preference data alongside a full account of how that data got collected. The largest recent public releases come from NVIDIA's Nemotron team, the HelpSteer line: HelpSteer2 and HelpSteer3-Preference. Even these don't expose the full collection methodology behind them. Given the cost and the opacity, it's no surprise that a lot of teams skip human preference data entirely and lean on AI feedback, off-the-shelf reward models, or implicit preference signals instead. That's a rational tradeoff for most teams, not a shortcut they should feel bad about, but it does mean the industry is training on proxies for human judgment far more often than the marketing around "human feedback" suggests.
Requirements for the iterative loop to work at production scale
Running this loop for real, at production scale, takes five pieces working together, and each one carries its own failure mode.
An on-policy generation pipeline has to exist: the current model needs to be queryable during training so it can produce the outputs that get evaluated. A preference collection layer has to sit behind it, whether that's human annotators, AI feedback, or implicit signals, each with its own latency and cost tradeoffs. A reward model update mechanism has to run, ideally one that operates in constant time rather than scaling linearly with each new round. Distribution shift monitoring has to be in place, because without it, reward hacking and overoptimization occur quietly during training and only become visible in performance metrics once performance has already slipped. And a policy optimization algorithm has to get chosen deliberately, PPO, online DPO, GRPO, or a hybrid, since each one carries its own compute profile and its own stability quirks.
The loop's defining property cuts both ways. On-policy data keeps the reward model calibrated to what the model is doing right now, which is the entire point of running it this way. But that same dependency is a liability: a reward model that drifts off-calibration doesn't just spoil one round, it corrupts the training signal for every round that follows. Errors compound here. They don't average out the way they might in a system built with more independent, redundant checks. Active exploration is the main lever available to fight this, giving practitioners a way to steer data collection toward the regions that actually need attention, instead of letting the policy's existing habits decide what gets labeled next.
Feasibility comes down to compute, in the end. Methods like the one-pass, constant-time reward modeling approach change what's possible by removing the linear cost penalty iteration count used to carry. Without something like that in place, teams face a blunt tradeoff: run enough iterations to get the statistical benefit the loop promises, or stay inside budget. Right now, most can't do both, and pretending otherwise is how projects run out of runway two rounds before the gains would have shown up.

Sources
- NeurIPS Poster Provably Efficient Online RLHF with One-Pass Reward Modeling
- Aligning Large Language Models with Implicit Preferences from User-Generated Content
- arxiv.org
- Demystifying the unreasonable effectiveness of online alignment methods
- Provably Efficient Online RLHF with One-Pass Reward Modeling
- shivu-agr.medium.com


