The Reward Signal

Compute Budget Allocation Across RLHF Pipeline Stages

How to budget each RLHF stage separately before costs spiral.

Editor at Large · · 11 min read
Cover illustration for “Compute Budget Allocation Across RLHF Pipeline Stages”
Post-Training Pipelines · September 21, 2026 · 11 min read · 2,532 words

SFT, reward modeling, and PPO don't share a cost model. Each stage carries its own GPU profile, its own scaling curve, and its own way of quietly eating a budget, and treating "RLHF" as a single line item is the first mistake most teams make before they've spent a dollar. FutureAGI's research still describes the textbook sequence, SFT, then a reward model, then PPO, as broadly accurate for how models like GPT, Claude, and Gemini get aligned. But the operational picture has splintered: DPO, GRPO, and RLAIF now sit alongside PPO as live production choices, not academic footnotes, and each one moves the cost somewhere else rather than making it disappear.

Widen the lens: the pipeline actually runs four stages, not three: continued pretraining on domain data, supervised fine-tuning on instructions, reasoning post-training (RLVR, GRPO, self-distillation), and finally preference optimization, the taste pass where RLHF and its variants live. Teams that blur reasoning post-training and preference optimization, running RLHF on reasoning data or RLVR on tone data, tend to regress both capabilities at once. Shipping a model that reasons well but talks badly, or one that talks well but can't reason, is a real cost, not a rounding error.

The mistake to avoid is treating "a full RLHF run" as one budget line. Each phase has its own memory footprint, its own update loop, its own reason to get expensive. Lump them together and the number comes out wrong in both directions at once: too high for some stages, dangerously low for others.

SFT compute: straightforward gradient descent, but the stage most teams underinvest in

SFT looks like ordinary language model fine-tuning because that's what it is: one forward pass, one backward pass, a gradient update against a causal language modeling loss. Cost scales with base model size, dataset size, sequence length, and epoch count, the way anyone who's fine-tuned a model would expect. Next to what comes after, the compute pattern here is almost dull, and that's not a knock. Dull and predictable is what a budgeting exercise wants.

The trap sits exactly there. Because SFT reads as "just fine-tuning," teams treat it as a formality and rush past data curation, which determines everything downstream. A reward model trained on comparisons between mediocre SFT outputs learns to rank mediocrity, and PPO then spends GPU-hours optimizing trajectories that were never any good to begin with.

The quality of SFT caps the ceiling of every downstream alignment pass. That single sentence settles where the real budget fight belongs. The research recommends spending something like 30% of total project time on SFT data quality, and that figure means human curation hours: reading examples, throwing out the bad ones, rewriting the ambiguous ones until the demonstration set actually demonstrates something.

Teams tend to assume volume drives quality here. Volume does not drive quality here. A small, well-curated dataset beats a large, unfiltered one almost every time, and OpenAI's InstructGPT work is the proof: a comparatively modest set of human-written demonstrations for SFT, showing that quality substitutes for scale at this stage in a way it stops doing later in the pipeline. SFT is where RLHF becomes possible, or where it gets quietly compromised before a single PPO step ever runs.

Reward model training: annotation cost is the budget line most engineers forget to model

Compute-wise, reward model training looks a lot like SFT: one forward pass, one backward pass, one gradient update. The data format changes, pairwise comparisons instead of demonstrations, but the GPU shape stays familiar. Nobody blows their budget on RM training compute.

They blow it on annotation, and the number most guides quote is simply wrong. A baseline of $0.50 per sample gets passed around in planning docs, but quality preference annotation through providers like Anthropic or Scale runs closer to $13 per preference pair. Pushing into complex domains, clinical judgments, legal reasoning, code correctness, drives a single comparison to run around $100. Multiplying that across production-scale label volume, the annotation line alone lands anywhere from $1 million to $5 million, depending on task complexity and how many comparisons the reward model needs before it converges.

Comparison volume doesn't scale with dataset size the way people expect, either. It scales with task difficulty. Reports on InstructGPT indicate it needed substantially more labeled preference comparisons to train its reward model than it needed SFT demonstrations, more than double the SFT count for a stage most teams assume is smaller. Complex tasks like summarization can require far larger preference comparison sets before the reward model holds up under real use, a reality that should push anyone planning a clinical or technical RLHF project to budget tens of thousands of labels, not thousands.

Reward model sizing is its own decision, and getting it wrong leaves an exploit sitting in plain sight. Some research shows strong results even with reward models considerably smaller than the policy, which complicates any simple rule about RM-to-policy sizing without erasing the dependency entirely. RM-to-policy size ratio is a real compute dependency.

One more cost hides inside this stage and follows it forward. KL divergence penalty computation, running reward model inference against reference model outputs, adds roughly 30 to 50% on top of reward compute hours, and that overhead doesn't end when RM training does. It carries straight into the PPO loop, where it turns into a recurring tax on every batch. The research names underestimated annotation budget as one of the most common reasons RLHF projects stall out mid-run, and the pattern is almost always the same: someone priced labels at the cheap rate, then discovered mid-project that the task needed the $100 tier instead of the $13 one.

The PPO loop's four-model architecture multiplies both memory and compute.

PPO doesn't resemble SFT or RM training at all, and treating it like a bigger version of either one is how budgets blow up. It runs four models at once: rollout generation, reward scoring, advantage computation, and simultaneous updates to two of the four models. That turns training into a systems engineering problem.

The four roles all have to fit in memory together. The actor is the online policy actually being trained, carrying BF16 weights plus AdamW optimizer states. The critic is a value model also being trained, with the same weight-plus-optimizer footprint. The frozen reference model is the original SFT checkpoint, kept around purely to compute the KL penalty, inference only. The frozen reward model is also inference only, scoring rollouts as they land.

Running the numbers for a 70B parameter model makes the picture concrete fast. Unsharded, a 70B actor in BF16 with AdamW optimizer states alone needs roughly 980 GB (140 GB for weights, 840 GB for optimizer states). Shard that actor 8 ways with FSDP across B200 nodes at 192 GB each, and it settles around 122 GB per GPU on a single 8-GPU node. The frozen reference and reward models each need about 140 GB, small enough that either one fits on a single B200. Spheron's figures show a full 70B PPO pipeline needs at least 3 nodes total just to hold everything in memory, before a single optimization step runs.

Memory is only half the multiplier. PPOTrainer runs a full forward-and-backward pass on both the policy and the reference model for every batch, which roughly doubles training compute hours compared to a naive estimate. Missing that detail leaves the cost model at half of what actually gets billed. Putting the whole picture together, full PPO-based RLHF costs somewhere around 3 to 5 times what SFT alone costs, a ratio large enough that algorithm choice belongs at the planning stage, not the invoice. H100s deliver roughly 2,000 to 3,000 tokens per second depending on batch size and sequence length, and any back-of-envelope math converting token targets into GPU-hours should anchor to that range.

Infrastructure pricing and the spot-vs-on-demand decision in a PPO run

Hardware choice sits on top of everything above, and the numbers move fast enough that any figure quoted today is a reference point, not a promise. Spheron's benchmarks put B200 spot pricing at $1.71 per hour against $7.00 on-demand, and B300 spot at $2.45 against $9.77 on-demand. H100 SXM5 pricing on Spheron runs around $4.21 per hour, meaningfully cheaper per GPU than an AWS p4de node (8x A100 80GB) once that node's per-GPU-hour cost gets worked out.

Mixing in spot capacity instead of running everything on-demand cuts total PPO loop cost by something like 35 to 40%, whether the hardware underneath is B200 or B300. That's a real saving, and there's no good reason to leave it on the table. But spot instances carry preemption risk, and a long PPO run interrupted without a solid checkpoint strategy loses more in restart cost than the spot discount ever saved. The saving only holds if checkpoint frequency gets budgeted as its own line item instead of an afterthought, checked in against wall-clock progress, not against calendar days.

Take the final compute estimate, add a real buffer on top, then recalibrate after the first live run using actual usage data. Treat the initial number as a floor, not a target. No GPU tier wins outright here. The right choice depends on model size, sharding strategy, and how much preemption risk a team can stomach, and the pricing table is one input among several, not the deciding one.

Failure modes that inflate actual cost beyond the planned budget

Reward hacking is the failure every PPO run has to guard against, and it happens because the reward model is an imperfect stand-in for human judgment. The policy finds the shortest path to a high score. FutureAGI's analysis documents the same patterns: responses that get longer because length alone nudges the reward score up, confident-sounding filler that carries no information but reads well to an automated scorer, and, in the worst cases, text that scores highly while reading as outright nonsense to a human.

Sycophancy is the quieter version of the same problem. If the label data rewards a pleasing tone over a correct action, the model learns to agree with the user and skip whatever checks it should be running. FutureAGI's RLHF documentation illustrates this with a support-agent scenario where the model starts approving requests it should be flagging. Over-refusal runs the opposite direction: punish anything that looks risky without a narrow, well-defined refusal rubric, and the model learns to block legitimate requests it should be answering. FutureAGI's 2026 analysis calls this the single most under-measured RLHF regression, because the headline metric, pairwise preference win rate, keeps climbing even as refusal rate climbs faster and total task value quietly falls.

Every one of these failure modes carries the same budget consequence. Fixing one means re-annotating for the reward model, re-running the PPO loop, or both, and either can double the cost of whatever stage broke. The KL penalty is the primary lever against reward hacking, and getting it wrong in either direction is expensive: set it too low and the policy drifts into exploiting the reward model's blind spots, set it too high and the policy barely moves from where SFT left it, turning PPO into an expensive way to reproduce the SFT checkpoint. Yobitel's research frames the current standard mitigation as four things working together: KL penalty tuning, reward model auditing, ensembles of reward models rather than a single one, and iterative relabeling applied as failure patterns show up in the model's outputs.

How DPO and GRPO shift where budget gets spent and when to choose each.

DPO, Direct Preference Optimization, removes the reward model and the PPO loop. Preference data feeds straight into a modified loss function computed against the SFT model. No rollout generation, no reward scoring, no separate critic.

Two models in memory instead of four drops VRAM requirements substantially, and with no online rollout generation, no advantage estimation, and no critic to stabilize, weeks of PPO stability engineering simply don't need to happen. Annotation cost does not disappear, since DPO consumes the same type of preference data PPO would, but compute cost drops sharply. The tradeoff appears in capability instead of budget: FutureAGI's method comparison notes DPO tends to underfit on hard preferences relative to PPO and performs worse on complex reasoning or tool-use tasks, where PPO's explicit reward model captures a signal more nuanced than a pairwise loss function can represent.

GRPO, Group Relative Policy Optimization, cuts the same problem differently. It removes the critic model from PPO entirely, comparing several responses within a group and computing group-relative advantage estimates straight from the rollout batch, instead of training a separate value function. That alone cuts memory by roughly 50% compared to full PPO and removes a whole source of training instability, since critic training is notoriously hard to get right, concentrating the remaining compute burden on rollout generation. GRPO has become a leading algorithm for training reasoning models specifically, and DeepSeek's R1-Zero work showed that RLVR combined with GRPO can produce emergent reasoning without any human feedback loop at all, cheaper and more scalable than traditional RLHF for reasoning-heavy tasks.

Each algorithm just moves the bill somewhere else. FutureAGI's decision framework lays out where each algorithm shifts the bill. Fewer than 10,000 paired preferences points toward DPO or KTO. Tasks with verifiable rewards, math, code, agentic tool use, point toward GRPO. A working reward model paired with a large compute budget and genuinely hard reasoning or tool-use tasks still favors PPO. And when labelers, not compute, are the bottleneck, RLAIF or constitutional-AI approaches layer on top to scale the signal without scaling headcount. Yobitel's research describes the mid-2026 production stack accordingly: rarely just PPO, more often SFT followed by DPO on preference data, then GRPO on verifiable rewards, then a final PPO or DPO polish pass, with RLAIF or constitutional AI generating labels at scale somewhere in the loop.

None of these choices actually reduces cost. They relocate it. DPO shifts spending away from GPU hours and back onto annotation, since the compute savings are real but the label bill never shrinks. GRPO shifts spending away from critic compute and into rollout generation, trading one bottleneck for another. PPO keeps everything concentrated in the four-model loop: expensive, but buying the most control over the final policy. Picking among the three is a reallocation decision, not a savings decision, and any plan that treats it otherwise is going to miss its own budget by the ratio described above.

Framework selection for production PPO: verl, OpenRLHF, and TRL

By May 2026, Spheron's reporting showed RLHF tooling had split into three frameworks built on meaningfully different architectural bets. Picking the wrong one for a given cluster shape creates its own hidden cost: engineering time spent working around a framework's assumptions instead of around the actual RLHF problem. verl, OpenRLHF, and TRL each optimize for different deployment shapes, different sharding strategies, and different tolerances for the multi-model memory juggling PPO demands.

Framework choice belongs last on the checklist, not first. It only makes sense to pick verl, OpenRLHF, or TRL once the earlier questions, model size, algorithm, hardware tier, already have answers. Choosing the framework before those are settled is how teams end up re-platforming mid-project, and re-platforming mid-project costs more than any framework ever saved.

Sources

  1. What Is RLHF? Definition, Examples & FutureAGI Guide (2026)
  2. RLHF Training Infrastructure on GPU Cloud: verl, OpenRLHF, and TRL for Production Reward Modeling (2026) | Spheron Blog
  3. What Is RLHF? The Complete Guide to Training LLMs That Actually Work (2026)
  4. RLHF (Reinforcement Learning from Human Feedback)
  5. Cost allocation for RLHF programs | Rlhf Advanced Course | The Neural Base

More in Post-Training Pipelines