The Reward Signal
RLHF ToolingLong read

Ray RLlib vs Custom Training Loops for LLM RL Fine-Tuning

RLlib's algorithm layer wasn't designed for billion-parameter model scaling.

Senior Writer · · 11 min read
Cover illustration for “Ray RLlib vs Custom Training Loops for LLM RL Fine-Tuning”
RLHF Tooling · October 4, 2026 · 11 min read · 2,431 words

Post-training has become the main site of work for turning open-weight language models into useful, task-ready systems, and reinforcement learning, specifically PPO, GRPO, and DPO, is at the center of that work. The thesis here is straightforward: RLlib and the purpose-built loops now used across the industry share the same algorithmic vocabulary but diverge completely once the conversation turns to infrastructure, and that divergence is the real decision teams face.

Why LLM Post-Training Makes RL Infrastructure a First-Class Problem

A single iteration of standard RL training for a language model requires coordinating up to four distinct models: Actor, Critic, Reference, and Reward. Each of these models has to move through generation, inference, and training stages in sequence, and the system running the show has to schedule all of it on the same hardware or on hardware sitting close enough together that moving weights and activations between stages doesn't become its own tax. That's a scheduling and memory-placement problem before it's anything else.

The algorithms themselves are not the hard part. PPO and GRPO carry over from classical reinforcement learning with little modification, and teams that have run policy optimization in robotics or games will recognize the math immediately. What classical RL infrastructure was never built for is generating text, at scale, from a model with billions of parameters, fast enough that the rest of the training loop doesn't sit idle waiting on it. That single fact, efficient generation from a multi-billion-parameter policy, is the bottleneck that defines whether an RL infrastructure choice will hold up under real post-training workloads or buckle under them.

So look at what that choice costs in the standard recipe for training reasoning models. The model gets domain knowledge first from supervised fine-tuning, and then reinforcement learning sharpens the policy toward reasoning strategies that earn higher reward. Research from ETH Zürich and Allen AI found that a correctly implemented SFT-then-RL pipeline beats every published mixed-policy interleaving method it was tested against. The same research traced the apparent advantage of those alternative methods to silent bugs in widely used frameworks, including TRL, OpenRLHF, and Llama-Factory, bugs that had been quietly suppressing the SFT baseline and making the newer methods look better than they were. Framework correctness turns out to carry as much weight as algorithm choice: a bug buried in a distributed training configuration can invalidate an entire line of published comparisons.

But infrastructure stability matters for another reason: it isn't just about getting numbers right. Research from Polytechnique Montréal and Mila shows that RL fine-tuning restores most of the out-of-distribution generalization a model loses during supervised fine-tuning, functioning less as pure discovery of new behavior and more as repair of damage SFT itself caused. That restorative effect has a limit: RL cannot rescue a model that SFT has already pushed into severe overfitting. The practical consequence is that the infrastructure running RL has to be stable enough to execute to sufficient depth, because a system prone to collapse or to the kind of silent degradation the ETH and Allen AI researchers found undercuts the very recovery RL is supposed to deliver.

What RLlib Was Built to Do

RLlib is a distributed policy-optimization framework built on top of Ray clusters, and it has earned its reputation honestly. It orchestrates multi-agent environments, it scales data-parallel training across many workers, and it manages RL dataflows that get complicated fast once a simulation involves more than a handful of interacting agents. For classical control tasks and simulation-heavy environments, these are real strengths, and no marketing claims.

Its distributed design follows a hierarchical but logically centralized control model. A single driver program delegates sub-tasks to worker processes, and because Ray's own distributed scheduler handles resource allocation, no single scheduling bottleneck forms even as the cluster grows. Each worker, implemented as a Ray actor, can spawn additional actors of its own, and that is part of what lets RLlib manage genuinely complex multi-agent setups without the driver becoming a traffic jam.

The constraint shows up once model size enters the picture. Scaling Learner actors in RLlib happens through distributed data parallel, and that is currently the only mechanism available for it. DDP-only scaling means the model has to fit on a single GPU, which rules out the multi-billion-parameter models that sit at the center of modern post-training work. This limitation appears in Ray's own documentation rather than in outside criticism of the framework, and the Ray team is actively building toward training very large models, including an RLHF prototype already in development. That a major framework's own maintainers are working to close this gap confirms both that the gap is real and that it hasn't closed yet in the production version teams can use today. RLlib's primitives were built around data-parallel training of networks in the hundreds-of-megabytes range, a different regime entirely from the tens of gigabytes current LLMs require.

A second mismatch sits alongside the size constraint. RLlib does not natively separate rollout generation from gradient updates, and that separation is exactly the architectural move that LLM-native frameworks rely on to keep GPUs busy while generation lengths vary wildly from one sample to the next. In a synchronous pipeline, the slowest chain-of-thought generation in a batch holds up the entire iteration, and purpose-built frameworks solve this with asynchronous designs that let each engine run at its own pace. RLlib's current design doesn't reach that async separation at the scale LLM training demands, which raises a reasonable objection: if Ray itself handles scheduling and fault tolerance so well, why not just build on Ray directly and skip RLlib's limitations?

Just Use Ray" vs. Using RLlib

That objection deserves a direct answer, because it rests on a distinction people collapse far too often. Ray, the underlying distributed runtime, and RLlib, the algorithm engine built on top of it, are two separable decisions. The leading purpose-built LLM RL frameworks have already made that separation explicit, and understanding how they did it clarifies what the real choice in front of a team actually is.

OpenRLHF is built on Ray, vLLM, DeepSpeed, and Hugging Face Transformers. It uses Ray for distributed scheduling and for placing models across hardware, but it bypasses RLlib's policy optimization loop entirely, running its own PPO and GRPO implementation on top of DeepSpeed ZeRO instead. A team running OpenRLHF is using Ray in every meaningful infrastructural sense without using RLlib. veRL follows the same pattern: Ray coordinates the cluster, while veRL's own HybridEngine handles the training-and-inference separation that RLlib's architecture does not provide.

Once that separation is visible, the real question comes into focus. It isn't "Ray or custom loops," since nearly every serious option runs on Ray somewhere in its stack, and the actual choice is between RLlib's algorithm layer and the algorithm layers built by purpose-built LLM RL stacks, which at the layer that matters for LLM workloads share almost no infrastructure in common. If a team already runs Ray for data processing or for model serving, it can adopt OpenRLHF or veRL without standing up a second orchestration layer, so it picks up LLM-native RL capability on infrastructure it has already paid for. But if RLlib specifically is the goal, a team with no existing Ray deployment faces a real setup cost, and that cost is hard to justify when the purpose-built alternatives also run on Ray and give you the LLM-specific machinery RLlib currently lacks.

How Purpose-Built LLM RL Stacks Solve the Generation Bottleneck

The defining architectural choice across purpose-built LLM RL frameworks is splitting rollout generation, the fast inference step that produces text, away from gradient updates, the training step that consumes it, and running the two on engines specialized for each job. Sometimes that split happens on the same GPUs; other times it happens across separate hardware pools. The specifics differ by framework, but the underlying move is consistent, and a handful of concrete implementations show what it looks like in practice.

veRL, built by ByteDance, handles this through its HybridEngine, which swaps weight layout in place between FSDP training shards and vLLM inference shards on the same GPUs. So you remove the memory overhead and latency cost of running training and inference as two permanently separate processes. OpenRLHF takes a different route to a similar outcome, integrating vLLM with automatic tensor parallelism and pipeline parallelism to push generation throughput higher, and layering an asynchronous design on top so each engine runs at its own pace rather than waiting in lockstep for the slowest sample in a batch. That asynchrony is what keeps hardware utilized even when chain-of-thought generations vary enormously in length within the same batch.

ROLL, introduced by Alibaba in May 2025, takes the architectural ambition further. It uses a single-controller design with a Rollout Scheduler that manages each sample's lifecycle individually during generation, combined with a resource-assignment scheme that lets different models in the pipeline draw on different hardware allocations at different stages. ROLL's team demonstrated this at serious scale: an in-house run training a mixture-of-experts model with more than 200 billion total parameters, across thousands of GPUs, running continuously for over two weeks.

DeepSeek-R1 makes the same underlying argument in a different way. DeepSeek-R1-Zero was trained with Group Relative Policy Optimization, a method built to simplify training and cut the resource cost PPO carries, and the team trained it with no neural reward model. Neural reward models introduce the risk of reward hacking at scale and demand additional retraining resources that complicate an already complex pipeline, so leaving one out was a deliberate simplification of the infrastructure, not merely an algorithmic preference. Arcee's Trinity model, released in 2026, shows the disaggregated pattern operating in a production setting: after an SFT stage, Trinity ran a short RL phase using prime-rl's asynchronous setup, where vLLM-backed workers generated rollouts while a separate distributed trainer applied updates using FSDP2.

Asynchronous training, where rollout generation and model updates overlap rather than running one after the other, is now supported across veRL, OpenRLHF, ROLL, and other frameworks, each handling staleness management and weight synchronization somewhat differently. This split between synchronous simplicity and asynchronous throughput is the defining engineering decision of the current generation of LLM RL frameworks. Synchronous pipelines are easier to reason about but waste GPU time waiting on slow generations; asynchronous pipelines keep hardware busier but require careful handling to avoid training on rollouts that have gone stale relative to the current policy. Purpose-built frameworks force a team to make this choice explicitly. RLlib's current design doesn't put a team in a position to make that choice at LLM scale.

The Algorithm Layer Is Narrower Than It Looks

The methods driving almost all current LLM post-training, PPO, GRPO, and critic-free policy gradients using Monte Carlo credit assignment, cover a narrow slice of what reinforcement learning as a field actually offers. A survey mapping the LLM RL literature onto the classical RL taxonomy found that value-based methods, off-policy actor-critic training, and bootstrapping-based credit assignment remain largely unexplored for language models, even though classical control and robotics research has had well-established counterparts for decades. The survey organizes the field around three layers: MDP creation (reward, state, action, termination, discount), exploration (temperature, entropy, curriculum, tree search), and learning (model-free versus model-based, value versus policy versus actor-critic, on-policy versus off-policy, credit assignment). Laid out this way, the taxonomy makes visible how many of these dimensions current practice simply skips over.

Part of that narrowness traces directly back to infrastructure. GRPO and PPO scaled first because they paired naturally with vLLM-backed generation, and the frameworks that made that pairing fast and reliable ended up defining the practical boundary of what teams could run, not just what researchers found theoretically interesting. Reward design shows the same pattern: rule-based and verifiable rewards, like unit tests for code or exact-match correctness for math, dominate current practice because neural reward models carry a documented risk of reward hacking at scale and demand retraining resources that complicate the pipeline, a tradeoff the DeepSeek-R1 team stated directly. That leaves a real decision for teams to make depending on where their algorithmic needs sit. A team whose work fits comfortably inside PPO, GRPO, and verifiable rewards can choose infrastructure on operational grounds alone, since the current generation of purpose-built frameworks already serves that band well. But if a team expects to explore off-policy methods, bootstrapped credit assignment, or multi-agent interaction down the line, it needs a framework whose abstractions don't quietly assume the narrow band is all there is. TRL fits teams for whom getting the training objective exactly right matters more than training speed, as long as prioritizing correctness over speed still trades off against it, provided the work stays within the established set of GRPO, DPO, PPO, and RLOO.

Training instability is a shared risk regardless of framework, and infrastructure choices determine how recoverable it is

Gradient explosions and outright training collapse are documented risks in any prolonged LLM RL run, occurring whether a team is using RLlib, a purpose-built framework, or a fully custom training loop built in-house. No framework choice makes a team immune to this category of failure.

The more dangerous version of instability produces no warning signs before it happens. The ETH Zürich and Allen AI research uncovered a CPU-offloaded optimizer bug in DeepSpeed that silently dropped intermediate microbatches during gradient accumulation, alongside a separate loss aggregation bug in OpenRLHF that incorrectly weighted per-mini-batch losses. Distributed training configurations specifically triggered both bugs, so a team could catch them only if it deliberately validated results across more than one framework. Together they were enough to systematically deflate SFT baselines across multiple independent published studies, affecting downstream frameworks including TRL, OpenRLHF, and Llama-Factory along the way. Distributed training configurations create failure modes that never surface in a local, single-GPU test run, so a team needs cross-framework validation and genuine observability tooling, not just as a nicety.

How recoverable a collapse turns out to be depends heavily on infrastructure decisions made well before anything goes wrong: how fault-tolerant the system is, how often it checkpoints, and how much visibility the team has into training dynamics in real time. ROLL ran for two weeks uninterrupted, training a large mixture-of-experts model across thousands of GPUs, and that stands as a fault-tolerance proof of concept as much as a throughput demonstration, since the run was built around continuity over raw speed. A team evaluating RLlib against a purpose-built stack should weigh this dimension alongside generation throughput and model-size support, because the framework that trains fastest on paper is not the one that matters if a silent bug or an unrecoverable collapse erases two weeks of compute partway through.

Sources

  1. RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs
  2. SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
  3. Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning
  4. arXiv:2506.06122v1 [cs.LG] 6 Jun 2025 ROLL 2025-06-13
  5. RLlib: Industry-Grade, Scalable Reinforcement Learning — Ray 2.58.0
  6. RLlib: Abstractions for Distributed Reinforcement Learning
Filed underRLHF Tooling

More in RLHF Tooling