Rejection Sampling Fine-Tuning as an RL Alternative
Rejection sampling matches reinforcement learning on math tasks without the complexity.

Rejection sampling fine-tuning (RSF, also called RAFT or RFT depending on the paper) gets a language model to reasoning-level performance by training only on the completions that already scored well. No critic, no clipped surrogate objective, no advantage estimation across a batch of partially-right answers. Generate a batch of completions, keep the ones a reward model or judge likes, and run ordinary supervised fine-tuning on that filtered set. A 2025 paper out of Salesforce AI Research and UIUC puts a number on how close this gets to full reinforcement learning on math reasoning tasks: the paper's own evidence argues against reaching for full RL as a first move. Production pipelines at Meta and DeepSeek reached the same conclusion in practice before the academic literature caught up to it.
RSF versus PPO and GRPO: mechanical differences
Start with what RSF actually asks the model to do. For each prompt, sample K completions from the current policy, score every one with a reward model or an LLM judge, keep the top-scoring subset, and fine-tune on that subset the same way you'd fine-tune on any labeled dataset. The generation step and the update step never touch each other. All K completions come from the policy as it existed before any gradient was computed, which makes the whole loop closer to an offline filtering pass than to anything resembling online RL.
That decoupling explains how RSF differs from PPO. PPO samples one rollout, sends it through a critic network to estimate advantage, and updates the policy with a clipped objective built to keep the new policy close to the old one. Call it a depth-first move: one sample, a lot of machinery wrapped around interpreting that one sample correctly. RSF runs breadth-first instead, many samples per prompt, no critic at all, no on-policy gradient coupling. Vanilla REINFORCE at least skips the critic; PPO adds it back to cut variance. RSF discards both the critic and the tight actor-update loop, and just filters.
GRPO sits in between, and this is where the comparison actually gets interesting. Like RSF, GRPO samples multiple completions per prompt. But GRPO keeps the rejected ones in the batch. It computes a relative advantage for every sample by normalizing rewards against the mean and standard deviation of responses to that same prompt, so a mediocre answer among terrible ones still gets a positive nudge. RSF applies none of that: it throws away everything below the acceptance bar and trains on the survivors as if they were curated data from the start. That's a structural difference, not a tuning choice, and it's exactly the distinction a recent ablation study set out to test.
What the Xiong et al. ablation study found about GRPO's advantage
Wei Xiong, Hanze Dong, and coauthors, in "A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce" (arXiv 2504.11343), ran RAFT, GRPO, and PPO head-to-head on math reasoning benchmarks. RAFT, the positive-only method, lands within a surprisingly small margin of GRPO and PPO on final performance. In the early phase of training, roughly the first 100 to 200 iterations, RAFT actually converges faster than either.
The researchers went looking for the source of GRPO's edge over plain rejection sampling. The obvious guess is reward normalization, the part where GRPO scales advantages by the within-prompt mean and standard deviation. That guess is wrong. What the ablation isolates instead is something far more mundane: GRPO implicitly throws out prompts where every single sampled response is wrong. Those prompts produce no usable gradient signal for a model at its current skill level, and GRPO's structure filters them out as a byproduct of how it's built, not because anyone designed it to. The filtering is doing the work. That's a stronger claim than "GRPO and RAFT are similar," and it's the one the field has mostly missed.
How production training pipelines at Meta and DeepSeek used RSF
RSF has been load-bearing infrastructure in at least two of the highest-profile model releases of the past two years, and in neither case did it stand alone.
Meta's Llama 2 used rejection sampling fine-tuning as the sole alignment method through RLHF-V4. Starting at RLHF-V5, the pipeline stopped treating RSF as a standalone step and instead ran PPO on top of the rejection-sampled results, sequentially. Multiple rounds of rejection sampling happened before RL-based RLHF entered the picture. That sequencing matters: RSF wasn't a substitute Meta tried once and dropped, it was the foundation the later RL stages were built on top of.
DeepSeek-R1 tells a similar story with a different shape. Its pipeline runs in four stages: a cold start, then a reasoning-focused RL phase, then a rejection-sampling-plus-SFT phase, then a final RL phase spanning all scenarios rather than reasoning alone. The interesting move happens in the middle. Once the reasoning RL checkpoint neared convergence, the DeepSeek team used rejection sampling to generate new SFT data, combined it with supervised data drawn from DeepSeek-V3 covering writing, factual QA, and self-cognition, and used that combined set to retrain DeepSeek-V3-Base. The resulting pipeline scored 79.8% on AIME 2024 and 97.3% on MATH-500.
Line those two pipelines up and neither Meta nor DeepSeek used RSF to replace RL, a pattern that's easy to miss if you only read "RSF is competitive with RL" as the headline. Both used it as a stabilizer, a filter that produces cleaner data to hand off to the next RL stage. Anyone reading the Xiong et al. result as license to drop RL entirely is reading past what the production evidence actually shows.
Where RSF falls short and what the lack of negative samples costs
The tradeoff is built into the design, not incidental to it. RAFT throws out every negative sample and trains only on responses that already cleared the bar. The model never sees a labeled example of what it got wrong. Over a long enough run, that has a predictable effect on the policy's output distribution: entropy drops as the model narrows onto whatever patterns keep getting rewarded, since nothing in the loss function pushes back against that narrowing.
Entropy collapse is the term for it, and the Xiong et al. paper treats it as a real risk rather than a footnote. Without exposure to failure modes, a policy trained purely on positive filtering can converge prematurely, locking into a narrow set of behaviors before it has actually explored the space of good answers. The authors still recommend RAFT as a baseline, but they stop short of calling the entropy concern solved.
The ablation results sharpen this into something more precise than "positive-only good, negative-only bad." Completely wrong responses, every sample in a batch scoring zero, carry no useful gradient and should get filtered out, which is exactly the mechanism credited for GRPO's advantage above. But completely correct responses are just as uninformative in the other direction: if every sample is already right, there's nothing left to learn from that prompt either. Reinforce-Rej, the method the paper proposes, filters both extremes and keeps the middle, the prompts where the model is genuinely uncertain. That version matches or edges out both GRPO and plain RAFT. For anyone deploying RSF today, the lesson is to filter with more judgment about which negatives are worth the compute. It's to filter with more judgment about which negatives are worth the compute, and Reinforce-Rej already shows that narrower fix works.
The alignment tax problem and RSF's place in the stack, as revealed by model averaging
None of this happens in isolation from a broader cost that follows RLHF-style fine-tuning in general, RSF included. Lin et al. (arXiv 2309.06256, a collaboration spanning Princeton, HKUST, UIUC, and NVIDIA) document what they call the alignment tax: models fine-tuned for alignment lose ground on the diverse capabilities they picked up during pretraining. The paper measures this across ARC, RACE, PIQA, SQuAD, DROP, and WMT translation benchmarks, and the pattern holds across every one of them.
The uncomfortable part is the correlation Lin et al. report between reward and forgetting: models that hit a higher reward during RLHF tend to pay a larger alignment tax. Better alignment and more forgetting move together, and that's a property of the optimization pressure itself, not a bug specific to RSF. But RSF sits squarely inside the paper's scope, for a reason the authors state directly: RSF and DPO account for nearly all of the open-source LLMs sitting on public leaderboards. Anyone who wants to study the alignment tax where it appears in the field has to look at RSF.
Regularization approaches and a lightweight fine-tuning method were tested against this problem and came up short. What worked better than either was almost embarrassingly simple: model averaging, interpolating the weights of the pre-RLHF and post-RLHF checkpoints. That interpolation produced the strongest Pareto front between alignment and forgetting of anything tested in the paper. RSF doesn't operate as a closed system: it sits inside a stack of decisions about how the final weights get assembled, and some of the most effective interventions happen after the RSF step is technically finished.
RSF versus full RL as a fine-tuning strategy
Default to RSF when the reward is verifiable, math, code, structured output formats where correctness is checkable, when the training run is short, when interpretability during debugging matters, and when standing up a critic network and an on-policy gradient pipeline isn't something the team has the infrastructure to support yet. The engineering case is direct: no critic, no KL penalty to tune against a reference policy, no actor-critic coupling to get right. It's the same loss function used in ordinary SFT, pointed at a dataset that's been filtered instead of hand-labeled.
Full RL earns its overhead under narrower conditions than most teams assume. Longer training runs make entropy maintenance matter more, since that's exactly the failure mode RSF is structurally exposed to. Dense, continuous reward signals give PPO and GRPO more to work with than a binary accept-or-reject filter can capture. And once RSF has already run as a baseline, GRPO or PPO become tools for squeezing out whatever marginal gains are left.
The Xiong et al. result should change how that decision gets made. If GRPO's real advantage over rejection sampling comes from implicitly filtering out prompts where every sampled response is wrong, rather than from its reward normalization, then a practitioner who builds that same filtering logic directly into an RSF pipeline, the way Reinforce-Rej cuts both the all-wrong and all-right batches, captures most of what GRPO offers without taking on GRPO's complexity. Full RL's expensive machinery contributed less to GRPO's edge than the field assumed. The filtering did the work.


