GRPO vs PPO for Large Language Model Fine-Tuning
GRPO eliminates the critic model, cutting memory costs without sacrificing performance.

Supervised fine-tuning gets a language model to imitate good answers, but it does not teach the model what to do when two answers are both plausible and one is quietly better than the other. That gap, between mimicking a dataset and actually satisfying a human preference, is what reinforcement learning from human feedback was built to close. The now-standard recipe runs in three stages: supervised fine-tuning first, then training a reward model on human preference comparisons, then using that reward model to drive a reinforcement learning update on the policy. A KL penalty gets bolted onto the reward signal at that third stage, holding the updated model close to its supervised starting point so it doesn't wander off into outputs that score well on the reward model but read as broken or strange to an actual person.
Proximal Policy Optimization became the default engine for that last stage almost by default. It appears in the formative work that defined this whole approach: Ziegler et al.'s early experiments applying reinforcement learning to language models, OpenAI's InstructGPT paper, Bai et al.'s work on aligning language models at another lab, and the LLaMA 2 release. By the time DeepSeek and other labs started looking for alternatives, PPO wasn't just one option among many. It was the accepted default. Any real alternative had to prove itself against PPO's track record, not against some abstract ideal.
PPO's mechanics and costs at LLM scale
PPO's central trick is a clipped surrogate objective: it caps how far a single policy update is allowed to move the model's behavior, so training stays stable without needing the hard trust-region math that earlier algorithms relied on. That clip is doing real work. Without it, a single batch of unusually good or bad rewards could push the policy into a corner it can't recover from.
To know whether a given response deserves reinforcement, PPO needs a baseline: some sense of what return was expected at each point in the sequence, so it can measure the actual outcome against that expectation. It gets that baseline through Generalized Advantage Estimation, or GAE, which leans on a learned value function, a model trained to predict expected future reward from any given token position. That value model is typically its own full-size network, initialized from a pretrained LLM roughly the same size as the policy model, and trained alongside it, update for update. It's typically its own full-size network, initialized from a pretrained LLM roughly the same size as the policy model, and trained alongside it, update for update.
PPO's training loop needs at least three large models running at once: the policy being updated, the reward model scoring its outputs, and the critic estimating value. Practitioners sometimes describe this as three large models to improve one, and the phrase is not an exaggeration so much as a fair summary of the memory bill. Every one of those models needs to sit in accelerator memory, and the value model in particular roughly doubles the footprint of the policy model on its own.
GRPO's structural changes and the role of the group
A group-based policy optimization method, introduced in DeepSeek's work on mathematical reasoning and later brought to wider attention through their work on DeepSeek-R1, starts from a simple question: what if the baseline didn't need to be learned? Instead of training a separate value network to estimate expected return, GRPO gets its baseline directly from data it already has to generate anyway.
The mechanism is straightforward. For a given prompt, GRPO samples a group of G responses, not just one. A reward model scores each of them. Then, instead of comparing each score to a learned prediction, GRPO computes a z-score: it subtracts the group's mean reward and divides by the group's standard deviation, turning each response's score into a measure of how it did relative to its own peer group. Responses that beat the group average get reinforced; responses that fall short get pushed down. No critic, no separate value network, no second large model trained in parallel just to produce a number the group could have produced on its own.
The architectural differences that determine training behavior side by side
Lay the two algorithms next to each other and the differences are structural. PPO requires a critic model roughly the same size as the policy; GRPO drops the critic. PPO estimates advantage through GAE, a learned process that relies on a value function to assess expected return across a sequence; GRPO estimates advantage through a group-relative z-score, computed fresh for each batch of sampled responses with no learned component.
That difference affects memory usage directly. PPO's use of a full second value model roughly doubles the memory needed relative to the policy model alone. GRPO, having no critic to store or update, keeps its footprint close to the size of the policy model by itself. Even the KL penalty is handled differently: PPO incorporates it into the reward signal, while GRPO handles the KL term within its optimization objective. Small detail, but it reflects the same underlying philosophy: GRPO keeps things simple and computes what it needs on the fly, rather than maintaining a separate learned system to estimate it in advance.
DeepSeek-R1 as the case that proved GRPO's practical ceiling
DeepSeek didn't pick GRPO as an academic exercise. The choice was about cost: eliminating a critic model of comparable size to the policy meant a meaningfully lighter training setup, and that mattered at the scale DeepSeek was operating on. The clearest test of what that architecture could actually do came with DeepSeek-R1-Zero, a model trained with pure reinforcement learning applied straight to the base model, with no supervised fine-tuning warmup at all, an unusual choice. That is an unusual choice. Skipping SFT means the model enters RL training without any demonstration of what good reasoning even looks like, which puts real pressure on whatever RL algorithm is steering the process.
GRPO held up under that pressure, and what came out the other side was not something anyone wrote into the training data. Chain-of-thought reasoning emerged unprompted during the RL process. So did self-verification behavior, reflection on prior steps, and backtracking when the model's own reasoning hit a dead end. None of that was demonstrated to the model beforehand; the RL process itself produced it.
The numbers back up how far that took the model. On AIME 2024, DeepSeek-R1-Zero's pass@1 score climbed from 15.6% to 71.0% purely through RL training, and majority voting across samples pushed that figure to 86.7%, putting it in the same range as OpenAI's o1-0912. For an algorithm whose main selling point was supposed to be cost savings, that is a striking result on capability grounds alone.
Where GRPO works well beyond flagship reasoning models
GRPO's group-relative design fits most naturally with tasks where the reward signal is unambiguous. Mathematical reasoning, code generation, and structured question answering all share a property that makes this work: a rule-based or binary check, correct or incorrect, passes the test suite or doesn't, can generate a clean reward without any of the ambiguity that plagues subjective judgments. That's the domain where GRPO was proven out, but it isn't the only place it functions. The algorithm also works with general-purpose reward models trained on helpfulness judgments, not just hard binary signals, which opens the door to less rule-bound applications.
Researchers have already carried GRPO's reward design into agentic settings, adapting it to optimize tool use in small language models, extending the approach beyond pure reasoning tasks into scenarios where a model has to decide when and how to call an external function. And the small-model evidence is concrete: a 3B-parameter Llama model fine-tuned with GRPO improved its accuracy on Zig, a programming language with little representation in most pretraining data, from 3% to 27% in just 200 fine-tuning steps. It's a data point suggesting GRPO's efficiency gains matter even when the model in question is nowhere near frontier size, not a flagship-scale result. It's a data point suggesting GRPO's efficiency gains matter even when the model in question is nowhere near frontier size.
The failure modes GRPO introduces and their design tradeoffs
None of this makes GRPO free of problems, and the problems trace directly back to the same design choices that make it efficient. Because GRPO's advantage function normalizes by group standard deviation, it ends up biased toward reward functions with high variance. In multi-objective training, where several reward signals are combined, that bias means GRPO tends to lock onto optimizing whichever objective produces the noisiest signal, often at the direct expense of the others. MO-GRPO was proposed specifically to fix this, introducing a normalization method that automatically reweights reward functions according to their variance, so no single noisy objective can dominate the update.
There's a second issue that isn't unique to GRPO but hits harder because of what GRPO lacks. When the reward model can't offer a comprehensive assessment of response quality, both PPO and GRPO become prone to reward hacking, especially in later stages of training, as the policy learns to exploit gaps in what the reward model actually measures. PPO's learned value function provides some buffering against this, smoothing out noisy reward signals over time. GRPO has no such buffer. Its baseline is recomputed fresh from each sampled group, so a flawed reward model's blind spots pass through more directly.
Response length bias is the clearest everyday example. Reward models, across many training setups, tend to favor longer responses regardless of whether the extra length adds anything. GRPO applies its computed advantage across a full response without token-level correction, so if the group-level reward signal favors length, that bias propagates through the entire response.
How the GRPO variant ecosystem addresses those limitations
A handful of variants have grown out of these specific weak points, each targeting one part of the original design. DAPO builds directly on GRPO and introduces several changes, including a clip-higher mechanism to stop entropy from collapsing during training and overlong reward shaping to handle responses that run past expected length. DAPO reached state-of-the-art results on AIME 2024. Weigh that tradeoff carefully against the performance gain before treating DAPO as a drop-in upgrade.
Other variants take the opposite approach, stripping the algorithm down rather than adding to it, targeting the specific sensitivity to reward noise and length bias that standard normalization introduces.
MO-GRPO, already mentioned above, solves the variance that arises when balancing multiple competing objectives through automatic reweighting, making GRPO usable in settings where several competing reward signals need to be balanced rather than left to compete unsupervised. And GRPO-TTA carries the group-relative idea somewhere the original paper never addressed at all: test-time adaptation for vision-language models, reframing class-specific prompt prediction as a group-wise policy optimization problem. That a text-generation algorithm generalizes into multimodal test-time settings illustrates how broadly the group-relative mechanism can be applied beyond its original reasoning-focused context.
Choosing between PPO and GRPO for a given fine-tuning scenario
The decision mostly comes down to what infrastructure is available and what kind of task is being trained. Anyone working without hyperscale compute should treat GRPO's critic-free architecture as the practical default: PPO's roughly doubled memory footprint, driven by that full-size value model, can be prohibitive outside of well-funded lab settings, and GRPO was built precisely to remove that constraint.
Task type matters just as much as compute budget. Math problems, code generation, structured QA: anything with a verifiable, rule-based outcome plays to GRPO's strengths, since the group-relative reward signal needs exactly that kind of clean, checkable output to work well. Tasks that call for fine-grained, partial-sequence value estimation, where knowing the expected return at each token affects how accurately credit is assigned across a sequence, may still be better served by PPO's GAE.
Reward model quality is not a place where the two algorithms diverge. Both are exposed to reward hacking when the underlying reward model is weak, and neither one is a substitute for getting reward design right. Multi-objective training is the one setting where the gap reopens: base GRPO struggles there without help, and needs MO-GRPO or comparable reward shaping to hold multiple objectives in balance, whereas PPO's learned value function can, in some configurations, manage mixed objectives more gracefully on its own. Choosing between the two, in the end, is less about which algorithm is better and more about which set of tradeoffs actually fits the problem sitting in front of you.
Sources
- Advancing SLM Tool-Use Capability using Reinforcement Learning
- GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- cameronrwolfe.substack.com
- cameronrwolfe.substack.com
- aipapersacademy.com


