The Reward Signal

KL Divergence Penalty Tuning in PPO for LLMs

Setting this hyperparameter wrong dooms alignment to either inertia or reward hacking.

Senior Writer · · 12 min read
Cover illustration for “KL Divergence Penalty Tuning in PPO for LLMs”
RLHF Foundations · September 4, 2026 · 12 min read · 2,794 words

The coefficient β sets how far a language model can wander from its supervised fine-tuned starting point while chasing reward during PPO training. Get it wrong and the whole RLHF setup breaks in one of two directions: a model that refuses to budge, or one that learns to trick its own reward model into applauding gibberish. This piece looks at what β does, why it shows up twice in a single training run wearing two different hats, and why practitioners across labs have landed on wildly different numbers for something that sounds, on paper, like it should have one correct answer. Spoiler: it doesn't, and the confusion around this hyperparameter mostly starts from expecting a single universal setting to exist.

The PPO phase in RLHF pushes a policy π_θ toward outputs that score well under a learned reward model, while a frozen reference policy, πSFT, sits in the background like a leash clipped to a fence post. Left off that leash, a policy finds the shortest path to a high reward score, and that path often runs straight through nonsense: repetitive phrasing, truncated answers, weird tokens repeated for no reason, all of it can score well under an imperfect reward model even as a human reader winces. That's reward hacking, and it's why the standard RLHF objective looks like this: maximize E[r(y|x) − β log(πθ(y|x) / π_SFT(y|x))]. The second term is a KL divergence penalty, and β is the price tag on wandering from home. At convergence, the resulting policy works out to be proportional to π_SFT(o|c) times exp(R(c,o)/β), which means β sits in the exponent: nudge it even slightly and the whole output distribution shifts, not just at the edges.

Two distinct KL mechanisms operating inside PPO at once

Here's the detail that trips up almost everyone learning PPO for the first time: there are two separate KL-related mechanisms running at the same time, distinct in purpose and behavior despite both carrying the letters K and L in their names.

The first is the PPO clip objective itself, the trust-region mechanism that keeps a single training step from moving too far from π_old, the checkpoint that generated the current batch of rollouts, not the original SFT model. This runs through a probability ratio clip; no explicit β shows up anywhere in it. Its job is short and narrow: don't let one gradient step blow the policy up.

The second mechanism is the KL penalty against π_ref, the frozen SFT model, and this is where β actually lives. It applies token by token as a dense reward-shaping signal, so the full per-token reward equals the reward model's score plus β times the KL term at that position. That distinction matters because the reward model itself only fires at the final token of a sequence; everything before that is sparse silence. The KL term is the only signal running dense through the whole generation, token after token, holding the leash taut the entire way.

So tuning the clip epsilon and tuning β solve two different problems, each addressing a distinct failure mode with its own mechanics. Turning up the clip won't fix long-run drift, and tightening β won't stabilize one noisy update. Worth flagging: GRPO handles the KL term as a standalone loss component instead of folding it into the per-token reward, so a β reported in a GRPO paper isn't directly comparable to a β from a PPO paper. Same letter, different currency, and mixing them up is an easy way to misread a paper's hyperparameter table.

What β too high and β too low each produce

Diagram: β Too High vs. β Too Low: Two Distinct Failure Modes. Visualizes: Visualize a single curve showing proxy reward and gold reward diverging as KL budget (controlled by β) increases, illustrating the two failure modes described in the article.

Set β too high and the policy barely moves off its SFT parent. The reward model can be shouting useful signal the whole time; the policy still creeps forward, training burns hours, and the alignment gains at the end are thin. The constraint has eaten the objective whole.

Set β too low and the opposite failure shows up, just later and louder. The policy explores hard, blows through whatever KL budget existed, and proxy reward keeps climbing even as actual quality, the "gold" reward a human or stronger judge would assign, peaks and then falls off a cliff. Outputs start repeating themselves, cutting off early, or drifting into a style that scores well on the RM and reads terribly to anyone else. A parametric study comparing PPO and GRPO found that in GRPO specifically, smaller β tracked with shorter, less comprehensive responses, a countable form of reward hacking rather than a vague impression of one.

The sweet spot sits in the middle, and it functions much like early stopping: it caps how long optimization gets to run before it starts overfitting to the proxy reward model instead of to actual quality. Peak gold reward tends to land at a moderate KL budget, not at either extreme. Worth flagging here, since later sections build on it directly: a uniform KL penalty treats every token position the same, including spots where the reference model is genuinely unsure and where the policy most needs room to move. That's a structural blind spot, and it gets its own section further down.

Reward overoptimization and why KL divergence is the natural budget measure

Reward overoptimization has a name because it has a shape: proxy reward climbs smoothly and keeps climbing, while gold reward rises, peaks, and then drops. Gao et al. (2022), in "Scaling Laws for Reward Model Overoptimization," mapped this out in detail and made KL divergence from the reference policy the standard x-axis for plotting the curve. That's why β comes up in every conversation about overoptimization: it's the dial that decides how far along that curve training gets to travel.

But the same paper complicates the tidy story people like to tell about it. Gao et al. found that adding a KL penalty can, under some conditions, widen the gap between proxy and gold reward rather than close it. So β doesn't uniformly suppress overoptimization; it moves where along the KL axis the overoptimization shows up. That's a subtler claim than "more penalty equals safer training," and it matters, because it means β schedules risk rather than insuring against it.

The overoptimization coefficients scale with reward model parameter count and dataset size, and scale more weakly with the size of the policy being trained (Gao et al., 2022). Gao et al. also found RL-based optimization runs slower than best-of-n sampling at both optimizing and overoptimizing, which complicates comparing the two methods on a KL axis alone. The upshot for anyone running these systems: the KL budget β enforces is a proxy for overoptimization risk, with no guarantee of acting as a hard ceiling on it. That gap is exactly what motivates the adaptive controllers and token-level methods covered later.

How reward model size changes the β you need

Here's a finding that turns β from a fixed knob into a knob whose right setting depends on a totally different hyperparameter. Research unpacking DPO and PPO found the optimal β shifts with the size of the reward model doing the scoring, and the shift isn't small.

With a smaller reward model, something in the 13B range, performance shows large drops as β decreases. The policy needs a tight leash because the RM itself is less trustworthy, and any freedom to explore just gets spent hunting for the reward model's blind spots. A larger reward model, around 70B, behaves differently: it is more robust across β values, and it hits higher performance at smaller β than the small RM ever manages at any setting. Bigger reward models are harder to game, so the policy can run with a longer leash without hanging itself on it.

That has a practical, occasionally expensive consequence: picking β without first knowing the reward model's scale is close to a coin flip. Pair a small RM with a low β and expect training to look fantastic on the reward curve and rough on the actual outputs. The same research also found larger RMs are simply easier to tune around, exactly because their performance cares less about where β lands. Bigger reward model, wider margin for error, a relationship worth checking before touching β at all.

Empirical β ranges reported across major implementations

Ask five teams what β they used and expect five different numbers, none of them wrong, all of them shaped by a setup that isn't quite like anyone else's. A widely cited hyperparameter survey lists a typical range of β = 0.01 to 0.2, with a default around 0.05 and adaptive schemes aiming for a target KL of 5 to 8.

AlpacaFarm (Dubois et al., 2023) used β = 0.02, fixed rather than adaptive, for human-feedback PPO runs, and β = 0.002 for simulated PPO runs, a tenfold gap the authors trace to differences in early stopping criteria and the scale of the surrogate reward values in play. Experiments on Pythia 1B and 2.8B in 2025 used a baseline of β = 0.05, though β = 0.03 produced higher gold scores in that particular study; a Qwen 2.5 1.5B setup tightened things to β = 0.1 to hold stability at smaller scale. GRPO baselines commonly report a reverse-KL penalty of β = 0.04, and GSM8K experiments run with LoRA use λ_KL = 10⁻³ for both TPO and GRPO. Llama-2 7B production-style runs used β = 0.1 for harmlessness objectives and β = 0.05 for other objectives, with target KL values running from 4.0 on faithfulness tasks up to 15.0 on helpfulness tasks.

The pattern across all of it: task type and model scale both shove the acceptable range around, and safety-flavored objectives like harmlessness want a tighter leash than capability-flavored ones like helpfulness. There's no universal β. Anyone quoting one without naming the reward model, the task, and the model scale is quoting a number stripped of the context that made it mean anything.

Adaptive KL controllers versus fixed β schedules

Fixed β is the simple option: pick a number before training starts, hold it steady, and hope the KL trajectory behaves the way it did in whatever paper the number came from. Easy to implement, easy to reason about, and brittle the moment reality diverges from the paper, which happens early and often.

The better default, when it's available, is an adaptive KL controller, the approach used in OpenAI's original PPO variant for RLHF. This controller checks the observed KL from the last update against a target and adjusts: push β up when the policy drifts faster than intended, pull it down when it drifts too slowly. It aims for a band, not a frozen number. Some systems start this process with a small initial β, around 0.001, and let adaptive control do the rest.

OpenAI's later work moved away from the adaptive KL mechanism toward the clipped surrogate objective (L^CLIP), which gets a trust-region-style update without needing any of the adaptive machinery. That's a genuine simplification of the algorithm, but it costs something: explicit KL targeting disappears, and with it a diagnostic signal adaptive controllers hand over almost for free. Current practice splits down the middle. Some RLHF stacks still favor adaptive KL with a small initial β; GRPO setups often prefer a larger, fixed, explicit penalty around 0.04 for training stability. There's a useful side effect of adaptive control beyond its main job, too: if β keeps getting cranked upward run after run, that's a fairly reliable sign the reward model's signal is stronger, or noisier, than the policy can actually absorb.

Token-level KL penalties and why uniform penalization misallocates constraint budget

Most implementations sum the KL penalty uniformly across every token in a sequence, treating each position as equally deserving of restraint. That raises an obvious question: is the reference model actually equally unsure at every position? It is not, not even close.

Some tokens push the policy into territory the reference model never saw clearly: novel reasoning steps, unusual formatting, an answer structure the SFT data barely touched. At exactly those high-entropy, high-impact positions, a uniform penalty clamps down hardest right where the reward signal most needs room to work. Meanwhile at low-entropy positions, common function words, boilerplate phrasing, that same uniform penalty spends its budget guarding territory that was never at risk to begin with.

Vassoyan et al. (2025) found, empirically, that uniform KL penalties can block exactly the exploration that matters most, at what the paper calls "critical tokens": positions with high impact on whether the final output is correct and elevated entropy under the reference model's own predictions. Related work on token-level DPO makes a similar point: DPO scores KL divergence at the sentence level, but generation happens one token at a time, and divergence stacks up token by token rather than arriving in one lump at the end. A moderate uniform β, then, can be too tight where the model needs room and too loose where it needs none, at the same time, in the same sequence. Token-adaptive approaches try to fix this by varying constraint strength by position instead of applying one number everywhere. Worth noting too: PPO's per-token objective structure clashes with reward signals that only exist once per full response, a mismatch that's part of what drives sequence-level clipping alternatives like GSPO.

When removing the KL penalty is a deliberate design choice

The whole KL penalty framework leans on one assumption: that π_SFT represents reasonably well-calibrated, human-like language, and drifting from it is inherently risky. That holds up fine when the goal is aligning a model to general human preferences. It holds up much worse, according to DAPO (ByteDance, 2025), when the goal is training long chain-of-thought reasoning.

DAPO's argument goes like this: learning extended reasoning chains needs the policy distribution to diverge substantially from its SFT starting point, because the SFT model was never trained to reason step by step across long chains in the first place. In that regime, the KL penalty actively fights the thing training is trying to do rather than safeguarding it. So DAPO drops the KL term entirely. β equals zero, on purpose, not by accident.

The results are worth sitting with: DAPO reaches 50 points on AIME 2024 using Qwen2.5-32B, against 47 points for DeepSeek-R1-Zero-Qwen-32B, and gets there in half the training steps. That said, there's a gap worth naming directly: the DAPO paper never isolates the KL removal on its own, so there's no clean way to credit those gains specifically to dropping β to zero rather than to the other changes bundled alongside it. DAPO and Dr. GRPO both sit inside a growing cluster of work arguing the KL term isn't needed for reasoning-heavy tasks, but that claim is scoped tightly to long chain-of-thought training, and shouldn't be read as a broader license to strip KL out of every RLHF pipeline. Separate GRPO experiments confirm β = 0 causes outright training collapse under sparse terminal rewards, so whether removing KL works depends heavily on task type and how dense the reward signal is to start with. Context decides this one, and treating it as a style choice is how teams get burned.

Theoretical limits of KL regularization under heavy-tailed reward error

There's a formal ceiling on what KL regularization can promise, and it comes out of work on what's called "Catastrophic Goodhart" (Kwa et al., NeurIPS 2024), which pins down exactly when KL constraints stop protecting against reward hacking at all.

The distinction turns on the tail behavior of the reward model's error. When reward error is light-tailed, meaning extreme errors stay rare and bounded in a predictable way, relaxing the KL constraint stays safe: policies optimized under looser penalties still reach arbitrarily high true utility as the penalty loosens further. That's the comfortable regime, the one most intuitions about β get built on.

Heavy-tailed reward error breaks that comfort completely. In that regime, a policy can hit arbitrarily high proxy reward while gaining no more real utility than the untouched base model ever had. The mechanism is exponential tilting: since the optimal policy under the RLHF objective is proportional to π_SFT times exp(reward/β), a heavy tail in the reward error means the exponential can find and massively amplify a handful of extreme, essentially fluke, high-scoring outputs, no matter how small β gets set. Shrinking β delays the problem rather than solving it, because the tail is still sitting there waiting. That's a sobering footnote to everything above: β can be tuned with real precision, informed by reward model scale, task type, and empirical ranges pulled from a dozen implementations, and still run into a wall if the underlying reward model's error distribution has a fat enough tail. Tuning the knob assumes the knob works at all, and this result is the reminder that, under the wrong conditions, it might not.

Sources

  1. apxml.com
  2. apxml.com
  3. apxml.com
  4. cameronrwolfe.substack.com
  5. arxiv.org
Filed underRLHF Foundations

More in RLHF Foundations