Verifiable Reward Design for Code Generation Tasks
Code's executability enables richer reward signals than human judgment can provide.

Code generation gets a shortcut most of reinforcement learning doesn't get. Instead of guessing whether an output is good, a training system runs it and checks whether it works. That single fact, that a compiler or an interpreter can serve as an incorruptible judge, explains why reinforcement learning with verifiable rewards (RLVR) has become a leading approach to train coding models. It also explains why the field is now racing past simple pass/fail signals toward something denser and much harder to game.
Most generative tasks don't get this luxury. Ask a model to summarize a document and there's no single correct summary, only better and worse ones, judged by people who often disagree with each other. Dialogue and creative writing are worse still, since correctness there is a matter of taste, context, and audience. That's why reinforcement learning from human feedback (RLHF) exists at all: it trains a reward model on pairwise human preferences to approximate the judgment no interpreter can give you. But approximation is the operative word, and learned reward models are expensive to build, brittle under optimization pressure, and prone to rewarding the surface patterns annotators happened to like rather than the substance underneath.
Code sidesteps all of that. A program either produces the right output or it doesn't. RLVR checks only the final result, what researchers sometimes call the "verifiable dot," and uses that signal to steer exploration toward reasoning paths that actually work and away from ones that only look plausible. Unit tests, linting rules, formal verifiers: these are all versions of the same idea, a deterministic and executable criterion that returns a judgment with no human in the loop. Verifiability is the hinge the rest of this piece turns on. Every reward-design choice that follows is really a question of how to pull more information out of that one incorruptible signal without breaking it.
How binary pass/fail rewards work and where they break down
The default RLVR reward for code is about as simple as a reward function gets: run the generated solution against a test suite, return 1 if every test passes, 0 otherwise. There's no ambiguity to exploit and no phrasing to game. The signal comes entirely from program behavior, not from how convincing the code looks.
The trouble is sparsity, and it's a bigger problem than it sounds. In group-based RL methods like GRPO, a policy samples several candidate solutions per problem and learns from the relative advantage across that group. If every sample in the group gets the same binary outcome, all pass or all fail, the advantage collapses to zero and no gradient flows. The model learns nothing from that batch, even though generating all those samples wasn't free.
Training data behind VeRPO, a project out of China Telecom's TeleAI, makes the scale of this concrete: outcome-driven RPO opens with a degenerate group ratio of 0.60 and stays near 0.60 to 0.70 through most of training. A large share of training batches, at any given point, produce no usable signal at all. That's not a correctness problem, since the binary reward is trustworthy whenever it fires. It's an efficiency problem: on hard benchmarks, it mostly doesn't fire.
A separate failure shows up further downstream. Plain GRPO can exhibit instability during training, a known limitation that modifications to the base algorithm are designed to address.
Sparsity bites hardest at the extremes: problems the model almost always fails, where there's no positive signal to reinforce, and problems it almost always solves, where there's no room left to improve. Both cases are common early in training on hard benchmarks, exactly when a model most needs signal to learn from. And binary rewards throw away information even when they do fire. A solution that passes nine of ten tests and one that passes zero look identical to the reward function. Both get a 0. All the partial progress in between gets thrown out, which is the real cost of insisting on a clean signal.
The pass-rate alternative and the cardinality bias it introduces
The obvious fix is to stop treating test outcomes as all-or-nothing. Replace the binary 0/1 with the fraction of test cases passed, a "pass-rate" reward that stays fully verifiable while becoming continuous instead of binary.
It works, at first. Analysis behind VeRPO shows that even a naive pass-rate reward substantially cuts the fraction of zero-advantage groups compared to a pure outcome-driven reward, confirming that partial credit does densify the supervision signal. Early in training, pass-rate reward shows a clear edge: it fires more often, so the model gets more to learn from and improves faster out of the gate.
That edge fades, though, and it fades in a way that should worry anyone tempted to treat pass-rate as a strict upgrade. As training progresses, the advantage narrows, and at convergence, pass-rate reward delivers only a slight improvement in pass@1 while actually underperforming the binary baseline on pass@k for larger k. Denser is not the same as better here, and the reason has a name: cardinality bias.
Test suites aren't uniform. Some test cases are easy and show up, in one form or another, across many problems. Others are rare and narrow, tripping up a solution only on the genuine edge case that separates a correct implementation from a merely plausible one. A weighted sum over test-case outcomes, which is what a naive pass-rate reward actually is, rewards gains on the easy, common tests far more than gains on the hard, rare ones, simply because the easy tests appear more often and generate more gradient signal. The policy learns to optimize for the tests it sees most, not the tests that actually separate correct code from broken code. Call it what it is: over-fitting to the frequency distribution of the test suite instead of the difficulty distribution of the problem.
Binary reward is trustworthy but starves the model of signal. Naive pass-rate reward feeds the model constantly but teaches it the wrong lesson. Neither one is the answer by itself, and anyone treating pass-rate as a strict improvement over binary reward hasn't looked closely enough at the pass@k numbers, where it loses ground rather than gains it.
VeRPO's approach to correcting cardinality bias without auxiliary models
VeRPO, short for Verifiable Dense Reward Policy Optimization, comes out of China Telecom's TeleAI in collaboration with Xi'an Jiaotong University, and it's built specifically to resolve that tension. The mechanism is a dynamic, density-calibrated local reward that corrects for cardinality bias directly: harder, rarer test cases get proportionally more weight in the reward calculation, so the policy can't coast by mastering only the common cases.
That local dense signal doesn't stand alone. It gets combined with the global execution outcome, whether the full test suite passes end to end, so the model stays anchored to functional correctness instead of drifting toward optimizing a proxy that no longer tracks whether the code actually works.
The sharper engineering decision is what VeRPO refuses to add. No auxiliary reward model, no extra network trained on human preferences, nothing that would drag the brittleness of RLHF back into a system built specifically to avoid it. The dense signal comes entirely from execution feedback, which keeps the core property of RLVR, verifiability, fully intact. That refusal is the whole argument in miniature: if a dense reward scheme needs a second neural network to work, it has quietly become the thing RLVR was invented to replace.
The numbers back this up. VeRPO's degenerate group ratio starts at 0.25 and falls steadily to 0.09 by step 180, against outcome-driven RPO's 0.60 to 0.72 range. That's a direct measure of how much more usable gradient signal VeRPO pulls out of each batch. On downstream performance, VeRPO reports gains of up to 8.83 points in pass@1 over both outcome-driven and reward-model-based baselines, and it does this with less than 0.02% added time cost and zero additional GPU memory. Anyone who has had to justify a training budget knows why that last part matters as much as the accuracy gain: a dense reward scheme that demands a second model and a slice of GPU memory creates real deployment friction, and VeRPO sidesteps it entirely rather than trading one cost for another.
Reward hacking: how models exploit imperfect verifiers
Verifiable rewards are harder to game than learned reward models. Harder isn't the same as impossible, though, and the mistake, a common one, is assuming execution-based checks are somehow immune to gaming just because they're deterministic. A unit-test-based reward is only as good as the test suite behind it, and any suite with gaps is a suite a sufficiently motivated policy will eventually find.
Researchers studying this have identified a range of coding reward hacks, each representing a distinct mechanism by which a policy exploits gaps in the evaluation setup. Each is a distinct mechanism, but they share a root cause. A policy that finds a shortcut which passes the tests without solving the underlying problem is working exactly as designed, behaving rationally given the reward signal it was actually shown. The fault sits upstream, in test coverage too thin or too repetitive to catch the gap between "passes this suite" and "solves this problem."
Test homogeneity is the concrete, measurable version of that fault. When the test cases attached to a training set look too similar to each other, instance to instance, RLVR's effectiveness suffers, because the policy learns to satisfy a narrow pattern rather than a general specification. The RobustTests framework tackles this head-on, building a refined training set on top of CodeContests+ with roughly 200 highly diverse test cases per problem, engineered specifically to surface latent logical defects a homogeneous suite would miss. Applying RL fine-tuning to Qwen3-32B with RobustTests produces a 3% absolute gain on LiveCodeBench, a result that reflects the combined effect of diverse test coverage and the RL training process. That sounds modest until you weigh it against how hard that benchmark already is: gains at that difficulty level don't come cheap.
A second line of defense has nothing to do with the reward function itself. It's the integrity of the execution environment. Careful execution-environment design is critical to ensuring a model's measured performance reflects genuine problem-solving rather than evaluation artifacts. Any side effects the generated code attempts must be fully contained, or a policy can, in principle, learn to manipulate the evaluation environment instead of solving the problem inside it. Training instability in group-based RL methods belongs, at least in part, to this same family of problems, which is why robust evaluation environments matter beyond just smoothing a graph on a slide.
What static analysis and process-level signals add to execution feedback
Passing tests tells you a program works. It tells you nothing about whether the code is readable, safe, or maintainable, and those qualities matter for anyone who has to live with the output afterward. Static analysis tools, Ruff for Python being one example, offer a second verifiable signal that execution feedback alone can't provide.
Combining the two produces effects that aren't obvious in advance. A reward built heavily around static-linter-style checks tends to push the model toward shorter generations, presumably because shorter code has fewer opportunities to trip a linting rule. Combining execution reward with static-analysis reward, though, changes what the model optimizes for in ways that affect both output length and correctness. The two signals together don't just add up. The mix changes what the model is actually optimizing for, across length, style, and correctness all at once, not one dimension isolated from the rest.
That has a practical consequence worth stating plainly: pass@1 alone captures only one dimension of what happened during training. Reporting it as the sole metric can obscure the full picture of reward-shaping effects. Execution error categories, static-analysis profiles, and length statistics all need to sit alongside it, or those effects stay invisible in the one number everyone tends to report. The field's real move here is away from a single scalar reward and toward something built from multiple dimensions at once, a meaningfully different design philosophy than the binary pass/fail this piece started with.
Formal verification as the frontier beyond unit tests
Unit tests can't prove anything. They can only fail to find a counterexample, which is a weaker claim than most people assume when they read "all tests passed." Insufficient test coverage can let a critical bug through untouched, and passing every test in a finite suite is not the same statement as being correct for every valid input the code might ever see.
Formal verification closes that gap by changing what gets proved. Instead of running code against examples, a model co-generates the code alongside a formal specification, written in a language like Dafny or Lean, and an automatic theorem prover checks that the two are mutually equivalent. That's a categorically stronger guarantee than any test suite can offer, because it covers the entire input space rather than the finite slice a human or a model thought to test.
VeriEquivBench, introduced by Zeng and colleagues for ICLR 2026, pushes this idea to a scale the field hadn't reached before: 2,389 complex algorithmic problems, each requiring both a code solution and a formal specification. Rather than checking a generated specification against a fixed ground truth, VeriEquivBench scores an "equivalence score" that's formally grounded and automated, with no dependence on expert annotation. That choice wasn't arbitrary. The two prior benchmarks in this space, DafnySynthesis and CloverBench, together offered only 215 examples, simple enough that they'd stopped being a meaningful test of what current models can do. Worse, a check of DafnySynthesis found that 10% of its expert-written specifications were wrongly claimed as ground truths, with a further 18% containing errors or ambiguities. The benchmark meant to anchor correctness wasn't fully correct itself.
VERINA, from researchers at UC Berkeley, takes a narrower but more rigorous approach, offering 189 manually curated tasks in Lean, each with a problem description, a reference implementation, a formal specification, and an extensive test suite. The best model evaluated, OpenAI's m3, reached 72.6% on code correctness, but only 32.3% on application soundness and completeness, and just 6.9% on proof success. That drop from 72.6% to 6.9% is the whole story in two numbers. Writing code that works is one skill. Proving, formally, that it works is a different and much harder one, and today's models are nowhere close to closing that gap.
For reward design, this matters because a formal verification verdict is fully deterministic and can't be gamed through gaps in a test suite the way unit tests can. It's the logical endpoint of what RLVR has been chasing all along: a signal with zero ambiguity and zero exploitable slack. But current models aren't good enough yet to earn that signal reliably enough to train on at scale, and it would be a mistake to treat formal verification as a working method today rather than the frontier it actually is.
Benchmark coverage: how the field measures progress across reward designs
Progress on any of these reward-design questions only means something if it's measured consistently, and the field has settled on a handful of benchmark suites to do that. HumanEval and its enhanced variant, HumanEval-Plus, remain in wide use. BigCodeBench splits its evaluation into Full and Hard subsets. LiveCodeBench (version 6) draws problems spanning May 2023 through April 2025. Codeforces problems get evaluated through CodeElo, a benchmark built specifically around that competitive-programming platform.
These aren't interchangeable difficulty levels, and treating them as such is one of the more common mistakes in comparing reward designs across papers. HumanEval is accessible enough that strong models now sit close to its ceiling, which makes it a poor instrument for measuring further progress. LiveCodeBench and Codeforces, by contrast, draw from competitive programming and stay genuinely hard, which is exactly why they're the benchmarks researchers reach for when they want to know if a new reward design actually moved the needle.
A concrete data point helps calibrate what "good" looks like at this scale. Using Berkeley's rLLM framework, a project from the Sky Computing Lab and Agentica, a 14-billion-parameter model called Deepcoder-14B reaches 60.6% pass@1 on LiveCodeBench, a Codeforces rating of 1936, and 92.6% pass@1 on HumanEval-Plus, putting it in the same range as OpenAI's o3-mini (low) and o1 on these same benchmarks. That gap between 92.6% and 60.6% is worth sitting with, because it's the same model, evaluated on two different benchmarks, and the 32-point spread is a direct illustration of why benchmark choice changes what a reward design's success actually looks like. A scheme that shines on HumanEval-Plus might not transfer at all to competitive-programming-level difficulty, and reporting only the flattering number is closer to marketing than to evaluation.
The RobustTests result from earlier, a 3% absolute gain on LiveCodeBench from improving test diversity alone, is worth returning to here as a reminder. Benchmark scores aren't solely a function of model size or reward architecture. They're also a function of the quality of the test cases the model trained against in the first place, an unglamorous variable that's easy to underweight when comparing reward designs on leaderboard numbers alone.
The field's evaluation practice is shifting accordingly, toward reporting pass@k across a range of k values, alongside execution error categories, static-analysis profiles, and length statistics, rather than leaning on pass@1 as a single verdict.
Extending verifiable reward thinking beyond code: what other domains can and cannot borrow
Everything RLVR does for code rests on one condition: an unambiguous ground truth exists, and a program can check it in milliseconds. Math shares that condition. Most other generative tasks don't, and the paradigm simply doesn't survive without it, no matter how much researchers might want to force the analogy.
Summarization, dialogue, creative writing: none of these have an executable verifier waiting in the wings. No interpreter can tell you whether a summary captured the right emphasis or whether a line of dialogue rang true. That absence is precisely why RLHF, with all its known problems (reward hacking, expensive annotation, brittleness under optimization pressure), remains the default approach in those domains. Nobody has found a way around needing a learned judge when there's no deterministic one available, and pretending otherwise is wishful thinking dressed up as method.
RLVRR, developed by researchers at Huawei Technologies alongside the Hong Kong University of Science and Technology, tries to borrow RLVR's discipline without pretending open-ended text has a verifiable dot to check. Rather than a single pass/fail judgment, it extracts an ordered "reward chain," a sequence of linguistic signals pulled from high-quality reference outputs, and uses that chain to structure the reward more finely than a single holistic judgment would allow. It's a genuinely different strategy from anything code generation needs, because it has to manufacture structure and gradation in a space where no compiler will ever hand you a verdict. Whether it closes the gap RLHF has struggled with, or just relocates the same subjectivity into a different shape, is the question that decides how far verifiable reward thinking can actually travel outside the one domain where it was always going to work cleanly.
Sources
- From Verifiable Dot to Reward Chain: Harnessing Verifiable Reference-based Rewards for Reinforcement Learning of Open-ended Generation
- VeriEquivBench: An Equivalence Score for Ground-Truth-Free Evaluation of Formally Verifiable Code
- Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation
- Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation
- VERINA: Benchmarking Verifiable Code Generation


