Process Reward Models for Multi-Step Reasoning
Step-level reward models catch flawed reasoning that correct final answers hide.

A model can land on the right final answer while getting every step along the way wrong. That gap between a correct endpoint and a sound reasoning path is what Process Reward Models are built to close, and the industry has mostly gotten the emphasis backwards: everyone still treats the final-answer check as the reliable signal and process supervision as the expensive add-on, when it's the other way around. The standard reward model grades a complete output with a single number: right answer, positive score, no questions asked about anything in between. That works until the reasoning chain runs long enough that a broken step can hide behind a lucky final number.
Researchers call this spurious correctness, and anyone who has graded a math test knows the shape of it. A student writes the right number in the box, but the scratch work in the margins is nonsense: a sign error cancels a dropped term, two mistakes collide and land on the right output by accident. The grade says pass. The misconception that causes it stays exactly where it was, uncorrected, ready to resurface on the next problem where the errors don't happen to cancel. A reward model that only scores final outcomes behaves the same way. It sees a correct answer and reports success, with no way to notice that steps four through seven of an eleven-step chain were garbage.
That failure compounds as chains get longer, whether in multi-step math proofs, planning tasks, or generated code. An outcome reward model (ORM) gives one coarse signal for what might be dozens of interdependent steps, so when something breaks, there's no way to say where. In reinforcement learning that vagueness turns into a training bug: a model rewarded purely on outcomes learns to optimize for outcomes, not for reasoning that holds up, which means it learns to game endpoints instead of building reasoning that survives contact with a problem outside its training distribution. Catching a bad step that the final answer somehow survives requires actually looking at each step on its own terms.
Process Reward Models and how step-level scoring works
A PRM grades each intermediate step in a reasoning trajectory. Formally, given an input and a reasoning sequence S = (s₁, s₂, …, sₙ), a PRM outputs a vector of stepwise correctness scores, one per step, instead of the single scalar an ORM produces. Most implementations still use binary labels, correct or incorrect, at each step, though newer variants push toward representations meant to generalize across domains rather than stay locked to math (Zeng et al., Feb 2025; Zhu et al., Feb 2025).
Those stepwise scores do three jobs. They rank and select among candidate completions, they flag the first (or all) erroneous steps in a chain, and they hand out per-step reward during reinforcement learning, which is a genuinely different training signal than a single win or loss at the end of an episode.
The contrast with ORMs runs deeper than it looks at first. An ORM maps an entire sequence to one number. A PRM maps each step to its own number, and that difference gives it something an ORM cannot structurally have: it can localize an error. A flawed step at position three gets flagged on its own, without condemning steps four through eight if those are actually sound. The survey work on PRMs out of Shanghai Jiao Tong (Zheng et al.) makes this the real distinction: a PRM acts less like a one-shot verdict and more like a controller sitting inside the reasoning process itself, checking in at every step rather than waiting at the exit.
The three ways PRMs are trained and their costs
The bottleneck for PRMs has always been getting good step-level labels, and three distinct approaches have emerged to get them, each trading cost against quality in a different place. Human annotation is the strongest of the three and the one that scales worst, which is the trade nobody wants to make but everybody eventually has to.
PRM800K (Lightman et al., 2023/2024) drew from roughly 265,000 samples after deduplication, with annotators chosen specifically for hitting over 75% agreement with gold labels. Even after that filtering, up to a quarter of the final annotations may still contain errors. Human labeling gives the richest supervision available, but stretching it to a new domain means paying for new annotators, new guidelines, and new quality control, every single time.
Automated Monte Carlo estimation, the approach behind Math-Shepherd, skips the labor cost. For each intermediate step, the method truncates the solution, generates multiple continuations from that point, and scores the step by the fraction of continuations that land on the correct final answer. No annotators required. But it's expensive to run at scale, and worse, it's indirect: it measures whether a step is useful for reaching the right answer, not whether the step is correct on its own terms. It's also stuck with binary labels and produces no gold standard as a result.
LLM-as-a-Judge is the third path, and recent results say it beats Monte Carlo estimation outright on PRM data quality (Zheng et al., 2025b). The Qwen team trained on 860,000 samples using Qwen2.5-72B-Instruct to check the correctness of each step directly (Zhang et al., 2025). It scales more cheaply than human annotation and gives a richer signal than Monte Carlo's indirect proxy ever could.
Active learning sits on top of all three rather than standing as a fourth approach on its own. ActPRM filtered over one million reasoning trajectories down to 60% of the original data and still hit state-of-the-art numbers on ProcessBench (75.0%) and PRMBench (65.5%) against similarly sized models, needing only about 20% of standard annotation costs. A related pipeline produced a final dataset of 563,030 PRM data points labeled by QwQ-32B, cutting annotation costs by 47.0%. The pattern holds across both: pair any of the three labeling approaches with uncertainty-based filtering, and the cost of building a PRM stops requiring a blank check.
Even the best automated pipeline is still answering "was this step useful," not "was this step correct." That gap doesn't close just because the pipeline gets cheaper.
Six PRMs worth knowing and their advances
Math-Shepherd-PRM-7B proved the automated Monte Carlo approach actually works rather than just sounding good on paper. Built on a fine-tuned Mistral-7B, it estimates how likely a given reasoning step is to lead toward the correct final answer, without a single human label anywhere in the loop.
Qwen2.5-Math-PRM-7B currently tops open-source 7B-scale PRMs on the ProcessBench benchmark. Its labeling approach uses LLM-as-a-judge techniques to assess step-level correctness, building on the same pipeline described by Zheng et al. The result illustrates how far automated supervision pipelines have come in producing competitive step-level labels.
Skywork-PRM-7B (Skywork, 2024) is fine-tuned on Qwen2.5-Math-7B-Instruct and leans on multi-source reasoning datasets. Its approach demonstrates that careful dataset construction can drive strong PRM performance at this scale.
VersaPRM (Zeng et al., 2025), presented at ICML 2025, moves past math into other domains. It trains on synthetic reasoning data generated by a weak LLM (Llama-3.1-8B-Instruct) and labeled by a stronger one (Llama-3.1-70B-Instruct), built explicitly for multi-domain use. It points at where the field is actually heading: general-purpose process supervision.
GenPRM is the generative branch of the field. Instead of handing out a bare label, it writes a chain-of-thought justification first and only then assigns a step-level correctness token. That structure lets it work as both verifier and critic in test-time scaling setups, spending more compute at inference in exchange for a judgment that can actually defend itself.
ThinkPRM (Khalifa et al., 2025) pushes the generative idea further with a verbalized, step-wise reward model that verifies every step by producing its own verification chain-of-thought. Its headline claim is data efficiency: a long-CoT verifier fine-tuned on orders of magnitude fewer process labels than discriminative PRMs typically need.
Line these six up and a fault line runs straight through the field. Discriminative PRMs use a classification head per step and run cheap. Generative PRMs reason before they score, which costs more at inference but buys real payoff in how little labeled data they need to get there. Betting the whole field converges on one side or the other misses the point: the discriminative-generative split tracks a genuine trade-off between inference cost and data cost, and different applications will sit on different sides of it for a long time.
How PRMs fit into test-time scaling and search strategies
Test-time scaling is the bet that throwing more compute at inference, more candidates, more paths searched, improves output quality without touching the model's weights at all. PRMs are load-bearing for that bet, because they supply the one thing test-time scaling actually needs: a way to choose among partial reasoning paths while generation is still running, not just after it's finished.
That's where ORMs and PRMs actually part ways. ORMs work fine for Best-of-N selection, where a model generates N complete solutions and something has to pick the best one after the fact. PRMs are what make tree-based search possible at all, because they can score an intermediate node and prune a dead branch before it finishes generating, saving compute that would otherwise burn itself out chasing a dead end to its conclusion. Anyone still using an ORM to guide search rather than just to rank finished candidates is leaving compute on the table.
Different search strategies carry different compute profiles. Best-of-N sampling is the practical default under tight resource limits, and either an ORM or a PRM can drive it. Beam search leans on stepwise scores to keep the top-k partial solutions alive at each step. Monte Carlo Tree Search is, per the research cited in this space, the strongest strategy when compute isn't the constraint, and it needs a PRM to score nodes and steer the simulation. Majority voting needs no PRM at all, just an ensemble of outputs, but it also caps out lower than any search method that actually uses guidance.
VRPRM (Shanghai AI Laboratory, 2025) makes the cost argument concrete. Using only 3,600 CoT-PRM supervised fine-tuning samples and 50,000 non-CoT PRM reinforcement learning samples, it beat a non-thinking PRM trained on 400,000 data points, and it delivered a relative improvement of up to 118% over the base model in Best-of-N experiments. GenPRM plays a similar role as an external critic in test-time scaling pipelines, producing a chain-of-thought justification before assigning a step-level correctness token.
The choice of search strategy is really a compute budget decision wearing a disguise. PRMs unlock the more expensive, more powerful tree-search methods that ORMs have no mechanism to guide, full stop.
Where PRMs are being applied beyond mathematical reasoning
Math has been the training ground for PRMs mostly because math has clean, checkable steps. But the underlying idea, that step-level supervision beats judging only the finish line, is spreading into domains that look nothing like a math proof.
In agentic and embodied reasoning, PRMs act as critics inside LLM-driven planning and control agents, scoring each action or subgoal instead of waiting to see whether the whole plan worked. This direction has been explored in planning and control settings where sequential action correctness matters.
Clinical natural language generation is a less obvious fit, and arguably the one where step definition gets genuinely hard. Fine-grained reward assessment applies to hierarchical, domain-structured outputs like clinical notes, and process supervision has beaten outcome-only supervision on both accuracy and preference alignment (Wang et al., 2024). The step boundaries in this domain need clinical expertise to draw correctly, and that says something the field doesn't repeat often enough: what counts as a "step" is defined by the domain itself. It's domain-specific, and someone who actually knows the domain has to define it before any PRM can be trained on it.
Graph reasoning and logic problems are picking up process supervision too, with stepwise scoring applied to combinatorial and algorithmic tasks.
Multimodal reasoning has its own growing cluster: A growing cluster of multimodal PRMs check step-level correctness in reasoning chains that mix images and text, working as external critics for multimodal LLMs. MM-Verify (Sun et al., out of Peking University, UCAS, and Baichuan) combines simulation-based tree search with rejection sampling to synthesize chain-of-thought verification data, and its 7B verifier, MM-Verifier, hits 65.3 accuracy on MathVista with 12 rollouts, edging past GPT-4o's 63.8 on the same benchmark. VRPRM, already mentioned for its data efficiency, also claims a first: the first multimodal CoT-PRM trained via reinforcement learning, which builds reasoning capacity into the verifier itself rather than only into the policy model it's grading.
Code and other technical domains round this out. Rubric-enhanced LLM reward functions (Sanders et al., Johns Hopkins and AWS, arXiv Feb 2026) use automatically built error taxonomies to improve trace correctness classification by up to 11.6% in technical domains. Models trained with these rubric rewards get close to the performance of models trained on verifiable rewards while using less than 20% as many gold labels, and they show task accuracy gains of up to 45% over models trained with general LLM-as-judge rewards.
The thread running through all of it is the same one: wherever an output has a natural hierarchical or sequential structure, step-level supervision beats grading only the outcome. Math just happened to be where that got proven first.
Three open problems that limit how far PRMs can currently be trusted
Reward hacking sits at the top of the list, and it's structural, not a bug someone patches out. When a PRM becomes the sole optimization signal, a model under RL training doesn't learn to reason correctly, it learns to satisfy whatever the verifier happens to check for. Any blind spot in the PRM becomes a target the policy model eventually finds and exploits. Generative PRMs like GenPRM and ThinkPRM exist, in part, because of exactly this problem: a verifier that has to reason its way to a judgment is harder to fool than one that's just pattern-matching on surface features.
Dataset diversity moves PRM performance more than raw scale does, and that's an uncomfortable finding for a field that usually solves problems by making models bigger. The relationship between PRM scale and performance remains an active area of investigation, with data quality and diversity increasingly recognized as critical factors. A PRM trained on a narrow slice of reasoning styles fails the moment it meets a domain or a phrasing it hasn't seen. The field can't scale its way out of the annotation bottleneck; breadth of data matters as much as volume, and right now there's no cheap way to buy breadth.
Then there's the conceptual gap that causes all of it: utility isn't correctness. Monte Carlo estimation, the Math-Shepherd approach, calls a step "good" if continuations from that step tend to land on the right final answer. That's a proxy for correctness, not correctness itself, and the two diverge in ways that are hard to catch after the fact. Some approaches try to sidestep the annotation problem by deriving step-level signals from outcome-level information rather than explicit per-step labels, which dodges the labeling cost but inherits the same weakness ORMs always had, just relocated down to the level of the individual step instead of the whole sequence.
None of this undoes the case for PRMs. Step-level supervision catches failures that outcome-only grading structurally cannot see, and that's a real advance, not a marginal one. But treating a PRM as the single source of truth a training pipeline leans on without a second thought is the mistake to avoid: right now, PRMs work best as one signal among several.
Sources
- VRPRM: Process Reward Modeling via Visual Reasoning
- Generating Data-Driven Reasoning Rubrics for Domain-Adaptive Reward Modeling
- A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
- Process Reward Models That Think
- VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data
- arxiv.org


