RAWR: Reward Assignment Without Rollouts in Verifiable Domains
Authors: Corentin Royer, Anna Hedström, Debarun Bhattacharjya, Gaetano Rossiello, Andrea Giovannini, Mennatallah El-Assady
Organizations: International Business Machines · Department of Computer Science, ETH Zurich, 8092 Zurich, Switzerland · ETH AI Center · Lirio · Department of Computer Science, ETH Zurich
Understanding and evaluating multi-step reasoning in LLMs at the level of individual steps remains a key challenge. Process reward models (PRMs) provide a solution by scoring each step, enabling fine-grained supervision and improved reliability. However, training them requires costly human annotation or computationally intensive rollout-based labeling. To solve this, we introduce MCNIG, a scalable method for automatically labeling the quality of individual reasoning steps in any verifiable domain. Its step score, net information gain (NetIG), improves upon single-reference information gain (IG) by comparing the most-supported correct answer against the most-supported incorrect one, yielding a robust signal even for long and structured outputs like code and SQL, where IG fails. We show that the signal produced by MCNIG correlates with human judgments of step quality, and we apply MCNIG labels to train PRMs that achieve the best average best-of-K accuracy across eight benchmarks spanning mathematics, code generation, text-to-SQL, and scientific QA. Crucially, MCNIG generates no rollouts, cutting labeling complexity to O(N) and making it up to X times cheaper than rollout-based methods at comparable label quality, which makes large-scale process supervision practical.
Figures & tables
Figure 1: Overview of RAWR. (1) An LLM samples K CoTs for question q . (2) A verifier Vq partitions them into correct ( Cq ) and incorrect ( Wq ) sets. (3) At every prefix r0,…,rn , we score each candidate answer and keep the max over Cq and Wq . (4) The step score, NetIG, is the change in the correct-incorrect gap relative to r0 , thresholded at τ to produce binary step labels.
Method
Label type
Complexity
Rollout-free
Cross-domain
PRM800K ( Lightman et al., 2023 )
Human expert
—
✓
✗
MathShepherd ( Wang et al., 2024a )
Rollout
O(N2)
✗
✗
OmegaPRM ( Luo et al., 2024 )
Rollout + binary search
O(NlogN)
✗
✗
QwenPRM ( Zhang et al., 2025 )
Rollout + consensus filtering
O(N2)
✗
✗
RAWR (ours)
Information Theoretic
O(N)
✓
✓
Table 1: Comparison of step-level supervision methods. RAWR requires no rollouts and achieves O(N) complexity. Relevant prior methods have been validated on mathematics only.
Figure 2: Left: Balanced accuracy of CoT -level labels for IG vs. RAWR across domains. Right: Accuracy stratified by the number of distinct correct answer formulations. IG degrades as formulation diversity increases, while RAWR remains stable.
Figure 3: ProcessBench Step F1 against labeling cost relative to RAWR ( Section A.1 ); top-left is best. Filtered QwenPRM adds an LLM-as-a-judge filtering stage to the same rollouts, raising F1 but discarding half of the labeled samples ( 155× ).
Figure 4: Best-of- K accuracy as a function of K , stratified by ground-truth answer length, for PRMs trained on IG labels and on RAWR labels. The RAWR- IG gap is negligible for short answers and widens as answers grow longer, confirming that single-reference IG degrades on structured, long-form outputs while RAWR remains robust. Shaded bands show ±1 standard deviation across repeated subsamples of K answers per problem.
Method
MATH
GSM
PubMed
HumanEval
BigCode
BIRD
AIME
Average
Single sampling
44.5%
73.3%
36.2%
55.5%
23.3%
19.6%
17.4%
38.5%
Majority voting
67.3%
91.2%
46.8%
66.1%
27.6%
37.6%
29.0%
52.2%
OVM
62.8%
92.0%
59.4%
76.3%
34.9%
45.1%
26.1%
56.6%
MathShepherd
56.1%
89.5%
49.7%
52.5%
29.5%
35.9%
23.0%
48.0%
ImplicitPRM
59.3%
89.9%
44.3%
77.8%
29.3%
35.6%
19.6%
50.8%
Granite PRM
62.5%
90.9%
49.8%
61.6%
31.4%
35.4%
27.0%
51.2%
Table 2: Best-of- K accuracy ( K=32 ) across eight benchmarks. Bold indicates the best result and underline the second best, computed separately within the 8B group and globally for RAWR 14B. Shaded rows are our methods; RAWR 8B achieves the highest average performance among 8B models.
Method
UGPhysics
ProcessBench Trace
ProcessBench Step
Single sampling
8.0%
—
—
Majority voting
11.5%
—
—
OVM
9.2%
33.3
18.6
MathShepherd
10.0%
41.1
31.5
ImplicitPRM
10.6%
68.6
21.2
Granite PRM
11.0%
51.9
51.4
Table 3: External evaluation. UGPhysics : best-of- K accuracy (OOD domain). ProcessBench Trace: F1 for classifying whether a trace contains an error (introduced by us). ProcessBench Step: official F1 for localizing the first incorrect step. Bold indicates the best result and underline the second best, computed separately within the 8B group and globally for RAWR 14B.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Provider (API share)
Model
Input
Output
Output / input
Anthropic (40%)
Claude Fable 5.1
10.00
50.00
5.0×
Claude Opus 5.5
4.00
20.00
5.0×
Claude Sonnet 5
2.00
10.00
5.0×
Claude Haiku 4.5
1.00
5.00
5.0×
OpenAI (27%)
GPT-6 Astra
10.00
50.00
5.0×
GPT-6 Sol
2.00
10.00
5.0×
Appendix
Table 4: API list prices (USD per million tokens; standard tier, short context) of the three providers that together account for 88% of enterprise LLM API usage ( Tully et al., 2025 ) , accessed 26 September 2026 ( Anthropic, 2026 ; OpenAI, 2026 ; Google, 2026 ) . Every model bills output at 5 – 8.3× the input rate (median 5× ), the multiplier used in Table 5 .
Qwen PRM
MathShepherd
OmegaPRM
Granite PRM
RAWR
Input tokens
1.8×108
0.44×108
0.44×108
0.44×108
1.5×108
Output tokens
45.8×108
22.9×108
6.6×108
6.6×108
0
Cost-weighted (out =5× )
230.6×108
114.9×108
33.6×108
33.6×108
1.5×108
relative to RAWR
155×
77×
23×
23×
1×
relative to RAWR, raw tokens (out =1× )
32×
15.7×
4.8×
4.8×
1×
Appendix
Table 5: Tokens processed by each labeling method to produce the same number of labeled samples on PRM800K , split into input (prefill/scoring) and output (generation), after normalizing by each method’s retention rate ρ . CoT generation is shared across all methods and excluded. Rollout-based methods spend 94 – 98% of their tokens on output; RAWR generates none. The cost-weighted total prices output at 5× input, the median list-price ratio of the major API providers ( Table 4 ); the last two rows give the cost relative to RAWR at that rate and, for reference, counting raw tokens ( 1× ). QwenPRM uses the MathShepherd rollouts plus a separate LLM-as-a-judge filtering stage; its judge verdict tokens are excluded, so its cost is a lower bound. Granite PRM uses Automatic Process Supervision (the OmegaPRM MCTS labeling method), so we estimate its labeling cost as equal to OmegaPRM ’s; this counts the labeling stage only, not its additional fine-tuning.
Dataset
# Problems
# Candidate answers
MATH500 ( Lightman et al., 2023 )
12,000
85,248
GSM8K ( Cobbe et al., 2021 )
8,790
31,768
MathQA ( Amini et al., 2019 )
29,837
204,974
AquaRat ( Ling et al., 2017 )
97,700
589,446
NumGLUE ( Mishra et al., 2022 )
41,018
265,532
PubMedQA ( Jin et al., 2019 )
1,000
3,727
Appendix
Table 6: Training dataset composition.
Train
Test
Dataset
Avg # step
BoK
% correct
BoK
% correct
MATH500
8.50
86.57%
33.37%
88.98%
51.10%
GSM8K
7.52
98.55%
75.81%
97.72%
84.30%
MathQA
8.08
81.22%
38.52%
-
-
AquaRat
8.29
71.65%
28.19%
-
-
NumGLUE
5.80
86.05%
38.34%
-
-
Appendix
Table 7: Statistics of the generated CoTs. The BoK columns show the oracle best-of- K : the percentage of problems where at least one of the K samples is correct, whereas the % correct columns show the percentage of generations that are correct.
Figure 5: Distribution of the number of steps in the CoTs of all the datasets combined.
Figure 6: Average best-of- K performance over the evaluation tasks for PRMs trained on different subsets of the training set, ranging from 100K to the full 1.2M samples.
Aggregation
Baseline anchor ( r0 )
Successive diff. ( ri−1 )
max
82.01
80.46
sum
81.60
80.09
Appendix
Table 8: Balanced accuracy (%) of RAWR labels under each aggregation rule ( max , sum) and baseline anchor, micro-averaged across the 13 training datasets. max outperforms sum under both anchors, and anchoring at the pre-reasoning baseline ( NetInfo0 ) outperforms the strictly local successive difference ( NetInfoi−1 ) under every aggregation rule.
i/N
0.0
0.2
0.4
0.6
0.8
1.0
Correct ( n=484 )
0.149
0.153
0.165
0.376
0.886
0.995
Incorrect ( n=757 )
0.140
0.133
0.113
0.107
0.069
0.027
Appendix
Table 9: Median length-normalized gold-answer probability pi(y⋆) against the normalized step index i/N , teacher-forced on 1,241 ProcessBench traces. Correct and incorrect traces are indistinguishable before reasoning and diverge monotonically as it accumulates.
Balanced acc.
Traces
Full set
83.0%±0.1
85,248
High-prior ( p0>0.9 )
89.9%±0.8
1,298
Appendix
Table 10: Balanced accuracy of RAWR labels on the full MATH training set versus the high-prior subset ( p0>0.9 ), same protocol as the main paper. Spreads are ±1 std over a bootstrap. Labels on saturated priors are more accurate than on the full set.
Verifier noise ε
BIRD
Δ vs. clean
MATH
Δ vs. clean
0.00 (clean)
85.34%
—
81.59%
—
0.01
84.41%
−0.93
81.50%
−0.09
0.05
82.22%
−3.13
79.30%
−2.29
0.10
80.15%
−5.19
76.94%
−4.65
Appendix
Table 11: Balanced accuracy of RAWR labels under a verifier that mislabels each candidate with probability ε (7 seeds per level). Degradation is sub-linear and the labels remain usable across the full range tested.
s
5
10
20
30
40
512 (ref)
Balanced acc.
69.8%
72.7%
73.9%
74.2%
74.1%
77.9%
Gap to reference
8.1
5.2
4.0
3.7
3.8
—
Appendix
Table 12: Downstream balanced accuracy as a function of the number of subsampled candidates s (100 MATH problems, 512-candidate reference pool). Accuracy plateaus by s≈20 - 30 ; the paper’s K=32 sits on the plateau.
Figure 7: Example of a sample that is misclassified under IG but correctly identified under RAWR, illustrating the sharper discriminative behavior of the max-based information signal. The right axis shows the NetIG score of RAWR.
Figure 8: Error step offset distributions on ProcessBench . Left: QwenPRM 7B, tightly concentrated at offset 0, indicating precise step-level localization. Right: RAWR 8B, with the mode still at offset 0 but wider spread predominantly in [−2,+2] . The right skew (about 56 traces at positive offsets vs. 36 at negative ones) indicates a tendency to flag errors later than ground truth.
GSM8K
MATH
Olympiad
OmniMath
82.8%
77.2%
81.5%
80.2%
Appendix
Table 13: Percentage of flagged incorrect ProcessBench solutions whose first flagged step precedes the final step. An outcome-level signal would place its flag at the last step, giving values near 0 .
Reinforcement Learning with Verifiable Rewards (RLVR) is widely used to improve reasoning in large language models, but rewards only final-answer correctness with no supervision over intermediate steps. Rubric-based methods such as Rubrics as Rewards (RaR) introduce finer-grained supervision by scoring rollouts against structured criteria, yet the rubric scores are still aggregated into a single scalar applied to the entire response, causing three weaknesses: loss of multi-criterion structure, uniform supervision of correct and incorrect steps, and reward hacking through unbounded self-correction. On 1,000 problems, we find 18.2% of steps in correct-answer responses are wrong yet positively rewarded, while 49.9% of steps in incorrect-answer responses are correct yet penalized. We introduce Step-wise Rubrics as Rewards (SRaR), an RLVR framework that (i) uses an LLM judge to attribute each rubric item to a specific reasoning step, (ii) normalizes per-step rubric scores across rollouts so only steps whose quality varies produce a learning signal, and (iii) combines the per-step reward with the outcome reward through a decoupled advantage estimator that keeps the outcome baseline stable. We further build a 16K-problem rubric dataset by contrastively distilling rubric items from correct and flawed reasoning paths sampled from a strong model. Across six mathematical reasoning benchmarks, SRaR improves average accuracy over RaR by 3.57 points on Qwen3-8B and 2.75 points on Qwen3-32B, raises the Faithful Reasoning Rate on AIME 2025 from 34.5% to 46.7%, and reduces self-correction looping from 48.1% to 26.5%.
Weichu Xie, Haozhe Zhao, Wenpu Liu +15
Peking University · JD Explore Academy · Shanghai Jiao Tong University +2
Training process reward models (PRMs) requires step-level correctness labels, obtained either through expensive human annotation or by relying on ground-truth answers, limiting the ability to scale process-level supervision. We propose ScalePRM, which scales verification compute as an alternative: given a problem and a candidate solution, we generate multiple independent verifications of each reasoning step and aggregate their judgments to produce synthetic step-level labels without ground truth. We explore two representative inference-time scaling strategies, parallel scaling through self-consistency and sequential scaling through meta-critique, and train generative PRMs on the resulting synthetic data. On ProcessBench, a benchmark for identifying erroneous steps in mathematical reasoning, PRMs trained on step-level self-consistency data achieve 67.5 F1, surpassing reference-guided training with ground-truth access (66.4 F1) and GPT-4o as a critic (61.9 F1). When deployed as reward signals in RL training with Qwen2.5-Math-7B, our best PRM achieves 47.4% average accuracy across six mathematical reasoning benchmarks, outperforming ground-truth-based RLVR (43.9%). We also identify and address reward exploitation patterns unique to generative PRM-based RL. Our results demonstrate that scaling verification compute is a viable alternative to ground-truth supervision for training process reward models.
Salman Rahman, Sruthi Gorantla, Arpit Gupta +3
2UCLA · Work done while as an intern at Amazon AGI. · 1Amazon AGI
The reliability of large language models (LLMs) during test-time scaling is often assessed with \emph{external verifiers} or \emph{reward models} that distinguish correct reasoning from flawed logic. Prior work has studied both outcome reward models (ORMs), which assess only the final answer, and process reward models (PRMs), which score intermediate reasoning steps. Although PRMs are often viewed as advantageous due to their finer-grained supervision, much of the supporting evidence comes from math-adjacent settings, and their relative benefits across broader domains remain unclear. We present the first unified evaluation of four reward model variants, discriminative ORM and PRM (dORM, dPRM) and generative ORM and PRM (gORM, gPRM), across 14 diverse domains. Contrary to conventional wisdom, we find that (i) dORM performs on par with dPRM, (ii) gPRM is not competitive, and (iii) overall, gORM is the most robust, yielding significant and consistent gains across every tested domain. We attribute the worse performance of gPRM to the stepwise scoring process, which inherits label noise from LLM-based automatic labeling, leading to difficulties in evaluating long reasoning trajectories, including those involving self-correcting reasoning. Both our theoretical analysis and empirical observations indicate that stepwise aggregation compounds errors as reasoning length increases. These findings challenge the common assumption that fine-grained supervision is always better and support generative outcome verification for multi-domain deployment. Our \href{https://github.com/db-Lee/Multi-RM}{\underline{code}} is publicly available to facilitate future research in multi-domain settings.