RAWR: Reward Assignment Without Rollouts in Verifiable Domains
Authors: Corentin Royer, Anna Hedström, Debarun Bhattacharjya, Gaetano Rossiello, Andrea Giovannini, Mennatallah El-Assady
Organizations: International Business Machines · Department of Computer Science, ETH Zurich, 8092 Zurich, Switzerland · ETH AI Center · Lirio · Department of Computer Science, ETH Zurich
Understanding and evaluating multi-step reasoning in LLMs at the level of individual steps remains a key challenge. Process reward models (PRMs) provide a solution by scoring each step, enabling fine-grained supervision and improved reliability. However, training them requires costly human annotation or computationally intensive rollout-based labeling. To solve this, we introduce MCNIG, a scalable method for automatically labeling the quality of individual reasoning steps in any verifiable domain. Its step score, net information gain (NetIG), improves upon single-reference information gain (IG) by comparing the most-supported correct answer against the most-supported incorrect one, yielding a robust signal even for long and structured outputs like code and SQL, where IG fails. We show that the signal produced by MCNIG correlates with human judgments of step quality, and we apply MCNIG labels to train PRMs that achieve the best average best-of-K accuracy across eight benchmarks spanning mathematics, code generation, text-to-SQL, and scientific QA. Crucially, MCNIG generates no rollouts, cutting labeling complexity to O(N) and making it up to X times cheaper than rollout-based methods at comparable label quality, which makes large-scale process supervision practical.
Figures & tables
Figure 1: Overview of RAWR. (1) An LLM samples K CoTs for question q . (2) A verifier Vq partitions them into correct ( Cq ) and incorrect ( Wq ) sets. (3) At every prefix r0,…,rn , we score each candidate answer and keep the max over Cq and Wq . (4) The step score, NetIG, is the change in the correct-incorrect gap relative to r0 , thresholded at τ to produce binary step labels.
Method
Label type
Complexity
Rollout-free
Cross-domain
PRM800K ( Lightman et al., 2023 )
Human expert
—
✓
✗
MathShepherd ( Wang et al., 2024a )
Rollout
O(N2)
✗
✗
OmegaPRM ( Luo et al., 2024 )
Rollout + binary search
O(NlogN)
✗
✗
QwenPRM ( Zhang et al., 2025 )
Rollout + consensus filtering
O(N2)
✗
✗
RAWR (ours)
Information Theoretic
O(N)
✓
✓
Table 1: Comparison of step-level supervision methods. RAWR requires no rollouts and achieves O(N) complexity. Relevant prior methods have been validated on mathematics only.
Figure 2: Left: Balanced accuracy of CoT -level labels for IG vs. RAWR across domains. Right: Accuracy stratified by the number of distinct correct answer formulations. IG degrades as formulation diversity increases, while RAWR remains stable.
Figure 3: ProcessBench Step F1 against labeling cost relative to RAWR ( Section A.1 ); top-left is best. Filtered QwenPRM adds an LLM-as-a-judge filtering stage to the same rollouts, raising F1 but discarding half of the labeled samples ( 155× ).
Figure 4: Best-of- K accuracy as a function of K , stratified by ground-truth answer length, for PRMs trained on IG labels and on RAWR labels. The RAWR- IG gap is negligible for short answers and widens as answers grow longer, confirming that single-reference IG degrades on structured, long-form outputs while RAWR remains robust. Shaded bands show ±1 standard deviation across repeated subsamples of K answers per problem.
Method
MATH
GSM
PubMed
HumanEval
BigCode
BIRD
AIME
Average
Single sampling
44.5%
73.3%
36.2%
55.5%
23.3%
19.6%
17.4%
38.5%
Majority voting
67.3%
91.2%
46.8%
66.1%
27.6%
37.6%
29.0%
52.2%
OVM
62.8%
92.0%
59.4%
76.3%
34.9%
45.1%
26.1%
56.6%
MathShepherd
56.1%
89.5%
49.7%
52.5%
29.5%
35.9%
23.0%
48.0%
ImplicitPRM
59.3%
89.9%
44.3%
77.8%
29.3%
35.6%
19.6%
50.8%
Granite PRM
62.5%
90.9%
49.8%
61.6%
31.4%
35.4%
27.0%
51.2%
Table 2: Best-of- K accuracy ( K=32 ) across eight benchmarks. Bold indicates the best result and underline the second best, computed separately within the 8B group and globally for RAWR 14B. Shaded rows are our methods; RAWR 8B achieves the highest average performance among 8B models.
Method
UGPhysics
ProcessBench Trace
ProcessBench Step
Single sampling
8.0%
—
—
Majority voting
11.5%
—
—
OVM
9.2%
33.3
18.6
MathShepherd
10.0%
41.1
31.5
ImplicitPRM
10.6%
68.6
21.2
Granite PRM
11.0%
51.9
51.4
Table 3: External evaluation. UGPhysics : best-of- K accuracy (OOD domain). ProcessBench Trace: F1 for classifying whether a trace contains an error (introduced by us). ProcessBench Step: official F1 for localizing the first incorrect step. Bold indicates the best result and underline the second best, computed separately within the 8B group and globally for RAWR 14B.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Provider (API share)
Model
Input
Output
Output / input
Anthropic (40%)
Claude Fable 5.1
10.00
50.00
5.0×
Claude Opus 5.5
4.00
20.00
5.0×
Claude Sonnet 5
2.00
10.00
5.0×
Claude Haiku 4.5
1.00
5.00
5.0×
OpenAI (27%)
GPT-6 Astra
10.00
50.00
5.0×
GPT-6 Sol
2.00
10.00
5.0×
Appendix
Table 4: API list prices (USD per million tokens; standard tier, short context) of the three providers that together account for 88% of enterprise LLM API usage ( Tully et al., 2025 ) , accessed 26 September 2026 ( Anthropic, 2026 ; OpenAI, 2026 ; Google, 2026 ) . Every model bills output at 5 – 8.3× the input rate (median 5× ), the multiplier used in Table 5 .
Qwen PRM
MathShepherd
OmegaPRM
Granite PRM
RAWR
Input tokens
1.8×108
0.44×108
0.44×108
0.44×108
1.5×108
Output tokens
45.8×108
22.9×108
6.6×108
6.6×108
0
Cost-weighted (out =5× )
230.6×108
114.9×108
33.6×108
33.6×108
1.5×108
relative to RAWR
155×
77×
23×
23×
1×
relative to RAWR, raw tokens (out =1× )
32×
15.7×
4.8×
4.8×
1×
Appendix
Table 5: Tokens processed by each labeling method to produce the same number of labeled samples on PRM800K , split into input (prefill/scoring) and output (generation), after normalizing by each method’s retention rate ρ . CoT generation is shared across all methods and excluded. Rollout-based methods spend 94 – 98% of their tokens on output; RAWR generates none. The cost-weighted total prices output at 5× input, the median list-price ratio of the major API providers ( Table 4 ); the last two rows give the cost relative to RAWR at that rate and, for reference, counting raw tokens ( 1× ). QwenPRM uses the MathShepherd rollouts plus a separate LLM-as-a-judge filtering stage; its judge verdict tokens are excluded, so its cost is a lower bound. Granite PRM uses Automatic Process Supervision (the OmegaPRM MCTS labeling method), so we estimate its labeling cost as equal to OmegaPRM ’s; this counts the labeling stage only, not its additional fine-tuning.
Dataset
# Problems
# Candidate answers
MATH500 ( Lightman et al., 2023 )
12,000
85,248
GSM8K ( Cobbe et al., 2021 )
8,790
31,768
MathQA ( Amini et al., 2019 )
29,837
204,974
AquaRat ( Ling et al., 2017 )
97,700
589,446
NumGLUE ( Mishra et al., 2022 )
41,018
265,532
PubMedQA ( Jin et al., 2019 )
1,000
3,727
Appendix
Table 6: Training dataset composition.
Train
Test
Dataset
Avg # step
BoK
% correct
BoK
% correct
MATH500
8.50
86.57%
33.37%
88.98%
51.10%
GSM8K
7.52
98.55%
75.81%
97.72%
84.30%
MathQA
8.08
81.22%
38.52%
-
-
AquaRat
8.29
71.65%
28.19%
-
-
NumGLUE
5.80
86.05%
38.34%
-
-
Appendix
Table 7: Statistics of the generated CoTs. The BoK columns show the oracle best-of- K : the percentage of problems where at least one of the K samples is correct, whereas the % correct columns show the percentage of generations that are correct.
Figure 5: Distribution of the number of steps in the CoTs of all the datasets combined.
Figure 6: Average best-of- K performance over the evaluation tasks for PRMs trained on different subsets of the training set, ranging from 100K to the full 1.2M samples.
Aggregation
Baseline anchor ( r0 )
Successive diff. ( ri−1 )
max
82.01
80.46
sum
81.60
80.09
Appendix
Table 8: Balanced accuracy (%) of RAWR labels under each aggregation rule ( max , sum) and baseline anchor, micro-averaged across the 13 training datasets. max outperforms sum under both anchors, and anchoring at the pre-reasoning baseline ( NetInfo0 ) outperforms the strictly local successive difference ( NetInfoi−1 ) under every aggregation rule.
i/N
0.0
0.2
0.4
0.6
0.8
1.0
Correct ( n=484 )
0.149
0.153
0.165
0.376
0.886
0.995
Incorrect ( n=757 )
0.140
0.133
0.113
0.107
0.069
0.027
Appendix
Table 9: Median length-normalized gold-answer probability pi(y⋆) against the normalized step index i/N , teacher-forced on 1,241 ProcessBench traces. Correct and incorrect traces are indistinguishable before reasoning and diverge monotonically as it accumulates.
Balanced acc.
Traces
Full set
83.0%±0.1
85,248
High-prior ( p0>0.9 )
89.9%±0.8
1,298
Appendix
Table 10: Balanced accuracy of RAWR labels on the full MATH training set versus the high-prior subset ( p0>0.9 ), same protocol as the main paper. Spreads are ±1 std over a bootstrap. Labels on saturated priors are more accurate than on the full set.
Verifier noise ε
BIRD
Δ vs. clean
MATH
Δ vs. clean
0.00 (clean)
85.34%
—
81.59%
—
0.01
84.41%
−0.93
81.50%
−0.09
0.05
82.22%
−3.13
79.30%
−2.29
0.10
80.15%
−5.19
76.94%
−4.65
Appendix
Table 11: Balanced accuracy of RAWR labels under a verifier that mislabels each candidate with probability ε (7 seeds per level). Degradation is sub-linear and the labels remain usable across the full range tested.
s
5
10
20
30
40
512 (ref)
Balanced acc.
69.8%
72.7%
73.9%
74.2%
74.1%
77.9%
Gap to reference
8.1
5.2
4.0
3.7
3.8
—
Appendix
Table 12: Downstream balanced accuracy as a function of the number of subsampled candidates s (100 MATH problems, 512-candidate reference pool). Accuracy plateaus by s≈20 - 30 ; the paper’s K=32 sits on the plateau.
Figure 7: Example of a sample that is misclassified under IG but correctly identified under RAWR, illustrating the sharper discriminative behavior of the max-based information signal. The right axis shows the NetIG score of RAWR.
Figure 8: Error step offset distributions on ProcessBench . Left: QwenPRM 7B, tightly concentrated at offset 0, indicating precise step-level localization. Right: RAWR 8B, with the mode still at offset 0 but wider spread predominantly in [−2,+2] . The right skew (about 56 traces at positive offsets vs. 36 at negative ones) indicates a tendency to flag errors later than ground truth.
GSM8K
MATH
Olympiad
OmniMath
82.8%
77.2%
81.5%
80.2%
Appendix
Table 13: Percentage of flagged incorrect ProcessBench solutions whose first flagged step precedes the final step. An outcome-level signal would place its flag at the last step, giving values near 0 .