Learning Process Rewards via Reasoning State Propagation
Authors: Kai Gan, Zi-Hao Zhou, Bo Ye, Jian Zhao, Min-Ling Zhang, Tong Wei
Organizations: School of Computer Science and Engineering, Southeast University, Nanjing 210096, China · Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China · Zhongguancun Academy · Zhongguancun Institute of Artificial Intelligence
Process reward models (PRMs) have demonstrated notable effectiveness in test-time scaling and reinforcement learning by providing fine-grained signals for evaluating intermediate reasoning states, but their training relies heavily on costly process annotations. A natural way to alleviate this dependence is to complement limited process supervision with scalable outcome supervision. However, existing PRMs often model reasoning prefixes independently, providing no explicit mechanism for effectively using final outcome to guide the learning of intermediate reasoning states. We introduce Reasoning State Propagation (RSP), which represents each reasoning prefix with a binary validity state and models transitions between successive states across the reasoning trajectory. Specifically, RSP predicts a break probability that a valid state becomes invalid and a repair probability that an invalid state returns to valid. By propagating these transitions, RSP connects intermediate states to the final state, allowing process annotations to supervise intermediate states while outcome labels supervise the final state and can provide learning signals to preceding steps. Across reasoning search, response selection, and reinforcement learning, RSP consistently outperforms representative PRM baselines, with average improvements over Qwen2.5-Math-PRM of 5.6% in beam search and 2.1% in reinforcement learning.
Figures & tables
Figure 1: (a) Empirical performance across evaluation settings. (b) RSP propagates a binary reasoning state across steps using break and repair probabilities. G denotes a Good or Valid state with no unresolved error, while B denotes a Bad or Invalid state with an unresolved error. Transitions G→B and B→G correspond to error introduction ( break ) and error recovery ( repair ), respectively.
PRM Backbone
Reward Model
MATH500
Gaokao
e=4
e=8
e=12
e=4
e=8
e=12
Qwen2.5-Math-7B
Qwen2.5-Math-PRM
74.8
74.0
75.8
58.5
55.6
59.3
Qwen3-1.7B
Supervised PRM
54.8
53.6
55.4
44.0
45.7
48.3
OVM
67.9
68.4
69.5
54.5
58.3
58.3
Pseudo-Label PRM
65.3
64.8
67.7
53.1
52.5
55.0
Joint-Supervised PRM
67.4
64.6
64.0
54.4
57.2
55.9
Table 1: Beam search accuracy (%) on MATH500 and Gaokao. The beam width is fixed to 4 , and e denotes the expansion width. The best and second-best results are shown in bold and underlined .
Method
Proc.
Out.
LLaMA3 70B
LLaMA3.3 70B
Qwen3 1.7B
Qwen3 4B
Qwen3 8B
DeepSeek-R1-0528 Qwen3-8B
Avg.
Pass@1
–
–
43.6
69.0
65.8
77.4
79.2
51.6
64.4
Supervised PRM
✓
–
57.2
74.6
76.0
80.8
79.9
46.0
69.1
OVM
–
✓
60.5
75.5
77.9
84.1
83.2
62.4
73.9
CRM
✓
–
58.2
77.8
79.1
83.1
82.5
52.1
72.1
Qwen2.5-Math-PRM
✓+
–
60.6
77.1
79.9
83.5
82.4
61.0
74.1
Pseudo-Label PRM
✓
✓
43.6
69.2
65.8
77.6
80.2
59.9
66.1
Table 2: Average Best-of- N accuracy (%) over N∈{8,16,32,64,128} . Proc./Out. denote process/outcome supervision, and ✓+ indicates additional large-scale process supervision.
Method
AIME
AMC
GSM8K
MATH500
Minerva
Olympiad
Avg.
Base Policy
23.3
70.0
92.5
78.6
38.2
52.4
59.2
GRPO
23.3
72.3
94.2
88.2
40.2
59.6
63.0
Supervised PRM
16.7
68.3
92.3
83.4
38.2
54.8
59.0
OVM
31.1
71.3
92.6
86.2
37.6
58.4
62.9
Qwen2.5-Math-PRM
31.1
68.3
93.8
86.8
39.8
59.8
63.3
Pseudo-Label PRM
26.7
72.3
92.6
85.8
35.8
57.3
61.8
Table 3: Avg@16 accuracy (%) after reinforcement learning. All methods use the same policy initialization, training data, and optimization budget.
Figure 2: (a) Distributions of predicted break scores for G→G and G→B . (b) Distributions of predicted repair scores for B→B and B→G . (c) Performance gains over Process-Only under increasing amounts of outcome supervision for beam search and reinforcement learning.
Variant
BoN
Beam Search
ProcessBench
RL
Process-Only
72.0 (-2.4)
59.7 (-12.5)
58.3 (-7.5)
59.8 (-5.6)
Outcome-Only
74.2 (-0.2)
69.5 (-2.7)
56.2 (-9.6)
61.6 (-3.8)
No Repair
73.6 (-0.8)
68.9 (-3.3)
64.5 (-1.3)
63.2 (-2.2)
Current Representation Only
73.5 (-0.9)
70.1 (-2.1)
64.2 (-1.6)
62.3 (-3.1)
Shared Token
73.5 (-0.9)
69.7 (-2.5)
63.5 (-2.3)
63.8 (-1.6)
No Outcome Propagation
73.9 (-0.5)
65.5 (-6.7)
61.8 (-4.0)
64.2 (-1.2)
Table 4: Ablation study across evaluation settings. Colored numbers indicate changes relative to Ours , with darker colors denoting larger degradation.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
GSM8K
MATH
OlympiadBench
Omni-MATH
Avg.
Supervised PRM
69.8
60.3
48.5
47.9
56.6
OVM
49.9
44.1
28.1
28.2
37.6
Qwen2.5-Math-PRM
81.8
67.3
66.8
66.2
70.5
Pseudo-Label PRM
1.9
2.1
2.0
2.2
2.1
Joint-Supervised PRM
70.1
66.2
57.0
52.8
61.5
CRM †
73.8
56.2
30.7
24.4
46.3
Appendix
Table 5: F1 scores (%) on ProcessBench. Our method classifies a step as valid when ptG≥0.5 . The last column reports the macro-average across all subsets.
Figure 3: (a) Distribution of κt=1−αt−βt . The signed value indicates the direction and strength of the dependence of the updated propagated state on the preceding state. (b) Empirical cumulative distribution of ∣κt∣ . Smaller ∣κt∣ indicates weaker dependence of the updated propagated state on the preceding propagated state.
Figure 4: Performance with different PRM backbone sizes. We compare RSP with Process-Only on beam search, Best-of- N selection, reinforcement learning, and ProcessBench.
Figure 5: Performance with different amounts of outcome supervision. The horizontal axis denotes the process-to-outcome data ratio within each training batch. For example, 1:3 indicates that the numbers of process-annotated and outcome-annotated examples are mixed at a ratio of 1:3 in a batch. The process supervision is kept fixed while the amount of outcome supervision is varied.
Process Reward Models (PRMs) are a powerful mechanism for steering large language model reasoning by providing fine-grained, step-level supervision. However, this effectiveness comes at a significant cost: PRMs require expert annotations for every reasoning step, making them costly and difficult to scale. Here, we propose a method for training unsupervised PRMs (uPRM) that requires no human supervision, neither at the level of step-by-step annotations nor through ground-truth verification of final answers. The key idea behind our approach is to define a scoring function, derived from LLM next-token probabilities, that jointly assesses candidate positions of first erroneous steps across a batch of reasoning trajectories. We demonstrate the effectiveness of uPRM across diverse scenarios: (i) uPRM achieves up to 15% absolute accuracy improvements over the LLM-as-a-Judge in identifying first erroneous steps on the ProcessBench dataset; (ii) as a verifier for test-time scaling, uPRM performs comparably to supervised PRMs and outperforms the majority voting baseline by up to 6.9%, and (iii) when used as a reward signal in reinforcement learning, uPRM enables more robust policy optimization throughout training compared to a supervised PRM trained using ground-truth labels. Overall, our results open a path toward scalable reward modeling for complex reasoning tasks.
Training process reward models (PRMs) requires step-level correctness labels, obtained either through expensive human annotation or by relying on ground-truth answers, limiting the ability to scale process-level supervision. We propose ScalePRM, which scales verification compute as an alternative: given a problem and a candidate solution, we generate multiple independent verifications of each reasoning step and aggregate their judgments to produce synthetic step-level labels without ground truth. We explore two representative inference-time scaling strategies, parallel scaling through self-consistency and sequential scaling through meta-critique, and train generative PRMs on the resulting synthetic data. On ProcessBench, a benchmark for identifying erroneous steps in mathematical reasoning, PRMs trained on step-level self-consistency data achieve 67.5 F1, surpassing reference-guided training with ground-truth access (66.4 F1) and GPT-4o as a critic (61.9 F1). When deployed as reward signals in RL training with Qwen2.5-Math-7B, our best PRM achieves 47.4% average accuracy across six mathematical reasoning benchmarks, outperforming ground-truth-based RLVR (43.9%). We also identify and address reward exploitation patterns unique to generative PRM-based RL. Our results demonstrate that scaling verification compute is a viable alternative to ground-truth supervision for training process reward models.
Salman Rahman, Sruthi Gorantla, Arpit Gupta +3
2UCLA · Work done while as an intern at Amazon AGI. · 1Amazon AGI
Process rewards have been widely used in deep reinforcement learning to improve training efficiency, reduce variance, and prevent reward hacking. In LLM reasoning, existing works also explore various solutions for learning effective process reward models (PRM) with or without the help of an expert policy. However, existing methods either rely on strong assumptions about the expert policies (e.g., requiring their reward functions) or suffer intrinsic limitations (e.g., entropy collapse), resulting in weak PRMs or limited generalizability. In this paper, we introduce rePIRL, an inverse RL-inspired framework that learns effective PRMs with minimal assumptions about expert policies. Specifically, we design a dual learning process that updates the policy and the PRM interchangeably. Our learning algorithm has customized techniques to address the challenges of scaling traditional inverse RL to LLMs. We theoretically show that our proposed learning framework can unify both online and offline PRM learning methods, justifying that rePIRL can learn PRMs with minimal assumptions. Empirical evaluations on standardized math and coding reasoning datasets demonstrate the effectiveness of rePIRL over existing methods. We further show the application of our trained PRM in test-time training, test-time scaling, and providing an early signal for training hard problems. Finally, we validate our training recipe and key design choices via a detailed ablation study.
Xian Wu, Kaijie Zhu, Ying Zhang +2
Meta AI · Department of Computer Science, University of California, Santa Barbara · Independent Researcher +1