Learning Process Rewards via Reasoning State Propagation
Authors: Kai Gan, Zi-Hao Zhou, Bo Ye, Jian Zhao, Min-Ling Zhang, Tong Wei
Organizations: School of Computer Science and Engineering, Southeast University, Nanjing 210096, China · Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China · Zhongguancun Academy · Zhongguancun Institute of Artificial Intelligence
Process reward models (PRMs) have demonstrated notable effectiveness in test-time scaling and reinforcement learning by providing fine-grained signals for evaluating intermediate reasoning states, but their training relies heavily on costly process annotations. A natural way to alleviate this dependence is to complement limited process supervision with scalable outcome supervision. However, existing PRMs often model reasoning prefixes independently, providing no explicit mechanism for effectively using final outcome to guide the learning of intermediate reasoning states. We introduce Reasoning State Propagation (RSP), which represents each reasoning prefix with a binary validity state and models transitions between successive states across the reasoning trajectory. Specifically, RSP predicts a break probability that a valid state becomes invalid and a repair probability that an invalid state returns to valid. By propagating these transitions, RSP connects intermediate states to the final state, allowing process annotations to supervise intermediate states while outcome labels supervise the final state and can provide learning signals to preceding steps. Across reasoning search, response selection, and reinforcement learning, RSP consistently outperforms representative PRM baselines, with average improvements over Qwen2.5-Math-PRM of 5.6% in beam search and 2.1% in reinforcement learning.
Figures & tables
Figure 1: (a) Empirical performance across evaluation settings. (b) RSP propagates a binary reasoning state across steps using break and repair probabilities. G denotes a Good or Valid state with no unresolved error, while B denotes a Bad or Invalid state with an unresolved error. Transitions G→B and B→G correspond to error introduction ( break ) and error recovery ( repair ), respectively.
PRM Backbone
Reward Model
MATH500
Gaokao
e=4
e=8
e=12
e=4
e=8
e=12
Qwen2.5-Math-7B
Qwen2.5-Math-PRM
74.8
74.0
75.8
58.5
55.6
59.3
Qwen3-1.7B
Supervised PRM
54.8
53.6
55.4
44.0
45.7
48.3
OVM
67.9
68.4
69.5
54.5
58.3
58.3
Pseudo-Label PRM
65.3
64.8
67.7
53.1
52.5
55.0
Joint-Supervised PRM
67.4
64.6
64.0
54.4
57.2
55.9
Table 1: Beam search accuracy (%) on MATH500 and Gaokao. The beam width is fixed to 4 , and e denotes the expansion width. The best and second-best results are shown in bold and underlined .
Method
Proc.
Out.
LLaMA3 70B
LLaMA3.3 70B
Qwen3 1.7B
Qwen3 4B
Qwen3 8B
DeepSeek-R1-0528 Qwen3-8B
Avg.
Pass@1
–
–
43.6
69.0
65.8
77.4
79.2
51.6
64.4
Supervised PRM
✓
–
57.2
74.6
76.0
80.8
79.9
46.0
69.1
OVM
–
✓
60.5
75.5
77.9
84.1
83.2
62.4
73.9
CRM
✓
–
58.2
77.8
79.1
83.1
82.5
52.1
72.1
Qwen2.5-Math-PRM
✓+
–
60.6
77.1
79.9
83.5
82.4
61.0
74.1
Pseudo-Label PRM
✓
✓
43.6
69.2
65.8
77.6
80.2
59.9
66.1
Table 2: Average Best-of- N accuracy (%) over N∈{8,16,32,64,128} . Proc./Out. denote process/outcome supervision, and ✓+ indicates additional large-scale process supervision.
Method
AIME
AMC
GSM8K
MATH500
Minerva
Olympiad
Avg.
Base Policy
23.3
70.0
92.5
78.6
38.2
52.4
59.2
GRPO
23.3
72.3
94.2
88.2
40.2
59.6
63.0
Supervised PRM
16.7
68.3
92.3
83.4
38.2
54.8
59.0
OVM
31.1
71.3
92.6
86.2
37.6
58.4
62.9
Qwen2.5-Math-PRM
31.1
68.3
93.8
86.8
39.8
59.8
63.3
Pseudo-Label PRM
26.7
72.3
92.6
85.8
35.8
57.3
61.8
Table 3: Avg@16 accuracy (%) after reinforcement learning. All methods use the same policy initialization, training data, and optimization budget.
Figure 2: (a) Distributions of predicted break scores for G→G and G→B . (b) Distributions of predicted repair scores for B→B and B→G . (c) Performance gains over Process-Only under increasing amounts of outcome supervision for beam search and reinforcement learning.
Variant
BoN
Beam Search
ProcessBench
RL
Process-Only
72.0 (-2.4)
59.7 (-12.5)
58.3 (-7.5)
59.8 (-5.6)
Outcome-Only
74.2 (-0.2)
69.5 (-2.7)
56.2 (-9.6)
61.6 (-3.8)
No Repair
73.6 (-0.8)
68.9 (-3.3)
64.5 (-1.3)
63.2 (-2.2)
Current Representation Only
73.5 (-0.9)
70.1 (-2.1)
64.2 (-1.6)
62.3 (-3.1)
Shared Token
73.5 (-0.9)
69.7 (-2.5)
63.5 (-2.3)
63.8 (-1.6)
No Outcome Propagation
73.9 (-0.5)
65.5 (-6.7)
61.8 (-4.0)
64.2 (-1.2)
Table 4: Ablation study across evaluation settings. Colored numbers indicate changes relative to Ours , with darker colors denoting larger degradation.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
GSM8K
MATH
OlympiadBench
Omni-MATH
Avg.
Supervised PRM
69.8
60.3
48.5
47.9
56.6
OVM
49.9
44.1
28.1
28.2
37.6
Qwen2.5-Math-PRM
81.8
67.3
66.8
66.2
70.5
Pseudo-Label PRM
1.9
2.1
2.0
2.2
2.1
Joint-Supervised PRM
70.1
66.2
57.0
52.8
61.5
CRM †
73.8
56.2
30.7
24.4
46.3
Appendix
Table 5: F1 scores (%) on ProcessBench. Our method classifies a step as valid when ptG≥0.5 . The last column reports the macro-average across all subsets.
Figure 3: (a) Distribution of κt=1−αt−βt . The signed value indicates the direction and strength of the dependence of the updated propagated state on the preceding state. (b) Empirical cumulative distribution of ∣κt∣ . Smaller ∣κt∣ indicates weaker dependence of the updated propagated state on the preceding propagated state.
Figure 4: Performance with different PRM backbone sizes. We compare RSP with Process-Only on beam search, Best-of- N selection, reinforcement learning, and ProcessBench.
Figure 5: Performance with different amounts of outcome supervision. The horizontal axis denotes the process-to-outcome data ratio within each training batch. For example, 1:3 indicates that the numbers of process-annotated and outcome-annotated examples are mixed at a ratio of 1:3 in a batch. The process supervision is kept fixed while the amount of outcome supervision is varied.