cs.AISep 30, 2026

Learning Process Rewards via Reasoning State Propagation

Authors: Kai Gan, Zi-Hao Zhou, Bo Ye, Jian Zhao, Min-Ling Zhang, Tong Wei

Organizations: School of Computer Science and Engineering, Southeast University, Nanjing 210096, China · Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China · Zhongguancun Academy · Zhongguancun Institute of Artificial Intelligence

Abstract

Process reward models (PRMs) have demonstrated notable effectiveness in test-time scaling and reinforcement learning by providing fine-grained signals for evaluating intermediate reasoning states, but their training relies heavily on costly process annotations. A natural way to alleviate this dependence is to complement limited process supervision with scalable outcome supervision. However, existing PRMs often model reasoning prefixes independently, providing no explicit mechanism for effectively using final outcome to guide the learning of intermediate reasoning states. We introduce Reasoning State Propagation (RSP), which represents each reasoning prefix with a binary validity state and models transitions between successive states across the reasoning trajectory. Specifically, RSP predicts a break probability that a valid state becomes invalid and a repair probability that an invalid state returns to valid. By propagating these transitions, RSP connects intermediate states to the final state, allowing process annotations to supervise intermediate states while outcome labels supervise the final state and can provide learning signals to preceding steps. Across reasoning search, response selection, and reinforcement learning, RSP consistently outperforms representative PRM baselines, with average improvements over Qwen2.5-Math-PRM of 5.6% in beam search and 2.1% in reinforcement learning.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Unsupervised Process Reward Models

    May 11, 2026Artyom Gadetsky, Maxim Kodryan, Siba Smarak Panigrahi +2Process Reward ModelLLM Reasoning Strategies

  2. ScalePRM: Training Process Reward Models by Scaling Verification Compute Without Ground Truth

    Dec 2, 2025Salman Rahman, Sruthi Gorantla, Arpit Gupta +3Process Reward ModelMathematical Reasoning Benchmarks

  3. rePIRL: Learn PRM with Inverse RL for LLM Reasoning

    Feb 8, 2026Xian Wu, Kaijie Zhu, Ying Zhang +2Process Reward ModelLarge Language Model Reinforcement Learning