cs.CLAug 31, 2026

PRO-Step: Step-level Process Reward Optimization for Retrieval-Augmented Generation

Authors: MinKeon Kim, Namjun Lee, Jaekwang Kim

Organizations: Department of Applied Artificial Intelligence, Convergence Program for Social Innovation, Sungkyunkwan University, Seoul, South Korea

Abstract

Retrieval-Augmented Generation enhances Large Language Models by grounding responses in external knowledge, but multi-hop reasoning remains vulnerable to error propagation, where early retrieval failures confound subsequent steps. Standard outcome-based optimization only rewards the final answer, leaving intermediate retrieval and reasoning errors undetected. While existing process-based methods introduce step-level signals, they still score each step against the final answer, rewarding spurious successes where flawed retrieval coincidentally produces the correct answer. Step-level supervision in RAG requires evaluating both logical validity and evidential grounding at each step. We introduce PRO-STEP: we train a generative PRM that evaluates both dimensions, employ PRM-guided value tree search to construct preference pairs contrasting valid steps against flawed ones, and optimize the policy via step-level Direct Preference Optimization. Experiments on single and multi-hop QA datasets demonstrate that PRO-STEP achieves the best average EM and F1 across five benchmarks. Code, models, and training data are publicly available at https://github.com/keemminnke/PRO-Step.

Explore similar work

CardsList
  1. Retrieval is Cheap, Show Me the Code: Executable Multi-Hop Reasoning for Retrieval-Augmented Generation

    May 13, 2026Jiashuo Sun, Jimeng Shi, Yixuan Xie +10Multi-Hop QARetrieval-Augmented Generation

  2. Latent Abstraction for Retrieval-Augmented Generation

    Apr 20, 2026Ha Lan N. T, Minh-Anh Nguyen, Dung D. LeRetrieval-Augmented GenerationAdaptive RAG

  3. Only Ask What You Don't Know: Grounded Delta Planning for Efficient Multi-step RAG

    Jun 21, 2026Wei-Chieh Chou, Xuanjun Chen, Jian-Ren Lin +3Multi-Hop QARetrieval-Augmented Generation