cs.CLOct 7, 2026

Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents

Authors: Wenyu Huang, Xinyu Hou, Pavlos Vougiouklis, Ruofei Lai, Jeff Z. Pan

Organizations: University of Edinburgh, UK · Huawei Technologies Research & Development (UK) Limited

Abstract

Search agents enable Large Language Models (LLMs) to iteratively retrieve and use information for complex multi-hop questions. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising approach for post-training such agents, but its reliance on sparse, outcome-based supervision can make credit assignment difficult and limit learning efficiency. In this paper, we systematically investigate how intermediate supervision can improve reinforcement learning for search agents. We study a range of reward-shaping and credit-assignment strategies that provide learning signals from intermediate retrieval steps. Building on these insights, we develop a training framework that combines intermediate signals with final outcome rewards to improve learning from multi-step search trajectories. Experiments across multiple benchmarks under matched training conditions demonstrate improvements in aggregate search-agent performance and show that both the choice of intermediate signal and where its credit is assigned affect training behaviour. These findings show that reward design and credit assignment are important design dimensions for training effective search agents.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?

    May 27, 2026Yibo Zhao, Zichen Ding, Jiayi Wu +2Search AgentsProcess Reward Model

  2. PiCA: Pivot-Based Credit Assignment for Search Agentic Reinforcement Learning

    May 10, 2026Dongyi Liu, Yifan Niu, Qinwen Wang +2Credit AssignmentAgentic Reinforcement Learning