Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents
Organizations: University of Edinburgh, UK · Huawei Technologies Research & Development (UK) Limited
Abstract
Search agents enable Large Language Models (LLMs) to iteratively retrieve and use information for complex multi-hop questions. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising approach for post-training such agents, but its reliance on sparse, outcome-based supervision can make credit assignment difficult and limit learning efficiency. In this paper, we systematically investigate how intermediate supervision can improve reinforcement learning for search agents. We study a range of reward-shaping and credit-assignment strategies that provide learning signals from intermediate retrieval steps. Building on these insights, we develop a training framework that combines intermediate signals with final outcome rewards to improve learning from multi-step search trajectories. Experiments across multiple benchmarks under matched training conditions demonstrate improvements in aggregate search-agent performance and show that both the choice of intermediate signal and where its credit is assigned affect training behaviour. These findings show that reward design and credit assignment are important design dimensions for training effective search agents.
Figures & tables
| Method | 2Wiki | HotpotQA | MuSiQue | Bamboogle | NQ | TriviaQA | PopQA | Avg |
|---|---|---|---|---|---|---|---|---|
| OO | 57.30 | 53.74 | 24.93 | 57.73 | 42.04 | 74.46 | 53.45 | 51.25 |
| Cov-scalar | 59.26 | 57.54 | 25.69 | 59.51 | 43.08 | 74.09 | 54.50 | 52.64 |
| Cov-Dep-scalar | 57.81 | 58.42 | 25.11 | 59.54 | 44.51 | 74.58 | 54.88 | 52.82 |
| AM-scalar | 59.27 | 57.11 | 25.89 | 60.21 | 43.25 | 74.45 | 54.16 | 52.66 |
| Cov-local | 62.82 | 58.91 | 28.84 | 61.91 | 43.43 | 75.86 | 54.37 | 54.34 |
| Cov-Dep-local | 60.96 | 59.49 | 25.49 | 58.29 | 42.01 | 76.09 | 52.42 | 52.96 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Value |
|---|---|
| Optimiser | AdamW |
| Learning rate | |
| Learning-rate schedule | Warmup followed by a constant learning rate |
| Warmup ratio | 0.285 |
| Rollouts per question | 5 |
| Training batch size | 128 |
| Rank reversal (%) | Advantage sign reversal (%) | |||
| Cov-Dep | AM | Cov-Dep | AM | |
| 0.05 | 0.19 | 0.38 | 0.14 | 0.42 |
| 0.10 | 0.38 | 0.76 | 0.42 | 0.92 |
| 0.25 | 0.57 | 1.49 | 1.02 | 2.49 |
| 0.50 | 0.73 | 2.45 | 2.03 | 4.39 |
| 1.00 | 1.99 | 3.90 | 4.62 | 7.29 |
| Construction category | Train | Dev |
|---|---|---|
| 1-hop, single | 1,200 | 150 |
| 1-hop, conjunctive | 1,300 | 150 |
| 2-hop, chain | 2,400 | 150 |
| 2-hop, wide | 900 | 100 |
| 3-hop-1, chain | 2,000 | 150 |
| 3-hop-1, wide | 700 | 50 |
| Check | Observed count |
|---|---|
| Selected train + development records | 15,200 |
| Grounded graph nodes (conjunctive subset) | 44,385 (8,050) |
| Missing prerequisites / cycles / missing node grounding | 0 / 0 / 0 |
| Gold annotation records with start-position coverage | 52,435 / 52,435 |
| Gold annotation records without full-mention chunk coverage | 265 / 52,435 |
| Trajectories affected by requiring full-mention coverage | 20 / 7,680 |
| Method | 2Wiki | HotpotQA | MuSiQue | Bamboogle | NQ | TriviaQA | PopQA | Avg |
|---|---|---|---|---|---|---|---|---|
| OO | 57.30 | 53.74 | 24.93 | 57.73 | 42.04 | 74.46 | 53.45 | 51.25 |
| Cov-scalar | 59.26 | 57.54 | 25.69 | 59.51 | 43.08 | 74.09 | 54.50 | 52.64 |
| Cov-Dep-scalar | 57.81 | 58.42 | 25.11 | 59.54 | 44.51 | 74.58 | 54.88 | 52.82 |
| AM-scalar | 59.27 | 57.11 | 25.89 | 60.21 | 43.25 | 74.45 | 54.16 | 52.66 |
| Cov-local | 62.82 | 58.91 | 28.84 | 61.91 | 43.43 | 75.86 | 54.37 | 54.34 |
| Cov-Dep-local | 60.96 | 59.49 | 25.49 | 58.29 | 42.01 | 76.09 | 52.42 | 52.96 |
| Method | 2Wiki | HotpotQA | MuSiQue | Bamboogle | NQ | TriviaQA | PopQA | Avg |
|---|---|---|---|---|---|---|---|---|
| OO | 45.90 | 41.80 | 13.28 | 46.40 | 24.22 | 62.89 | 45.51 | 39.22 |
| Cov-scalar | 48.44 | 45.51 | 15.82 | 44.80 | 26.76 | 63.09 | 46.68 | 41.19 |
| Cov-Dep-scalar | 46.48 | 46.48 | 15.04 | 46.40 | 27.73 | 64.45 | 47.85 | 41.54 |
| AM-scalar | 47.66 | 45.12 | 14.65 | 48.00 | 27.73 | 63.87 | 47.07 | 41.29 |
| Cov-local | 51.95 | 46.88 | 16.60 | 49.60 | 26.56 | 65.43 | 46.68 | 42.63 |
| Cov-Dep-local | 49.80 | 46.88 | 15.43 | 48.00 | 25.00 | 65.04 | 44.73 | 41.41 |
| Method | Multi-hop | Single-hop | All |
|---|---|---|---|
| OO | 46.26 | 56.65 | 51.25 |
| Cov-scalar | 48.40 | 57.22 | 52.64 |
| Cov-Dep-scalar | 48.05 | 57.99 | 52.82 |
| AM-scalar | 48.38 | 57.28 | 52.66 |
| Cov-local | 51.07 | 57.88 | 54.34 |
| Cov-Dep-local | 49.37 | 56.84 | 52.96 |
| Omitted dataset | Cov-local OO | AM-local OO |
|---|---|---|
| 2Wiki | +2.63 | +2.32 |
| HotpotQA | +2.70 | +1.66 |
| MuSiQue | +2.94 | +2.26 |
| Bamboogle | +3.05 | +2.30 |
| NQ | +3.42 | +2.47 |
| TriviaQA | +3.42 | +2.34 |