Search agents enable Large Language Models (LLMs) to iteratively retrieve and use information for complex multi-hop questions. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising approach for post-training such agents, but its reliance on sparse, outcome-based supervision can make credit assignment difficult and limit learning efficiency. In this paper, we systematically investigate how intermediate supervision can improve reinforcement learning for search agents. We study a range of reward-shaping and credit-assignment strategies that provide learning signals from intermediate retrieval steps. Building on these insights, we develop a training framework that combines intermediate signals with final outcome rewards to improve learning from multi-step search trajectories. Experiments across multiple benchmarks under matched training conditions demonstrate improvements in aggregate search-agent performance and show that both the choice of intermediate signal and where its credit is assigned affect training behaviour. These findings show that reward design and credit assignment are important design dimensions for training effective search agents.
Figures & tables
Figure 1: Observed coverage and schematic credit assignment. (a) An observed search matches v1 and v3 ; the missing father link at v2 blocks dependency credit for v3 , giving Cov 0.67 and Cov-Dep 0.33. Text is abridged. (b) Scalar changes the shared advantage; local retains the outcome advantage and adds credit on executed query tokens (Equations 4 – 5 ). Tool observations are masked. Colour shows support, not magnitude.
Method
2Wiki
HotpotQA
MuSiQue
Bamboogle
NQ
TriviaQA
PopQA
Avg
OO
57.30
53.74
24.93
57.73
42.04
74.46
53.45
51.25
Cov-scalar
59.26
57.54
25.69
59.51
43.08
74.09
54.50
52.64
Cov-Dep-scalar
57.81
58.42
25.11
59.54
44.51
74.58
54.88
52.82
AM-scalar
59.27
57.11
25.89
60.21
43.25
74.45
54.16
52.66
Cov-local
62.82
58.91
28.84
61.91
43.43
75.86
54.37
54.34
Cov-Dep-local
60.96
59.49
25.49
58.29
42.01
76.09
52.42
52.96
Table 1: Answer F1 (%) on all seven datasets. Avg summarises performance across all seven datasets. Cov measures grounded evidence coverage; Cov-Dep adds dependency constraints to the same coverage signal. AM uses retrieved-answer matching. Bold marks the highest value in each column. The blocks compare scalar and local incorporation for each signal. Figure 3 reports the local-credit interventions; complete F1 and EM results for all 13 conditions appear in Appendix C .
Figure 2: Scalar and local credit relative to OO across task groups. Each score averages over the group’s evaluation questions. Appendix C.1 gives full scores and benchmark composition.
Figure 3: F1 comparisons under three retrieval-credit interventions. (a) Removing closure helps local incorporation, but not scalar incorporation. (b) Aligned minus permuted F1 overall and within each task group; positive values favour aligned credit. (c) Using both outcome groups gives the highest aggregate F1 for Cov and AM; the dashed line denotes OO.
Figure 5
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Optimiser
AdamW
Learning rate
10−6
Learning-rate schedule
Warmup followed by a constant learning rate
Warmup ratio
0.285
Rollouts per question
5
Training batch size
128
Appendix
Table 2: Shared training hyperparameters.
Rank reversal (%)
Advantage sign reversal (%)
α
Cov-Dep
AM
Cov-Dep
AM
0.05
0.19
0.38
0.14
0.42
0.10
0.38
0.76
0.42
0.92
0.25
0.57
1.49
1.02
2.49
0.50
0.73
2.45
2.03
4.39
1.00
1.99
3.90
4.62
7.29
Appendix
Table 3: Effect of scalar reward coefficients on the primary sample. Rank reversal is measured over 2,617 within-group pairs with unequal original F1. Sign reversal is measured over 2,166 originally nonzero advantages in mixed-outcome groups. Both rates are computed by applying different coefficients to the same saved trajectories.
Construction category
Train
Dev
1-hop, single
1,200
150
1-hop, conjunctive
1,300
150
2-hop, chain
2,400
150
2-hop, wide
900
100
3-hop-1, chain
2,000
150
3-hop-1, wide
700
50
Appendix
Table 4: Training and development counts by construction category. Structural labels follow the naming convention of MuSiQue ( Trivedi et al., 2022 ) ; numbered suffixes distinguish graph templates at the same hop count. Categories account for shared supporting paragraphs and describe evidence structure, not search counts.
Gold annotation records with start-position coverage
52,435 / 52,435
Gold annotation records without full-mention chunk coverage
265 / 52,435
Trajectories affected by requiring full-mention coverage
20 / 7,680
Appendix
Table 5: Graph structure and chunk grounding in the training and development data, with retrieval-progress differences measured on 7,680 training trajectories.
Method
2Wiki
HotpotQA
MuSiQue
Bamboogle
NQ
TriviaQA
PopQA
Avg
OO
57.30
53.74
24.93
57.73
42.04
74.46
53.45
51.25
Cov-scalar
59.26
57.54
25.69
59.51
43.08
74.09
54.50
52.64
Cov-Dep-scalar
57.81
58.42
25.11
59.54
44.51
74.58
54.88
52.82
AM-scalar
59.27
57.11
25.89
60.21
43.25
74.45
54.16
52.66
Cov-local
62.82
58.91
28.84
61.91
43.43
75.86
54.37
54.34
Cov-Dep-local
60.96
59.49
25.49
58.29
42.01
76.09
52.42
52.96
Appendix
Table 6: Complete answer F1 (%) results for all 13 conditions on seven datasets. Avg follows the aggregation used in Table 1 ; bold marks the highest value in each column across all conditions.
Method
2Wiki
HotpotQA
MuSiQue
Bamboogle
NQ
TriviaQA
PopQA
Avg
OO
45.90
41.80
13.28
46.40
24.22
62.89
45.51
39.22
Cov-scalar
48.44
45.51
15.82
44.80
26.76
63.09
46.68
41.19
Cov-Dep-scalar
46.48
46.48
15.04
46.40
27.73
64.45
47.85
41.54
AM-scalar
47.66
45.12
14.65
48.00
27.73
63.87
47.07
41.29
Cov-local
51.95
46.88
16.60
49.60
26.56
65.43
46.68
42.63
Cov-Dep-local
49.80
46.88
15.43
48.00
25.00
65.04
44.73
41.41
Appendix
Table 7: Exact match (%) on all seven datasets. Avg follows the aggregation used in Table 1 ; bold marks the highest value in each column.
Figure 6: Differences in average EM and F1 from outcome-only training, on the same score scale as Table 1 . Both scalar and local incorporation improve EM and F1 over OO for all three signals; Cov-local gives the largest gains.
Figure 7: Per-dataset F1 differences from OO, on the same score scale as Table 1 . The heatmap shows how each retrieval signal and incorporation rule affects individual benchmarks.
Method
Multi-hop
Single-hop
All
OO
46.26
56.65
51.25
Cov-scalar
48.40
57.22
52.64
Cov-Dep-scalar
48.05
57.99
52.82
AM-scalar
48.38
57.28
52.66
Cov-local
51.07
57.88
54.34
Cov-Dep-local
49.37
56.84
52.96
Appendix
Table 8: F1 (%) within each QA task group and across all 3,197 examples. Each question has equal weight within the reported average. Bold marks the highest value in each column.
Omitted dataset
Cov-local − OO
AM-local − OO
2Wiki
+2.63
+2.32
HotpotQA
+2.70
+1.66
MuSiQue
+2.94
+2.26
Bamboogle
+3.05
+2.30
NQ
+3.42
+2.47
TriviaQA
+3.42
+2.34
Appendix
Table 9: Changes in average F1 after omitting one dataset. The full seven-dataset contrasts are +3.09 and +2.26 , respectively.
Search agents powered by large language models can autonomously decompose queries, retrieve information, and synthesize answers through multi-step reasoning. However, the rapid growth of training methods has outpaced controlled comparison: existing works differ in retrieval corpora, reward designs, and training protocols, making it unclear what actually drives improvements. We present a controlled empirical study that isolates three under-explored dimensions of search agent training. First, we identify a critical data-coverage issue in the widely used Wikipedia 2018 corpus and show that correcting it alone yields larger gains than the differences between training algorithms. Second, we systematically compare outcome-based and process-based reward methods across three base models, finding that the simplest outcome-based approach achieves competitive or superior performance in most settings, and that process-level credit assignment can over-correct agent behavior. Third, we analyze training data diversity, off-policy data utilization, and search budget scaling, distilling practical guidelines for training effective search agents. Our code is available at https://github.com/YiboZhao624/SearchAgentReview.
Yibo Zhao, Zichen Ding, Jiayi Wu +2
School of Data Science and Engineering, East China Normal University · Shanghai AI Laboratory
Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. However, existing methods for training these agents typically treat all steps within a trajectory uniformly during both supervised fine-tuning (SFT) and reinforcement learning (RL), failing to distinguish useful actions from erroneous or redundant ones. In this paper, we propose Answer-Backtracked Credit Assignment (ABC), a fine-grained credit assignment framework for training long-horizon search agents by converting sparse trajectory-level outcomes into dense step-level supervision that rewards useful actions (even in failed trajectories) while suppressing erroneous or redundant actions. Specifically, given a potentially obscure query and its corresponding ground-truth answer, ABC first performs Answer-Backtracked Clue Recovery, which traces back from the answer to recover intermediate clues required to solve the question. It then applies Clue-Anchored Step Scoring to evaluate each search step against these clues, converting sparse binary outcome supervision into dense step-level rewards. Based on these rewards, we develop ABC-SFT, which reweights the loss of each turn, and ABC-GRPO, which uses the step-level scores as rewards in GRPO. Building on this framework, we train ABSeeker based on Qwen3.5-4B with only 8.5k examples. ABSeeker achieves 37.3% on BrowseComp and 39.1% on BrowseComp-ZH. With context management, the scores further improve to 55.3% and 52.9%, respectively, significantly outperforming same-scale (4B) agents and even matching the performance of larger ones (approximately 30B). These results demonstrate the effectiveness of answer-backtracked step-level credit assignment for training long-horizon search agents.
Large Language Model (LLM)-based search agents trained with reinforcement learning (RL) have significantly improved the performance of knowledge-intensive tasks. However, existing methods encounter critical challenges in long-horizon credit assignment: (i) Reward Sparsity, where models receive only outcome feedback without step-level guidance to differentiate action quality; (ii) Isolated Credit, where credit is assigned to steps independently, failing to capture sequential dependencies; and (iii) Distributional Shift, where rewards are estimated on templates that deviate from the model's natural generative distribution. To address these issues, we propose Pivot-Based Credit Assignment (PiCA), a novel step reward mechanism that reformulates the search trajectory as a sequential process of cumulative search progress. Unlike prior isolated step rewards, PiCA defines process rewards as success probabilities dependent on the historical context based on Potential-Based Reward Shaping (PBRS). This approach identifies pivot steps, which comprise target golden sub-queries and sub-answers derived from historical trajectories, as information peaks that significantly boost the likelihood of a correct final answer. By anchoring these step rewards to the final task objective, PiCA provides dense, pivot-aware and trajectory-dependent guidance while maintaining distributional consistency. Extensive experiments show that PiCA outperforms existing strong baselines across seven knowledge-intensive QA benchmarks, achieving 15.2% and 4.4% improvements for 3B and 7B models. The consistent performance gains across various models show PiCA's robust generalization. The code is available at https://github.com/novdream/PiCA.
Dongyi Liu, Yifan Niu, Qinwen Wang +2
The Hong Kong University of Science and Technology (Guangzhou) · The Hong Kong University of Science and Technology