Training long-horizon agents to solve complex tasks requires effective supervision over extended interaction sequences. However, sparse terminal rewards obscure intermediate contributions, while on-policy distillation can lose informative teacher guidance as student-generated histories grow. To address this problem, we introduce SCAD, which organizes interactions into planning and bounded subtask execution, distills execution in local contexts, and refines planning credit through cross-rollout subtask prefix trees, with planning receiving full terminal credit and execution receiving positive terminal credit and teacher guidance. Across all evaluated benchmarks, SCAD improves macro-average accuracy over the strongest training baseline by 4.48 percentage points for text tasks and 4.19 points for multimodal tasks. SCAD effectively combines outcome-based credit assignment with teacher-guided distillation to improve planning and execution in long-horizon agents.
Figures & tables
Figure 1: Target challenges and SCAD
Figure 2: SCAD training overview (left) and benchmark accuracy (%) across text and multimodal tasks (right; SCAD in red). Axes use separate zero-based scales; † marks out-of-distribution datasets.
Method
Organizing structure
How supervision is constructed
HiPER ( Peng et al., 2026 )
Subgoals and actions
Hierarchical advantages for planning and execution
GiGPO ( Feng et al., 2025 )
Shared states
Shared-state action comparisons plus trajectory credit
Tree-GRPO ( Ji et al., 2026 )
Rollout trees
Branch sampling and tree groups induce implicit step credit
TCOD ( WANG et al., 2026 )
Trajectory segments
Expands student-controlled spans for token distillation
MemOPD ( Liu et al., 2026c )
Invocation states
Reconstructs compact invocation states for RL and OPD
ATOD ( Tan et al., 2026 )
Interaction turns
Anneals RL/OPD mixing with turn-weighted distillation
Table 1: Comparison of how representative methods organize supervision for long-horizon agents. SCAD uses subtask boundaries to define local teacher contexts and subtask–report prefixes to estimate planning credit across completed rollouts. Planning retains signed terminal credit, while execution receives its positive part together with teacher guidance within each bounded subtask context.
Figure 3: SCAD training framework with local distillation and subtask prefix-tree credit. Planning uses full terminal credit; execution uses positive terminal credit and teacher guidance.
Method
Browse [-1pt]Comp-Plus
WebWalker [-1pt]QA
Hotpot
2Wiki
MuS
NQ
Trivia
Bamb. †
PopQA †
GAIA †
xbench †
Avg.
Teacher
23.61 ±0.5
36.67 ±3.1
60.51 ±0.9
52.31 ±1.5
11.28 ±0.9
68.72 ±1.8
85.13 ±0.9
56.92 ±0.0
56.92 ±0.0
25.42 ±0.0
23.67 ±0.6
45.56 ±0.6
Vanilla
0.00 ±0.0
2.92 ±0.7
31.28 ±0.9
21.54 ±0.0
11.79 ±0.9
28.21 ±0.9
46.67 ±1.8
24.10 ±0.9
21.03 ±0.9
2.82 ±1.0
3.67 ±0.6
17.64 ±0.5
GRPO
19.44 ±4.6
28.75 ±1.3
40.00 ±1.5
46.15 ±1.5
10.77 ±1.5
63.59 ±1.8
75.90 ±1.8
32.82 ±5.4
42.05 ±6.2
15.25 ±0.0
25.00 ±3.6
36.34 ±1.4
FoldGRPO
17.22 ±0.5
34.58 ±5.1
52.31 ±5.5
52.31 ±1.5
13.33 ±0.9
63.08 ±0.0
71.28 ±1.8
41.54 ±1.5
55.38 ±1.5
13.56 ±6.1
21.00 ±1.0
39.60 ±0.9
Tree-GRPO
18.33 ±2.2
35.83 ±0.7
50.77 ±0.0
51.28 ±1.8
18.46 ±1.5
61.03 ±0.9
78.46 ±8.1
38.46 ±4.1
52.31 ±1.5
18.64 ±1.7
25.33 ±4.5
40.81 ±0.3
HiPER
21.67 ±2.5
32.92 ±3.1
54.36 ±4.7
46.15 ±4.1
16.92 ±1.5
66.15 ±1.5
73.85 ±1.5
43.59 ±0.9
53.85 ±1.5
18.64 ±1.7
28.33 ±1.2
41.49 ±0.5
Table 2: Text accuracy (%); Qwen3-4B(student) / Qwen3-4B-Instruct-2507 (teacher). Mean ± std over 3 seeds. Avg.: equal-weight mean over 11 ID+OOD sets; †: OOD. Bold: best student result.
Method
MMSearch- [-1pt]Plus
VDR-Bench
VisBrowse- [-1pt]Bench
MMSearch
SimpleVQA
ID [-1pt]Avg.
FVQA †
InfoSeek †
LiveVQA †
All [-1pt]Avg.
Teacher
16.19 ±1.0
16.79 ±0.9
5.64 ±1.8
40.61 ±1.0
45.71 ±2.9
24.99 ±1.0
56.19 ±1.6
37.14 ±2.9
28.89 ±2.2
30.90 ±1.0
Vanilla
0.95 ±0.0
0.74 ±0.7
1.54 ±1.5
23.64 ±1.8
20.00 ±2.9
9.37 ±1.2
25.71 ±2.9
11.43 ±0.0
8.89 ±0.0
11.61 ±0.4
GRPO
3.49 ±2.4
3.46 ±1.9
0.51 ±0.9
40.00 ±1.8
42.86 ±0.0
18.06 ±0.5
47.62 ±1.6
27.62 ±1.6
19.26 ±2.6
23.10 ±0.4
MMSearch-R1
9.52 ±1.0
12.35 ±1.9
5.13 ±0.9
40.61 ±4.6
46.67 ±10.0
22.85 ±3.5
56.19 ±1.6
32.38 ±1.6
22.22 ±2.2
28.13 ±2.6
VSearcher
10.16 ±3.3
14.32 ±6.0
6.15 ±1.5
37.58 ±2.1
47.62 ±1.6
23.17 ±1.5
51.43 ±2.9
36.19 ±1.6
23.70 ±2.6
28.39 ±1.7
SFT
4.76 ±4.4
6.67 ±0.7
2.56 ±2.4
35.76 ±6.4
39.05 ±1.6
17.76 ±0.9
48.57 ±0.0
24.76 ±1.6
17.04 ±7.8
22.40 ±0.7
Table 3: Multimodal accuracy (%); Qwen3-VL-2B(student) / Qwen3-VL-4B (teacher). Mean ± std over 3 seeds. ID/All Avg.: equal-weight means over five/eight datasets; †: OOD. Bold: best student.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Browse [-1pt]Comp-Plus
WebWalker [-1pt]QA
Hotpot
2Wiki
MuS
NQ
Trivia
Bamb. †
PopQA †
GAIA †
xbench †
Avg.
Teacher
23.61 ±0.5
36.67 ±3.1
60.51 ±0.9
52.31 ±1.5
11.28 ±0.9
68.72 ±1.8
85.13 ±0.9
56.92 ±0.0
56.92 ±0.0
25.42 ±0.0
23.67 ±0.6
45.56 ±0.6
Vanilla
1.11 ±0.5
2.92 ±0.7
17.95 ±0.9
15.38 ±1.5
5.13 ±0.9
26.15 ±1.5
31.28 ±0.9
12.82 ±0.9
14.36 ±0.9
3.95 ±1.0
6.00 ±1.0
12.46 ±0.2
GRPO
9.44 ±0.5
27.92 ±1.9
41.03 ±2.4
28.21 ±3.2
7.69 ±3.1
40.51 ±2.4
73.85 ±3.1
24.10 ±0.9
39.49 ±1.8
7.34 ±2.0
25.67 ±0.6
29.57 ±0.6
FoldGRPO
14.17 ±0.8
31.67 ±2.6
45.13 ±3.2
37.44 ±0.9
10.26 ±0.9
47.18 ±2.4
77.44 ±1.8
26.67 ±4.4
44.62 ±1.5
9.04 ±2.6
29.33 ±1.5
33.90 ±1.0
Tree-GRPO
15.00 ±0.8
29.17 ±1.4
48.72 ±3.2
39.49 ±1.8
10.26 ±3.2
43.59 ±1.8
79.49 ±2.4
29.74 ±3.6
43.59 ±0.9
11.30 ±2.0
29.33 ±0.6
34.52 ±0.3
HiPER
16.39 ±1.3
32.50 ±1.2
42.56 ±3.2
39.49 ±1.8
10.77 ±2.7
45.64 ±3.9
80.00 ±4.1
30.26 ±0.9
42.56 ±3.2
10.17 ±3.4
30.00 ±3.0
34.58 ±1.7
Appendix
Table 4: Text accuracy (%); Qwen3-1.7B(student) / Qwen3-4B-Instruct-2507 (teacher). Mean ± std over 3 seeds. Avg.: macro accuracy over 11 ID+OOD datasets. †: OOD. Bold: best student.
Figure 10: Qwen3-1.7B training dynamics: (a) SCAD and ATOD accuracy; (b) SCAD reverse KL on local execution tokens.
Figure 11: Visit-count distributions with logarithmic count axes. (a) Np for root and non-root parents. (b) npa for all observed actions and the subset belonging to multi-candidate parents.
Planning-credit setting
Added tree term
Mean ∣A∣
Avg. (%)
No tree credit
0
0.729
44.32 ±1.5
Shrinkage only
Cp
1.082
45.02 ±1.1
Sibling only: V(p)=Qˉp
Spa
0.828
45.21 ±1.0
SCAD
Spa+Cp
1.181
46.10 ±1.2
Appendix
Table 5: Credit decomposition. Mean ∣A∣ : paired replay on 1,146 decisions from 37 complete mixed-outcome question groups. Accuracy (%): 11-dataset macro mean ± std over 3 seeds.
Subtask τs
Report τr
Avg. (%)
0.75
0.90
45.31 ±1.3
0.83
0.90
46.10 ±1.2
0.90
0.90
45.72 ±1.3
0.83
0.85
45.49 ±1.4
0.83
0.95
45.84 ±1.3
Appendix
Table 6: Threshold and prior-count sensitivity. Accuracy (%): ID+OOD macro mean ± std over 3 seeds; shading: default. Mean ∣A∣ uses the fixed decisions in Table A.2 .
Training long-horizon LLM agents with reinforcement learning is challenging because sparse outcome rewards reveal whether a task succeeds, but not which intermediate actions caused the outcome or how they should be corrected. Recent methods alleviate this issue by generating rewards or textual hints from turn-level action-output signals, or by using feedback-conditioned self-distillation. However, generating feedback at every turn is inefficient when many intermediate turns are already successful or neutral, and applying feedback at a fixed or misaligned turn often fails to supervise the actions that contributed to the failure. To bridge this gap, we propose HINT-SD, a targeted self-distillation framework that uses full-trajectory hindsight to select failure-relevant actions and applies feedback-conditioned distillation only on targeted action spans. Experiments on BFCL v3 and AppWorld show that our method improves over the dense per-turn feedback baseline by up to 18.80 percent while achieving 2.26× lower time per training step, suggesting that selecting where to distill is a key factor for both effective and efficient long-horizon agent training.
Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. However, existing methods for training these agents typically treat all steps within a trajectory uniformly during both supervised fine-tuning (SFT) and reinforcement learning (RL), failing to distinguish useful actions from erroneous or redundant ones. In this paper, we propose Answer-Backtracked Credit Assignment (ABC), a fine-grained credit assignment framework for training long-horizon search agents by converting sparse trajectory-level outcomes into dense step-level supervision that rewards useful actions (even in failed trajectories) while suppressing erroneous or redundant actions. Specifically, given a potentially obscure query and its corresponding ground-truth answer, ABC first performs Answer-Backtracked Clue Recovery, which traces back from the answer to recover intermediate clues required to solve the question. It then applies Clue-Anchored Step Scoring to evaluate each search step against these clues, converting sparse binary outcome supervision into dense step-level rewards. Based on these rewards, we develop ABC-SFT, which reweights the loss of each turn, and ABC-GRPO, which uses the step-level scores as rewards in GRPO. Building on this framework, we train ABSeeker based on Qwen3.5-4B with only 8.5k examples. ABSeeker achieves 37.3% on BrowseComp and 39.1% on BrowseComp-ZH. With context management, the scores further improve to 55.3% and 52.9%, respectively, significantly outperforming same-scale (4B) agents and even matching the performance of larger ones (approximately 30B). These results demonstrate the effectiveness of answer-backtracked step-level credit assignment for training long-horizon search agents.
Long-horizon tool-use reinforcement learning learns from outcome verification, but trajectory-level advantages are broadcast over reasoning, API, and answer tokens. Direct self-distillation can supply a denser signal, but in our experiments it can also destroy tool use by rehearsing teacher behavior without identifying which actions the verifier rewards. We introduce Sibling-Guided Credit Distillation (SGCD), which uses distillation for bounded credit weighting rather than as a competing actor loss. Dynamic sampling produces mixed successful and failed sibling rollouts; an external LLM summarizes their contrast into a training-only credit reference; and detached teacher/student divergence reshapes GRPO token advantages. The deployed student receives only the clean task prompt. Across AppWorld and tau^3-airline, SGCD reports higher held-out point estimates than GRPO-family comparators: AppWorld TGC improves from 42.9 to 45.6 on test_normal and from 24.7 to 27.0 on test_challenge, and tau^3-airline held-out evaluator score improves from 0.583 to 0.602. These results support a narrow design rule for long-horizon tool-use agents: use distillation to guide credit assignment while keeping policy gradient in charge of the actor update.
Tianyu Ding, Jianhong Xin, Juan Pablo De la Cruz Weinstein