Training long-horizon agents to solve complex tasks requires effective supervision over extended interaction sequences. However, sparse terminal rewards obscure intermediate contributions, while on-policy distillation can lose informative teacher guidance as student-generated histories grow. To address this problem, we introduce SCAD, which organizes interactions into planning and bounded subtask execution, distills execution in local contexts, and refines planning credit through cross-rollout subtask prefix trees, with planning receiving full terminal credit and execution receiving positive terminal credit and teacher guidance. Across all evaluated benchmarks, SCAD improves macro-average accuracy over the strongest training baseline by 4.48 percentage points for text tasks and 4.19 points for multimodal tasks. SCAD effectively combines outcome-based credit assignment with teacher-guided distillation to improve planning and execution in long-horizon agents.
Figures & tables
Figure 1: Target challenges and SCAD
Figure 2: SCAD training overview (left) and benchmark accuracy (%) across text and multimodal tasks (right; SCAD in red). Axes use separate zero-based scales; † marks out-of-distribution datasets.
Method
Organizing structure
How supervision is constructed
HiPER ( Peng et al., 2026 )
Subgoals and actions
Hierarchical advantages for planning and execution
GiGPO ( Feng et al., 2025 )
Shared states
Shared-state action comparisons plus trajectory credit
Tree-GRPO ( Ji et al., 2026 )
Rollout trees
Branch sampling and tree groups induce implicit step credit
TCOD ( WANG et al., 2026 )
Trajectory segments
Expands student-controlled spans for token distillation
MemOPD ( Liu et al., 2026c )
Invocation states
Reconstructs compact invocation states for RL and OPD
ATOD ( Tan et al., 2026 )
Interaction turns
Anneals RL/OPD mixing with turn-weighted distillation
Table 1: Comparison of how representative methods organize supervision for long-horizon agents. SCAD uses subtask boundaries to define local teacher contexts and subtask–report prefixes to estimate planning credit across completed rollouts. Planning retains signed terminal credit, while execution receives its positive part together with teacher guidance within each bounded subtask context.
Figure 3: SCAD training framework with local distillation and subtask prefix-tree credit. Planning uses full terminal credit; execution uses positive terminal credit and teacher guidance.
Method
Browse [-1pt]Comp-Plus
WebWalker [-1pt]QA
Hotpot
2Wiki
MuS
NQ
Trivia
Bamb. †
PopQA †
GAIA †
xbench †
Avg.
Teacher
23.61 ±0.5
36.67 ±3.1
60.51 ±0.9
52.31 ±1.5
11.28 ±0.9
68.72 ±1.8
85.13 ±0.9
56.92 ±0.0
56.92 ±0.0
25.42 ±0.0
23.67 ±0.6
45.56 ±0.6
Vanilla
0.00 ±0.0
2.92 ±0.7
31.28 ±0.9
21.54 ±0.0
11.79 ±0.9
28.21 ±0.9
46.67 ±1.8
24.10 ±0.9
21.03 ±0.9
2.82 ±1.0
3.67 ±0.6
17.64 ±0.5
GRPO
19.44 ±4.6
28.75 ±1.3
40.00 ±1.5
46.15 ±1.5
10.77 ±1.5
63.59 ±1.8
75.90 ±1.8
32.82 ±5.4
42.05 ±6.2
15.25 ±0.0
25.00 ±3.6
36.34 ±1.4
FoldGRPO
17.22 ±0.5
34.58 ±5.1
52.31 ±5.5
52.31 ±1.5
13.33 ±0.9
63.08 ±0.0
71.28 ±1.8
41.54 ±1.5
55.38 ±1.5
13.56 ±6.1
21.00 ±1.0
39.60 ±0.9
Tree-GRPO
18.33 ±2.2
35.83 ±0.7
50.77 ±0.0
51.28 ±1.8
18.46 ±1.5
61.03 ±0.9
78.46 ±8.1
38.46 ±4.1
52.31 ±1.5
18.64 ±1.7
25.33 ±4.5
40.81 ±0.3
HiPER
21.67 ±2.5
32.92 ±3.1
54.36 ±4.7
46.15 ±4.1
16.92 ±1.5
66.15 ±1.5
73.85 ±1.5
43.59 ±0.9
53.85 ±1.5
18.64 ±1.7
28.33 ±1.2
41.49 ±0.5
Table 2: Text accuracy (%); Qwen3-4B(student) / Qwen3-4B-Instruct-2507 (teacher). Mean ± std over 3 seeds. Avg.: equal-weight mean over 11 ID+OOD sets; †: OOD. Bold: best student result.
Method
MMSearch- [-1pt]Plus
VDR-Bench
VisBrowse- [-1pt]Bench
MMSearch
SimpleVQA
ID [-1pt]Avg.
FVQA †
InfoSeek †
LiveVQA †
All [-1pt]Avg.
Teacher
16.19 ±1.0
16.79 ±0.9
5.64 ±1.8
40.61 ±1.0
45.71 ±2.9
24.99 ±1.0
56.19 ±1.6
37.14 ±2.9
28.89 ±2.2
30.90 ±1.0
Vanilla
0.95 ±0.0
0.74 ±0.7
1.54 ±1.5
23.64 ±1.8
20.00 ±2.9
9.37 ±1.2
25.71 ±2.9
11.43 ±0.0
8.89 ±0.0
11.61 ±0.4
GRPO
3.49 ±2.4
3.46 ±1.9
0.51 ±0.9
40.00 ±1.8
42.86 ±0.0
18.06 ±0.5
47.62 ±1.6
27.62 ±1.6
19.26 ±2.6
23.10 ±0.4
MMSearch-R1
9.52 ±1.0
12.35 ±1.9
5.13 ±0.9
40.61 ±4.6
46.67 ±10.0
22.85 ±3.5
56.19 ±1.6
32.38 ±1.6
22.22 ±2.2
28.13 ±2.6
VSearcher
10.16 ±3.3
14.32 ±6.0
6.15 ±1.5
37.58 ±2.1
47.62 ±1.6
23.17 ±1.5
51.43 ±2.9
36.19 ±1.6
23.70 ±2.6
28.39 ±1.7
SFT
4.76 ±4.4
6.67 ±0.7
2.56 ±2.4
35.76 ±6.4
39.05 ±1.6
17.76 ±0.9
48.57 ±0.0
24.76 ±1.6
17.04 ±7.8
22.40 ±0.7
Table 3: Multimodal accuracy (%); Qwen3-VL-2B(student) / Qwen3-VL-4B (teacher). Mean ± std over 3 seeds. ID/All Avg.: equal-weight means over five/eight datasets; †: OOD. Bold: best student.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Browse [-1pt]Comp-Plus
WebWalker [-1pt]QA
Hotpot
2Wiki
MuS
NQ
Trivia
Bamb. †
PopQA †
GAIA †
xbench †
Avg.
Teacher
23.61 ±0.5
36.67 ±3.1
60.51 ±0.9
52.31 ±1.5
11.28 ±0.9
68.72 ±1.8
85.13 ±0.9
56.92 ±0.0
56.92 ±0.0
25.42 ±0.0
23.67 ±0.6
45.56 ±0.6
Vanilla
1.11 ±0.5
2.92 ±0.7
17.95 ±0.9
15.38 ±1.5
5.13 ±0.9
26.15 ±1.5
31.28 ±0.9
12.82 ±0.9
14.36 ±0.9
3.95 ±1.0
6.00 ±1.0
12.46 ±0.2
GRPO
9.44 ±0.5
27.92 ±1.9
41.03 ±2.4
28.21 ±3.2
7.69 ±3.1
40.51 ±2.4
73.85 ±3.1
24.10 ±0.9
39.49 ±1.8
7.34 ±2.0
25.67 ±0.6
29.57 ±0.6
FoldGRPO
14.17 ±0.8
31.67 ±2.6
45.13 ±3.2
37.44 ±0.9
10.26 ±0.9
47.18 ±2.4
77.44 ±1.8
26.67 ±4.4
44.62 ±1.5
9.04 ±2.6
29.33 ±1.5
33.90 ±1.0
Tree-GRPO
15.00 ±0.8
29.17 ±1.4
48.72 ±3.2
39.49 ±1.8
10.26 ±3.2
43.59 ±1.8
79.49 ±2.4
29.74 ±3.6
43.59 ±0.9
11.30 ±2.0
29.33 ±0.6
34.52 ±0.3
HiPER
16.39 ±1.3
32.50 ±1.2
42.56 ±3.2
39.49 ±1.8
10.77 ±2.7
45.64 ±3.9
80.00 ±4.1
30.26 ±0.9
42.56 ±3.2
10.17 ±3.4
30.00 ±3.0
34.58 ±1.7
Appendix
Table 4: Text accuracy (%); Qwen3-1.7B(student) / Qwen3-4B-Instruct-2507 (teacher). Mean ± std over 3 seeds. Avg.: macro accuracy over 11 ID+OOD datasets. †: OOD. Bold: best student.
Figure 10: Qwen3-1.7B training dynamics: (a) SCAD and ATOD accuracy; (b) SCAD reverse KL on local execution tokens.
Figure 11: Visit-count distributions with logarithmic count axes. (a) Np for root and non-root parents. (b) npa for all observed actions and the subset belonging to multi-candidate parents.
Planning-credit setting
Added tree term
Mean ∣A∣
Avg. (%)
No tree credit
0
0.729
44.32 ±1.5
Shrinkage only
Cp
1.082
45.02 ±1.1
Sibling only: V(p)=Qˉp
Spa
0.828
45.21 ±1.0
SCAD
Spa+Cp
1.181
46.10 ±1.2
Appendix
Table 5: Credit decomposition. Mean ∣A∣ : paired replay on 1,146 decisions from 37 complete mixed-outcome question groups. Accuracy (%): 11-dataset macro mean ± std over 3 seeds.
Subtask τs
Report τr
Avg. (%)
0.75
0.90
45.31 ±1.3
0.83
0.90
46.10 ±1.2
0.90
0.90
45.72 ±1.3
0.83
0.85
45.49 ±1.4
0.83
0.95
45.84 ±1.3
Appendix
Table 6: Threshold and prior-count sensitivity. Accuracy (%): ID+OOD macro mean ± std over 3 seeds; shading: default. Mean ∣A∣ uses the fixed decisions in Table A.2 .