Effective multi-turn agents require interaction strategies that coordinate information gathering, actions, and feedback over long horizons. GRPO is a reinforcement learning algorithm used to train these agents, but sparse trajectory-level rewards limit early exploration in small models. Recent methods augment RL with on-policy distillation (OPD) from a stronger teacher. However, a fixed mixture assumes that teacher guidance and reward optimization should retain a constant relative role throughout training and across interaction turns. This assumption can fail at two scales. Globally, as training progresses, maintaining strong distillation pressure can constrain the model from moving beyond the teacher's capabilities. Locally, teacher--student disagreement identifies where the student departs from the teacher, but cannot tell whether that departure is exploration supported by better outcomes or low-quality policy drift. Our methodological insight is that teacher guidance and reward optimization should be dynamically rebalanced over training and jointly allocated across turns. We instantiate this insight in \tide. Globally, \tide uses the measured disagreement trend as a practical schedule signal, advancing an OPD-to-RL handoff when discrepancy reduction becomes slow but remains positive and progressively increasing the relative weight of RL. Locally, \tide jointly modulates teacher-guided and reward-driven updates: relative action value and disagreement prioritize the OPD signal, whereas relative action value supplies the RL advantage and normalized disagreement reweights it across turns. Coupled with the global handoff, \tide allocates stronger teacher guidance early and gives reward-driven updates greater relative weight later in training. Experiments across multiple benchmarks, student scales, and controlled ablations support the effectiveness of TIDE's adaptive OPD--RL coordination.
Figures & tables
Group
Trajectory success (%)
GPT-5.5 process quality
Teacher–student disagreement
Action Distribution (%)
Search Items
Open Item
Select Option
Go Back
Buy Item
All
44.9
2.0
0.131
18.5
24.9
39.0
10.6
6.8
Glow
56.7
2.2
0.069
55.7
6.5
26.7
1.9
9.2
Ghigh
36.8
1.9
0.189
3.3
37.6
43.2
11.4
4.5
Qlow
0.0
1.0
0.134
19.9
26.5
36.0
15.0
2.3
Qhigh
100.0
3.0
0.123
17.2
21.0
44.3
5.3
12.2
Table 1: Turn-level diagnostics of WebShop trajectories generated with OPD, grouped by teacher–student disagreement and GPT-5.5 process quality.
Method
Type
ALFWorld
WebShop
Pick
Look
Clean
Heat
Cool
Pick2
Avg.
Score
SR
Qwen2.5-1.5B-Instruct
Vanilla †
Prompt
11.1
0.0
6.2
0.0
0.0
4.2
5.5
17.8
5.5
GRPO
RL
80.0 ±2.9
53.8 ±7.7
50.6 ±4.3
37.5 ±6.3
68.0 ±4.0
44.4 ±4.8
58.8 ±0.8
82.4 ±1.2
62.8 ±2.0
GiGPO †
RL
94.4
67.5
94.8
94.4
79.8
76.4
86.7
83.1
65.0
OPD
Distill
88.6 ±5.0
66.7 ±4.4
84.0 ±4.0
72.9 ±5.1
61.3 ±4.7
56.9 ±5.5
73.6 ±2.1
79.4 ±2.3
67.8 ±3.1
Table 2: Main results across model scales on ALFWorld and WebShop (%). Prior-work values are marked † . Bold indicates the highest displayed value in each column at each scale.
Method
WebShop
ALFWorld
SR ↑
Score ↑
Seen ↑
Unseen ↑
GRPO
62.8 ±2.0
82.4 ±1.2
58.8 ±0.8
47.8 ±2.6
Success Handoff
71.4 ±2.5
82.2 ±1.6
73.6 ±2.5
73.1 ±2.7
Fixed 1:1 Mixture
72.0 ±1.7
84.4 ±0.8
77.1 ±1.4
79.9 ±1.5
Linear Handoff
72.2 ±2.2
86.3 ±1.3
85.7 ±1.9
81.3 ±2.0
Cosine Handoff
73.2 ±1.4
87.3 ±0.7
85.0 ±1.4
81.3 ±1.5
Table 3: TIDE ablations with Qwen2.5-1.5B-Instruct students (%). Bold indicates the highest displayed value.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Type
NQ
TriviaQA
PopQA
HotpotQA
2Wiki
MuSiQue
Bamboogle
Avg.
GRPO
RL
14.8 ±1.1
28.3 ±1.6
21.1 ±0.7
15.8 ±1.4
23.9 ±1.2
2.7 ±0.1
9.3 ±0.8
16.6 ±0.3
OPD
Distill
36.0 ±0.8
49.1 ±1.4
38.8 ±0.6
36.8 ±1.1
36.2 ±0.9
26.1 ±1.7
31.2 ±1.0
36.3 ±0.5
TIDE
Hybrid
40.8 ±1.2
51.5 ±0.9
44.3 ±0.5
42.4 ±1.4
38.3 ±1.0
27.7 ±1.6
30.6 ±0.8
39.4 ±0.3
Appendix
Table 4: SearchQA exact-match accuracy with Qwen2.5-1.5B-Instruct students (%). Bold indicates the highest displayed student result.
ζ
WebShop
ALFWorld
SR ↑
Score ↑
Seen ↑
Unseen ↑
0.01
76.0 ±1.9
89.1 ±1.2
85.7 ±1.5
84.3 ±1.7
0.02
77.2 ±1.8
89.8 ±1.2
86.0 ±0.8
85.1 ±1.5
0.05
75.6 ±2.0
88.4 ±1.3
85.0 ±1.6
82.8 ±1.9
Appendix
Table 5: Sensitivity of global handoff parameters (%).
Method
WebShop
ALFWorld
SR ↑
Score ↑
Seen ↑
Unseen ↑
GRPO
63.3 ±2.0
79.8 ±1.1
74.0 ±0.8
60.4 ±2.2
Success Handoff
74.2 ±2.3
84.5 ±1.4
79.3 ±1.9
77.6 ±2.2
Fixed 1:1 Mixture
74.9 ±1.8
85.8 ±1.0
82.1 ±1.2
80.3 ±1.7
Linear Handoff
75.2 ±2.1
87.5 ±1.2
87.1 ±1.4
84.6 ±1.7
Cosine Handoff
76.0 ±1.6
88.4 ±0.9
86.4 ±1.4
85.1 ±1.5
Appendix
Table 6: TIDE ablations with Qwen2.5-3B-Instruct students (%). Bold indicates the highest displayed value.