Iterative self-distillation enables LLM agents to learn from successive deployments, offering a path toward recursive self-improvement (RSI). Yet our experiments with existing methods reveal a collapse in deployment performance across cycles, while task performance with privileged information (PI) also declines. We address this collapse by prioritizing informative interaction steps for distillation and preserving PI-conditioned behavior as the student becomes the next teacher. We introduce Retentive and Selective Augmentation for Iterative Self-Distillation (ReSAIL), a plug-in augmentation for iterative PI-based self-distillation. ReSAIL selects interaction steps where PI most strongly changes the teacher's predictions and balances the resulting distillation losses across trajectories. It also regularizes the student's PI-conditioned output distributions toward those of the frozen teacher at selected and unselected steps to preserve PI-conditioned behavior for supervision in the next cycle. On ALFWorld and TextCraft, ReSAIL sustains substantial gains across model scales over three cycles, with an average absolute gain of 22.5% in final-cycle success rates when added to self-distillation baselines. Sensitivity-guided selection of offline data also improves action prediction accuracy for multimodal GUI agents on AITZ. These findings provide the first evidence that a more robust learning mechanism can effectively mitigate performance collapse in iterative agent self-distillation over deployment trajectories.
Figures & tables
Figure 1: Iterative self-distillation does not guarantee sustained improvement. On ALFWorld OOD with Qwen3-4B: (a) deployment success rate of SDPO and OEL with and without ReSAIL; (b) success rate of the original methods conditioned on privileged information (PI), with the same tasks and their PI held fixed across cycles. Cycle 0 denotes the shared initial model. Error bars show the standard deviation of success rates across three decoding seeds.
Figure 2: Overview of ReSAIL . (a) Base self-distillation matches the ordinary-view student to the privileged-view teacher, uniformly averaging losses over all interaction steps used for training. (b) ReSAIL uses Sensitivity-Guided Selection (SGS) to select the most PI-sensitive steps and Trajectory Loss Balancing (TLB) to balance the selected losses across trajectories. Privileged Retention (PR) regularizes the student’s privileged-view predictions toward those of the teacher at both selected and unselected interaction steps within each training batch.
Model
Method
ALFWorld (ID)
ALFWorld (OOD)
TextCraft
Cycle 1
Cycle 2
Cycle 3
Cycle 1
Cycle 2
Cycle 3
Cycle 1
Cycle 2
Cycle 3
Qwen3-4B
ReAct
40.1 (2.3)
–
–
39.6 (1.8)
–
–
61.3 (1.5)
–
–
RFT
43.2 (2.7)
44.8 (4.4)
41.9 (6.6)
39.3 (4.6)
42.2 (1.4)
43.2 (3.9)
59.7 (0.6)
62.3 (5.5)
60.7 (2.5)
GRPO
39.1 (2.1)
36.5 (1.2)
37.5 (2.3)
45.1 (1.2)
42.2 (1.6)
50.3 (2.0)
60.3 (1.2)
58.7 (2.3)
56.7 (2.3)
EPD
45.6 (1.2)
36.7 (3.4)
22.1 (0.5)
43.2 (2.7)
31.5 (3.2)
18.2 (3.5)
63.3 (0.6)
62.0 (2.6)
62.3 (3.5)
SDPO
52.3 (1.6)
53.4 (2.0)
58.9 (1.6)
44.8 (1.6)
50.5 (1.6)
44.3 (0.5)
62.7 (2.3)
64.7 (0.6)
60.3 (4.5)
Table 1: Main results across three deployment cycles. Task success rates (%) without PI, reported as mean (SD) over three decoding seeds. SD denotes standard deviation. Shaded rows add ReSAIL to the preceding baseline; bold marks the highest mean in each column within each model size. ReAct is the unchanged initial policy, shown once under Cycle 1 as a fixed reference across cycles.
Components
ALFWorld (ID)
ALFWorld (OOD)
SGS
TLB
PR
Cycle 1
Cycle 2
Cycle 3
Cycle 1
Cycle 2
Cycle 3
{\color[rgb]{0.6,0.6,0.6}\times}
{\color[rgb]{0.6,0.6,0.6}\times}
{\color[rgb]{0.6,0.6,0.6}\times}
54.9 (4.7)
56.5 (2.0)
47.4 (1.2)
47.1 (2.5)
46.1 (2.1)
38.5 (1.2)
✓
{\color[rgb]{0.6,0.6,0.6}\times}
{\color[rgb]{0.6,0.6,0.6}\times}
66.4 (4.8)
69.0 (4.6)
50.5 (3.2)
61.5 (5.9)
56.5 (2.5)
37.8 (2.5)
{\color[rgb]{0.6,0.6,0.6}\times}
✓
{\color[rgb]{0.6,0.6,0.6}\times}
56.5 (1.6)
60.2 (2.1)
59.9 (2.4)
51.3 (1.6)
50.8 (2.7)
50.8 (3.1)
{\color[rgb]{0.6,0.6,0.6}\times}
{\color[rgb]{0.6,0.6,0.6}\times}
✓
52.9 (4.4)
55.2 (0.9)
64.3 (1.6)
49.0 (2.0)
49.7 (1.6)
53.9 (2.1)
✓
✓
{\color[rgb]{0.6,0.6,0.6}\times}
65.9 (3.0)
68.5 (0.5)
56.5 (3.0)
60.4 (3.0)
58.9 (1.6)
45.6 (1.6)
Table 2: Component ablations with OEL and Qwen3-4B on ALFWorld. ✓ and × indicate enabled and disabled components; shading marks the full method. Entries report success rates (%) as mean (SD) over three decoding seeds. Bold marks the highest mean in each column. SGS: Sensitivity-Guided Selection; TLB: Trajectory Loss Balancing; PR: Privileged Retention.
Selection
ID
OOD
Random
49.7 (3.9)
45.8 (4.4)
Bottom
40.1 (3.2)
40.4 (4.3)
Top (ours)
62.2 (2.0)
57.3 (4.0)
Table 3: Selection strategy and retention view. OEL with Qwen3-4B on ALFWorld. (a) Cycle 1 selection at ρ=0.05 , fixing TLB and PR. Bottom and Top select the least and most PI-sensitive interaction steps. (b) Cycle 3 retention, fixing SGS and TLB. Entries report success without PI (%) as mean (SD) over three decoding seeds; bold marks the highest mean in each column.
Figure 3: Privileged retention sustains PI-conditioned task competence. OEL+TBSD with Qwen3-4B on ALFWorld OOD tasks, evaluated (a) without PI and (b) with PI. Task initial states and the PI in (b) are fixed across checkpoints and retention variants. C0 denotes the shared initial model. Points show means; error bars show ±1 SD over three decoding seeds.
Figure 4: Selection ratio and retention weight. OEL+ReSAIL with Qwen3-4B on ALFWorld. (a) Cycle 1 success at each selection ratio ρ , fixing λ=0.5 . (b) Cycle 3 success at each retention weight λ , fixing ρ=0.05 ; each weight remains constant across cycles. TLB is enabled in both sweeps. Points show means; horizontal error bars show ±1 SD over three decoding seeds.
Figure 5: Training cost and AITZ accuracy. (a) Approximate Cycle 1 training-loop costs on ALFWorld using eight H800 GPUs. Components are added cumulatively to OEL; the final bar is OEL+ReSAIL. Other* denotes residual training-loop cost. (b) AITZ step-level action accuracy on a fixed subset. OEL+SGS retains a fixed top-80% training subset.
Table 5: Training settings per cycle (one cycle for AITZ, three for text tasks). Batch size counts source trajectories before step filtering; ρ and λ denote selection ratio and retention weight.
Method
Response source
Training objective
ReAct
Initial policy
No update
RFT
Successful logged responses
Cross-entropy
GRPO
Policy-generated candidate groups
Clipped objective + KL
EPD
Cached privileged-teacher responses
Cross-entropy
Offline SDPO
Cached cycle-start student responses
Distillation with correction
OEL
Fresh current-student responses
Distillation
Appendix
Table 6: Baseline response sources and training objectives. All methods use ordinary input at evaluation. SDPO in the result tables refers to the offline adaptation below.
Top- ρ
ID
OOD
1%
60.9 (2.1)
54.2 (3.2)
5%
62.2 (2.0)
57.3 (4.0)
25%
60.4 (4.3)
53.6 (5.5)
50%
56.5 (1.8)
54.7 (2.8)
100%
56.8 (3.0)
50.3 (2.4)
Appendix
Table 7: Complete hyperparameter ablation results. Success rates (%) are reported as mean (sample standard deviation) over three decoding seeds, with 128 episodes per split in each evaluation. Shading marks the default parameter in each sweep. The two panels report different deployment cycles.
Configuration
Response
Scoring
Optimization
Other
Total
OEL
1.53
0.33
0.91
1.56
4.33
OEL + SGS
1.97
0.55
0.24
1.32
4.08
OEL + SGS + TLB
1.99
0.56
0.24
1.39
4.17
ReSAIL (+ PR)
1.73
0.72
1.13
1.73
5.32
Appendix
Table 8: Training-loop cost on ALFWorld (Cycle 1, GPU-hours on eight H800 GPUs). Scoring includes selection; Other is computed as the residual. Values are approximate; totals are computed before rounding.
Retention
Cycle
Ordinary view
Privileged view
Shared initial model
0
39.6 (1.8)
49.2 (2.1)
None
1
60.4 (3.0)
61.5 (3.3)
2
58.9 (1.6)
52.6 (0.9)
3
45.6 (1.6)
44.8 (3.9)
Ordinary
1
52.1 (1.6)
54.7 (1.6)
2
66.7 (0.9)
61.5 (1.8)
Appendix
Table 9: Task success under both input views. Mean success rates (%) and sample standard deviations over three decoding seeds on 128 ALFWorld OOD tasks. All trained configurations use OEL+TBSD and differ only in retention; the shared initial checkpoint is listed once.
Figure 6: Full PI-sensitivity dynamics across three deployment cycles. Current-student JSD on the fixed ALFWorld/Qwen3-4B probe, using response-token averaging. (a) OEL and component combinations. (b) SDPO with and without ReSAIL, plus EPD. All observations are shown without smoothing across 90 updates; vertical lines mark cycle boundaries and markers are spaced for readability.
On-policy self-distillation (OPSD) offers a promising approach for training large language models without relying on a separate teacher model. However, its effectiveness on complex agentic tasks remains largely unexplored. In this work, we instantiate Feedback-Augmented Self-Distillation (FA-SD), a self-distillation algorithm for agentic search that leverages successful demonstrations as privileged information. We identify that models can rely on recurring reasoning-and-search output templates, producing trajectories that appear diverse but are largely agnostic to the input question, making the KL-based self-distillation signal uninformative. We term this phenomenon decoding collapse, a failure mode that can be missed by existing evaluation metrics. To understand its underlying cause, we show that although the self-teacher achieves stronger performance, learning remains inherently unstable due to inconsistent supervision signals. We further decompose this inconsistency into model inconsistency and prompt inconsistency, and show that the latter can significantly degrade the quality of the supervision signal, limiting the effectiveness of self-teacher learning. To mitigate this inconsistency, we introduce an exponential moving average (EMA) teacher to stabilize the self-teacher and provide more consistent supervision signals. Although the EMA teacher requires a warm-up phase during which performance may temporarily regress, it ultimately improves model performance by providing more stable supervision.
Fan Yang, Rui Meng, Yuxin Wen
Chapman University, Orange, CA, USA · Lawrence Berkeley National Laboratory, Berkeley, CA, USA
Recent work on LLM agents is shifting from external capability elicitation to capability internalization, enabling agents to retain useful skills without retrieval at inference time. On-policy self-distillation (OPSD) offers a promising direction, but many existing methods typically supervise students by scoring actions along student-generated trajectories. Such supervision has two limitations: teacher preferences are not validated by environment outcomes, and action-level scores underuse information from student rollouts, teacher rollouts, and their behavioral relationship. We therefore advocate outcome-verified teacher supervision and comparative learning over teacher-student trajectories. Based on this view, we propose Outcome-Verified Comparative Self-Distillation (OVCSD). OVCSD organizes failed student rollouts into a prefix tree, adaptively invokes a skill-conditioned teacher from student-reached states, and retains only outcome-verified successful continuations. It then applies localized comparative learning at the first state-aligned divergence and distills the post-divergence teacher suffix to transfer completion behavior. Experiments on ALFWorld and WebShop across three model scales show that OVCSD consistently outperforms skill-free RL and existing self-distillation baselines, achieving up to 29.7 and 5.4 absolute success-rate gains over the strongest baselines on ALFWorld and WebShop, respectively, while adding less than 3% privileged interaction during training.
Xu Xia, Jinghua Piao, Min Yang +3
Southeast University · Zhongguancun Academy · Tsinghua University +2
On-policy self-distillation (OPSD) has become a popular recipe for post-training LLM agents. It supervises the agent model at the token level with a stronger teacher view of the same model, obtained by conditioning on privileged information (PI). In this work, we show that in multi-turn agents, this paradigm teaches the student to act with confidence but without the information behind it. The trained agent behaves as if it had privileged information it never observed, and its performance falls well short of plain RL, in the worst case below the untrained base model. Therefore, we propose Privileged Self-Practice (PSP), which keeps the PI and moves it from the loss to the sampler. When the student's rollouts on a task mostly fail, we inject a short per-task instruction written by an analyzer model, sample the task again with the instruction in context, and train on the result with an unchanged GRPO objective. The privileged information stays in the prompt and never enters the loss. Across AppWorld and SWE-bench Verified, with three different student models, PSP obtains the best average score in every setting and is the only method that consistently outperforms plain GRPO, improving task-goal completion by up to 65% on AppWorld and the resolved rate by up to 61% on SWE-bench Verified.