cs.LGSep 30, 2026

ReSAIL: Mitigating Collapse in Iterative Agent Self-Distillation

Authors: Shengjie Jin, Hengbo Xu, Zelong Sun, YuJie Guo, Zhiwu Lu

Organizations: Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China

Abstract

Iterative self-distillation enables LLM agents to learn from successive deployments, offering a path toward recursive self-improvement (RSI). Yet our experiments with existing methods reveal a collapse in deployment performance across cycles, while task performance with privileged information (PI) also declines. We address this collapse by prioritizing informative interaction steps for distillation and preserving PI-conditioned behavior as the student becomes the next teacher. We introduce Retentive and Selective Augmentation for Iterative Self-Distillation (ReSAIL), a plug-in augmentation for iterative PI-based self-distillation. ReSAIL selects interaction steps where PI most strongly changes the teacher's predictions and balances the resulting distillation losses across trajectories. It also regularizes the student's PI-conditioned output distributions toward those of the frozen teacher at selected and unselected steps to preserve PI-conditioned behavior for supervision in the next cycle. On ALFWorld and TextCraft, ReSAIL sustains substantial gains across model scales over three cycles, with an average absolute gain of 22.5% in final-cycle success rates when added to self-distillation baselines. Sensitivity-guided selection of offline data also improves action prediction accuracy for multimodal GUI agents on AITZ. These findings provide the first evidence that a more robust learning mechanism can effectively mitigate performance collapse in iterative agent self-distillation over deployment trajectories.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Why Does Feedback-Augmented Self-Distillation Fail to Improve Retrieval-Interleaved Search Agents?

    Jul 20, 2026Fan Yang, Rui Meng, Yuxin WenSelf-Distillation FrameworkSelf-Teacher

  2. From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents

    Jul 30, 2026Xu Xia, Jinghua Piao, Min Yang +3Unsupervised On-Policy Self-DistillationSelf-Distillation Framework