ReSAIL: Mitigating Collapse in Iterative Agent Self-Distillation
Organizations: Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China
Abstract
Iterative self-distillation enables LLM agents to learn from successive deployments, offering a path toward recursive self-improvement (RSI). Yet our experiments with existing methods reveal a collapse in deployment performance across cycles, while task performance with privileged information (PI) also declines. We address this collapse by prioritizing informative interaction steps for distillation and preserving PI-conditioned behavior as the student becomes the next teacher. We introduce Retentive and Selective Augmentation for Iterative Self-Distillation (ReSAIL), a plug-in augmentation for iterative PI-based self-distillation. ReSAIL selects interaction steps where PI most strongly changes the teacher's predictions and balances the resulting distillation losses across trajectories. It also regularizes the student's PI-conditioned output distributions toward those of the frozen teacher at selected and unselected steps to preserve PI-conditioned behavior for supervision in the next cycle. On ALFWorld and TextCraft, ReSAIL sustains substantial gains across model scales over three cycles, with an average absolute gain of 22.5% in final-cycle success rates when added to self-distillation baselines. Sensitivity-guided selection of offline data also improves action prediction accuracy for multimodal GUI agents on AITZ. These findings provide the first evidence that a more robust learning mechanism can effectively mitigate performance collapse in iterative agent self-distillation over deployment trajectories.
Figures & tables
| Model | Method | ALFWorld (ID) | ALFWorld (OOD) | TextCraft | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Cycle 1 | Cycle 2 | Cycle 3 | Cycle 1 | Cycle 2 | Cycle 3 | Cycle 1 | Cycle 2 | Cycle 3 | ||
| Qwen3-4B | ReAct | (2.3) | – | – | (1.8) | – | – | (1.5) | – | – |
| RFT | (2.7) | (4.4) | (6.6) | (4.6) | (1.4) | (3.9) | (0.6) | (5.5) | (2.5) | |
| GRPO | (2.1) | (1.2) | (2.3) | (1.2) | (1.6) | (2.0) | (1.2) | (2.3) | (2.3) | |
| EPD | (1.2) | (3.4) | (0.5) | (2.7) | (3.2) | (3.5) | (0.6) | (2.6) | (3.5) | |
| SDPO | (1.6) | (2.0) | (1.6) | (1.6) | (1.6) | (0.5) | (2.3) | (0.6) | (4.5) | |
| Components | ALFWorld (ID) | ALFWorld (OOD) | ||||||
|---|---|---|---|---|---|---|---|---|
| SGS | TLB | PR | Cycle 1 | Cycle 2 | Cycle 3 | Cycle 1 | Cycle 2 | Cycle 3 |
| {\color[rgb]{0.6,0.6,0.6}\times} | {\color[rgb]{0.6,0.6,0.6}\times} | {\color[rgb]{0.6,0.6,0.6}\times} | (4.7) | (2.0) | (1.2) | (2.5) | (2.1) | (1.2) |
| {\color[rgb]{0.6,0.6,0.6}\times} | {\color[rgb]{0.6,0.6,0.6}\times} | (4.8) | (4.6) | (3.2) | (5.9) | (2.5) | (2.5) | |
| {\color[rgb]{0.6,0.6,0.6}\times} | {\color[rgb]{0.6,0.6,0.6}\times} | (1.6) | (2.1) | (2.4) | (1.6) | (2.7) | (3.1) | |
| {\color[rgb]{0.6,0.6,0.6}\times} | {\color[rgb]{0.6,0.6,0.6}\times} | (4.4) | (0.9) | (1.6) | (2.0) | (1.6) | (2.1) | |
| {\color[rgb]{0.6,0.6,0.6}\times} | (3.0) | (0.5) | (3.0) | (3.0) | (1.6) | (1.6) | ||
| Selection | ID | OOD |
|---|---|---|
| Random | (3.9) | (4.4) |
| Bottom | (3.2) | (4.3) |
| Top (ours) | (2.0) | (4.0) |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Benchmark | Test set | Metrics |
|---|---|---|
| ALFWorld | 128 ID and 128 OOD episodes | Success rate |
| TextCraft | Full official test set: 100 episodes | Success rate |
| AITZ | Fixed monitoring subset: 101 episodes, 843 non- stop steps | Accuracy |
| Benchmark | Updates | Batch size | Trajectories | ||
|---|---|---|---|---|---|
| ALFWorld | 30 | 32 | 960 | 0.05 | 0.5 |
| TextCraft | 30 | 8 | 240 | 0.25 | 1.0 |
| AITZ (OEL) | 100 | 16 | 1,600 | – | – |
| AITZ (OEL + SGS) | 100 | 16 | 1,600 | 0.80 | – |
| Method | Response source | Training objective |
|---|---|---|
| ReAct | Initial policy | No update |
| RFT | Successful logged responses | Cross-entropy |
| GRPO | Policy-generated candidate groups | Clipped objective + KL |
| EPD | Cached privileged-teacher responses | Cross-entropy |
| Offline SDPO | Cached cycle-start student responses | Distillation with correction |
| OEL | Fresh current-student responses | Distillation |
| Top- | ID | OOD |
|---|---|---|
| 1% | (2.1) | (3.2) |
| 5% | (2.0) | (4.0) |
| 25% | (4.3) | (5.5) |
| 50% | (1.8) | (2.8) |
| 100% | (3.0) | (2.4) |
| Configuration | Response | Scoring | Optimization | Other | Total |
|---|---|---|---|---|---|
| OEL | 1.53 | 0.33 | 0.91 | 1.56 | 4.33 |
| OEL + SGS | 1.97 | 0.55 | 0.24 | 1.32 | 4.08 |
| OEL + SGS + TLB | 1.99 | 0.56 | 0.24 | 1.39 | 4.17 |
| ReSAIL (+ PR) | 1.73 | 0.72 | 1.13 | 1.73 | 5.32 |
| Retention | Cycle | Ordinary view | Privileged view |
|---|---|---|---|
| Shared initial model | 0 | 39.6 (1.8) | 49.2 (2.1) |
| None | 1 | 60.4 (3.0) | 61.5 (3.3) |
| 2 | 58.9 (1.6) | 52.6 (0.9) | |
| 3 | 45.6 (1.6) | 44.8 (3.9) | |
| Ordinary | 1 | 52.1 (1.6) | 54.7 (1.6) |
| 2 | 66.7 (0.9) | 61.5 (1.8) |