cs.LGSep 28, 2026

ROSS: Relearning from Self-Generated Rollouts through Selective Supervision

Authors: Zhiwei Zhang, Huayu Deng, Fei Zhao, Jiayan Fu, Bin Liang, Kam-Fai Wong, Mu Chuan

Organizations: AllSpark Team

Abstract

Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances. Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expresses reliably. However, they may also contain mistakes, abandoned attempts, and redundant actions that should not be imitated, motivating finer-grained selective supervision. We introduce ROSS (Relearning from Self-Generated Rollouts through Selective Supervision), which preserves the full historical trajectory as context while applying loss only to selected model-generated continuations. Across domain-specific reinforcement learning, multi-teacher on-policy distillation, and agentic reinforcement learning, ROSS consistently improves upstream checkpoints and outperforms baselines across mathematics, code generation, instruction following, and software engineering. On Qwen3.6-35B-A3B, ROSS improves the six-benchmark MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%. These results show that self-rollout training leaves behind reusable behavioral experience that can yield further gains through offline supervised fine-tuning (SFT), without additional policy rollouts.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. From Rollouts to Recipes: Self-Contained Post-Training for LLMs

    Sep 1, 2026Yifei Li, Lingling Zhang, Muye Huang +3Post-TrainingReinforcement Learning Post-Training

  2. OPSRD: On-Policy Self-Role Distillation

    Sep 30, 2026Weijie Ren, Yanwen Zhang, Hao Li +3Unsupervised On-Policy Self-DistillationInstruction-Tuned Models

  3. On-Policy Self-Distillation without Any Supervision

    Aug 6, 2026Yijiang Li, Bingyang Wang, Yijun Liang +3Unsupervised On-Policy Self-DistillationSelf-Distillation Framework