ROSS: Relearning from Self-Generated Rollouts through Selective Supervision
Organizations: AllSpark Team
Abstract
Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances. Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expresses reliably. However, they may also contain mistakes, abandoned attempts, and redundant actions that should not be imitated, motivating finer-grained selective supervision. We introduce ROSS (Relearning from Self-Generated Rollouts through Selective Supervision), which preserves the full historical trajectory as context while applying loss only to selected model-generated continuations. Across domain-specific reinforcement learning, multi-teacher on-policy distillation, and agentic reinforcement learning, ROSS consistently improves upstream checkpoints and outperforms baselines across mathematics, code generation, instruction following, and software engineering. On Qwen3.6-35B-A3B, ROSS improves the six-benchmark MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%. These results show that self-rollout training leaves behind reusable behavioral experience that can yield further gains through offline supervised fine-tuning (SFT), without additional policy rollouts.
Figures & tables
| Math RL | Code RL | ||||||
| Method | AIME 25 | AIME 26 | HMMT-Nov. | Math Avg. | LCB Gen | OJBench | Code Avg. |
| Base | 71.25 | 76.56 | 70.62 | 72.81 | 57.24 | 22.20 | 39.72 |
| Upstream | 74.84 | 77.50 | 73.23 | 75.19 | 58.95 | 25.22 | 42.09 |
| Continued RL | 74.64 | 78.85 | 72.81 | 75.43 | 57.24 | 27.37 | 42.31 |
| Positive-Rollout SFT | 75.47 | 79.06 | 73.44 | 75.99 | 58.19 | 23.49 | 40.84 |
| ROSS (ours) | 76.46 | 79.90 | 74.58 | 76.98 | 61.05 | 27.37 | 44.21 |
| Method | SWE-bench Verified |
| Base | 60.80 |
| Upstream | 64.20 |
| Positive-Rollout SFT | 65.20 |
| ROSS (ours) | 68.40 |
| Math RL | Code RL | ||||
| Method | AIME 25 | AIME 26 | HMMT-Nov. | LCB Gen | OJBench |
| Upstream | 74.84 | 77.50 | 73.23 | 58.95 | 25.22 |
| Positive-Rollout SFT | 75.47 | 79.06 | 73.44 | 58.19 | 23.49 |
| ROSS w/o mask | 76.67 | 79.06 | 73.23 | 59.14 | 26.94 |
| ROSS (ours) | 76.46 | 79.90 | 74.58 | 61.05 | 27.37 |
| Math RL | Code RL | MOPD | |||
| Math | Code | Math | Code | IF | |
| Base | 72.81 | 57.24 | 72.81 | 57.24 | 34.10 |
| Upstream | 75.19 | 58.95 | 73.89 | 57.43 | 44.58 |
| After ROSS on identical trajectories and masks | |||||
| Initialized from Base | 77.09 | 61.33 | 77.95 | 61.24 | 47.18 |
| Initialized from Upstream | 76.98 | 61.05 | 78.32 | 60.76 | 47.76 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Initialization / rollout window | Positive-Rollout SFT examples | ROSS examples |
| Math RL | Iteration 47 / steps 1–48 | 35,700 | 27,848 |
| Code RL | Iteration 187 / steps 1–187 | 152,454 | 144,702 |
| MOPD | Step 199 / first 200 steps | 42,558 | 31,228 |
| Parameter | Single-turn rollouts (no-thinking) | Agentic trajectories (thinking) |
| Optimizer | Adam, , | |
| Peak / minimum learning rate | / | |
| Schedule / warmup fraction | Cosine / 0.05 | |
| Weight decay | 0.1 | |
| Global batch size | 128 | |
| Maximum sequence length (tokens) | 32,768 | 65,536 |
| Setting | Checkpoint | AIME 25 | AIME 26 | HMMT- Nov. | Avg. Math | LCB Gen | OJBench | IFBench | BFCL Overall | BFCL Base | Miss. Func. | Miss. Param. | Long Context |
| Math RL | Upstream (s47) | – | – | – | – | 56.76 | 25.65 | 33.30 | 44.12 | 57.50 | 29.50 | 39.00 | 50.50 |
| ROSS | – | – | – | – | 58.95 | 29.53 | 34.60 | 45.62 | 58.50 | 32.50 | 40.50 | 51.00 | |
| – | – | – | – | ||||||||||
| Code RL | Upstream (s187) | 75.21 | 76.93 | 70.42 | 74.19 | – | – | 35.45 | 46.25 | 60.50 | 32.50 | 40.50 | 51.50 |
| ROSS | 73.44 | 78.54 | 73.12 | 75.03 | – | – | 35.07 | 49.38 | 59.50 | 43.50 | 45.00 | 49.50 | |
| – | – |
| Metric | Upstream | ROSS (from Upstream) | ROSS (from Base) |
| Mean output tokens | 13,895 | 9,778 | 9,606 |
| Median output tokens | 16,384 | 9,820 | 9,850 |
| Cap-hit rate (%) | 68–74 | 15–17 | |
| Mean boxed-answer count | 628.3 | 63.0 | 91.6 |
| Repetition rate (%) | 84.5 | 21.5 | 24.5 |
| Method | Budget | AIME 25 | AIME 26 | HMMT-Nov. | Avg. Math |
| Upstream | 16K | 45.42 | 48.54 | 62.19 | 52.05 |
| Upstream | 32K | 44.17 | 48.49 | 62.19 | 51.62 |
| ROSS (from Upstream) | 16K | 60.94 | 68.33 | 61.67 | 63.65 |
| ROSS (from Base) | 16K | 62.24 | 71.30 | 63.23 | 65.59 |