Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances. Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expresses reliably. However, they may also contain mistakes, abandoned attempts, and redundant actions that should not be imitated, motivating finer-grained selective supervision. We introduce ROSS (Relearning from Self-Generated Rollouts through Selective Supervision), which preserves the full historical trajectory as context while applying loss only to selected model-generated continuations. Across domain-specific reinforcement learning, multi-teacher on-policy distillation, and agentic reinforcement learning, ROSS consistently improves upstream checkpoints and outperforms baselines across mathematics, code generation, instruction following, and software engineering. On Qwen3.6-35B-A3B, ROSS improves the six-benchmark MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%. These results show that self-rollout training leaves behind reusable behavioral experience that can yield further gains through offline supervised fine-tuning (SFT), without additional policy rollouts.
Figures & tables
Figure 1 : Relearning from self-generated historical rollouts. (a) As the policy evolves, some useful behaviors may become underrepresented in its current rollout distribution. (b) Even when a rollout reaches a successful outcome, it may include redundant reasoning, or unnecessary detours, making full-trajectory replay suboptimal.
Figure 2 : When is historical experience worth revisiting? Its value depends on compatibility with the current policy and complementarity to its behavioral coverage.
Figure 3 : Historical rollouts contain compatible but under-consolidated behaviors. (a) Historical rollouts have response-level NLL comparable to the current policy and lower than external-teacher trajectories. (b) These historically successful behaviors remain unreliable under the current policy, with mean pass@1=37.7% . (c) Repeated sampling raises success to pass@32=99.0% , showing that the behaviors remain available but are not reliably expressed.
Figure 4 : Overview of ROSS. Starting from the final checkpoint, ROSS filters saved rollouts by outcome verification, reviews the positive trajectories, and applies imitation loss only to selected model-generated tokens while preserving the full history as context.
Math RL
Code RL
Method
AIME 25
AIME 26
HMMT-Nov.
Math Avg.
LCB Gen
OJBench
Code Avg.
Base
71.25
76.56
70.62
72.81
57.24
22.20
39.72
Upstream
74.84
77.50
73.23
75.19
58.95
25.22
42.09
Continued RL
74.64
78.85
72.81
75.43
57.24
27.37
42.31
Positive-Rollout SFT
75.47
79.06
73.44
75.99
58.19
23.49
40.84
ROSS (ours)
76.46
79.90
74.58
76.98
61.05
27.37
44.21
Table 1 : Main results with Qwen3.6-35B-A3B. All scores are higher-is-better. (a) Math and Code use independent domain-specific RL runs; Math Avg. and Code Avg. are the corresponding domain means. (b) MOPD spans all three domains, and Avg. is the mean of its six benchmarks. All relearning methods start from the corresponding Upstream checkpoint; bold marks the best result among them, including ties.
Method
SWE-bench Verified ↑(Δ)
Base
60.80
Upstream
64.20
Positive-Rollout SFT
65.20 (+1.00)
ROSS (ours)
68.40 (+4.20)
Table 2 : Agentic trajectory reuse. Δ is relative to Upstream.
Math RL
Code RL
Method
AIME 25
AIME 26
HMMT-Nov.
LCB Gen
OJBench
Upstream
74.84
77.50
73.23
58.95
25.22
Positive-Rollout SFT
75.47 +0.63
79.06 +1.56
73.44 +0.21
58.19 −0.76
23.49 −1.73
ROSS w/o mask
76.67 +1.83
79.06 +1.56
73.23 +0.00
59.14 +0.19
26.94 +1.72
ROSS (ours)
76.46 +1.62
79.90 +2.40
74.58 +1.35
61.05 +2.10
27.37 +2.15
Table 3 : Effect of trajectory filtering and token-level masking. Subscripts show percentage-point changes from Upstream; green/red denote gains/losses and bold marks the best score. ROSS w/o mask and ROSS use identical examples and differ only by token-level masking.
Math RL
Code RL
MOPD
Math ↑
Code ↑
Math ↑
Code ↑
IF ↑
Base
72.81
57.24
72.81
57.24
34.10
Upstream
75.19
58.95
73.89
57.43
44.58
After ROSS on identical trajectories and masks
Initialized from Base
77.09
61.33
77.95
61.24
47.18
Initialized from Upstream
76.98 −0.11
61.05 −0.28
78.32 +0.37
60.76 −0.48
47.76 +0.58
Table 4 : ROSS reaches similar scores from different initializations. Within each setting, the two runs use identical historical trajectories, masks, and SFT configurations. Math is Avg. Math, Code is LCB Gen, and IF is IFBench.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Initialization / rollout window
Positive-Rollout SFT examples
ROSS examples
Math RL
Iteration 47 / steps 1–48
35,700
27,848
Code RL
Iteration 187 / steps 1–187
152,454
144,702
MOPD
Step 199 / first 200 steps
42,558
31,228
Appendix
Table 6 : Historical data used for relearning. ROSS w/o mask and ROSS share the retained examples; only the token-level loss mask differs. Positive-Rollout SFT uses all verifier-positive examples in the corresponding window.
Parameter
Single-turn rollouts (no-thinking)
Agentic trajectories (thinking)
Optimizer
Adam, β1=0.9 , β2=0.98
Peak / minimum learning rate
4×10−6 / 2×10−7
Schedule / warmup fraction
Cosine / 0.05
Weight decay
0.1
Global batch size
128
Maximum sequence length (tokens)
32,768
65,536
Appendix
Table 7 : 35B relearning configurations. Single-turn Math, Code, and MOPD rollouts use no-thinking responses; agentic rollouts retain thinking traces and use a longer context and separate schedule. Within each setting, ROSS and its matched full-supervision control share the listed configuration.
Table 8 : A segment-level complementarity case. The earlier historical rollout recovers from a shared overcounting error, whereas the later rollout retains it.
Figure 5 : Full distribution of masked-content categories. Each column gives the percentage of sampled trajectories assigned to each primary category. Math early/late have n=300/300 ; Code early/late have n=290/295 . Early and late denote rollout steps 1–20 and 81–100. Percentages sum to approximately 100 within each column, subject to rounding.
Setting
Checkpoint
AIME 25
AIME 26
HMMT- Nov.
Avg. Math
LCB Gen
OJBench
IFBench
BFCL Overall
BFCL Base
Miss. Func.
Miss. Param.
Long Context
Math RL
Upstream (s47)
–
–
–
–
56.76
25.65
33.30
44.12
57.50
29.50
39.00
50.50
ROSS
–
–
–
–
58.95
29.53
34.60
45.62
58.50
32.50
40.50
51.00
Δ
–
–
–
–
+2.19
+3.88
+1.30
+1.50
+1.00
+3.00
+1.50
+0.50
Code RL
Upstream (s187)
75.21
76.93
70.42
74.19
–
–
35.45
46.25
60.50
32.50
40.50
51.50
ROSS
73.44
78.54
73.12
75.03
–
–
35.07
49.38
59.50
43.50
45.00
49.50
Δ
−1.77
+1.61
+2.70
+0.84
–
–
−0.38
+3.13
−1.00
+11.00
+4.50
−2.00
Appendix
Table 9 : Cross-domain retention and transfer after domain-specific ROSS. Scores are percentages, and Δ is ROSS minus Upstream. The first seven columns use the corresponding third-epoch ROSS models; the BFCL columns use Math i650 and Code i3389. Avg. Math averages AIME 2025, AIME 2026, and HMMT-November 2025. BFCL Base denotes the standard multi-turn subset.
Metric
Upstream
ROSS (from Upstream)
ROSS (from Base)
Mean output tokens ↓
13,895
9,778
9,606
Median output tokens
16,384
9,820
9,850
Cap-hit rate (%) ↓
68–74
≈15
15–17
Mean boxed-answer count ↓
628.3
63.0
91.6
Repetition rate (%) ↓
84.5
21.5
24.5
Appendix
Table 10 : Reduced output degeneration after ROSS on 4B Math RL. Measurements use the 16K output budget. Repetition denotes at least five boxed-answer markers per response.
Method
Budget
AIME 25
AIME 26
HMMT-Nov.
Avg. Math
Upstream
16K
45.42
48.54
62.19
52.05
Upstream
32K
44.17
48.49
62.19
51.62
ROSS (from Upstream)
16K
60.94
68.33
61.67
63.65
ROSS (from Base)
16K
62.24
71.30
63.23
65.59
Appendix
Table 11 : Mathematical accuracy and the output-budget control. All scores are percentages. Only Upstream is reevaluated with a 32K limit; both ROSS variants use 16K. Avg. Math is the unweighted mean of the three displayed benchmarks.
School of Computer Science and Technology, Xi’an Jiaotong University, China · Shaanxi Province Key Laboratory of Big Data Knowledge Engineering · Zhongguancun Academy, Beijing, China +1