Production RL for language models lets the sampler fall behind the learner and repairs the resulting mismatch with a truncated importance weight. We ask how long the sampler can go without a refresh under that correction, and find a cliff: on Qwen2.5-Math-1.5B and GSM8K, importance-corrected GRPO refreshed every 192 updates learns well for 180 steps and then degrades severely in all three data seeds before the refresh arrives. Published remedies for staleness act on the update; we act on the sampler instead. Decoupled cooling draws samples at temperature 0.8 while the learner, the reference model and the importance weights stay at temperature 1, with the behaviour probability recorded from the tempered distribution, so the learner's objective is unchanged. All corresponding cooled runs are stable, and the longer interval keeps what the short one delivered: at the same update budget, a cooled sampler refreshed every 192 steps matches an uncooled sampler refreshed every 96 at the end of training (0.857 for both) and averaged over it (0.79), whereas lowering the learning rate to a safe value ends 3-7 points lower. On Qwen2.5-Math-7B the degradation points at interval 192 predict that an interval of 144 is fatal without cooling and survivable with it; on two data seeds the uncooled runs degrade before their first refresh and the cooled runs pass it and end at 92-93% against 68-81%, with one cooled run degrading transiently late in the second cycle. The benefit has a window: at three times the safe interval and in a high-mismatch MATH setting cooling delays degradation without preventing it, stronger cooling is not better, and cooling without the correction collapses. Sampling temperature is a control on staleness tolerance, and temperature and refresh interval should be chosen together.
Figures & tables
Figure 1: Importance-corrected GRPO has a finite safe staleness; cooling the sampler extends it. Qwen2.5-Math-1.5B on GSM8K, sampler refreshed every 192 steps (grey line), learner at temperature 1, token-level Tis with cap 2 in all runs. (a) Validation accuracy (greedy, 1319 problems, every 10 steps). With Tb=1.0 all three data seeds collapse at step 185–190 and oscillate for the rest of the run; with Tb=0.8 all runs are stable (three data seeds with the learner at temperature 1, plus one run with the learner tied to 0.8); Tb=0.9 dips once at step 190 and recovers. (b) Per-step KL between learner and sampler policies for the seed-1 runs of each temperature (log scale, dashed line at 1 nat per token). The KL accumulates within a refresh cycle, resets at the refresh, and grows faster the hotter the sampler; only the Tb=1 run crosses one nat before the refresh.
Figure 2: Cooling lowers the learner–sampler mismatch in every setting. (a) KL at a fixed step of the first cycle (110, or 60 where the refresh is at 96), uncooled ( Tb=1 , orange) against cooled ( 0.8 , blue), one row per model, task and refresh interval; large dots are seed means, small dots single runs, the label is their ratio. (b) First step at which the KL of the same runs exceeds 0.1 nat per token; the grey tick marks the first refresh, and hollow markers did not reach 0.1 nat before it.
Figure 3: Qwen2.5-Math-7B on GSM8K. (a) Refresh every 144 steps, chosen before the runs from the two runs in (b); solid lines data seed 1, dashed seed 43. The uncooled runs degrade severely before the first refresh on both seeds; the cooled runs have no degradation before it, keep improving after it, and end at 0.92–0.93 (all realizations in Appendix B ). (b) Refresh every 192 steps: the uncooled run collapses to zero; the cooled run degrades 60 steps later and recovers after the refresh.
Figure 4: Lowering the learning rate or an update-side stabiliser ( μ -GRPO) also avoids the cliff; cooling learns more, and does not stack with μ -GRPO. 1.5B, GSM8K, refresh every 192, data seeds 43 and 44 for every arm. (a) Validation accuracy: Tis at Tb=1 with learning rate 2×10−6 collapses; at 1.5×10−6 and 10−6 it is stable but ends at 0.82–0.83 and 0.79–0.80; the cooled sampler at 2×10−6 is stable and ends at 0.86. μ -GRPO (green; solid Tb=1 , dashed 0.8 ) is stable in both seeds (last-five mean 0.83 and 0.85 alone, 0.82 with the cooled sampler). (b) Learner–sampler KL of the same runs.
Objective
Tb
N
Refreshes
Degraded
Last five: 43 / 44 / mean
Whole run
Tis
1.0
96
3
0/2
0.854
0.860
0.857
0.790
Tis
0.8
96
3
0/2
0.860
0.862
0.861
0.800
Tis
1.0
192
1
2/2
0.679
0.532
0.605
0.654
Tis
0.8
192
1
0/2
0.863
0.852
0.857
0.787
Tis , lr 1.5×10−6
1.0
192
1
0/2
0.818
0.830
0.824
0.761
Tis , lr 10−6
1.0
192
1
0/2
0.790
0.789
0.789
0.734
Table 1: Same budget, same seeds. 1.5B, GSM8K, 300 updates, learning rate 2×10−6 unless stated, data seeds 43 and 44: severe-degradation count, the end metric (last-five mean, steps 260–300) per seed and averaged, and the whole-run metric (trapezoid mean of the 31 validations, averaged over seeds). Cooling doubles the refresh interval at the end and whole-run quality of the short one; the alternatives avoid the cliff at a cost on both. Updates and prompt budget are identical across rows; fewer refreshes pay off where the refresh is the bottleneck.
Figure 5: Boundaries. (a) 1.5B, GSM8K, refresh every 288: every temperature degrades before or just after the refresh and the onset is not monotone in temperature. (b) 1.5B, MATH levels 3–5 with learning rate 4×10−6 and refresh every 288 (validation on MATH-500): uncooled runs collapse in both seeds; cooled runs degrade gradually and later; Tb=0.6 is worse than 0.8 . (c) MATH at the standard learning rate and refresh 192: no run degrades and cooling is neutral.
Figure 6: Cooling without the importance correction only delays collapse. (a) 1.5B, GSM8K, refresh every 96. Plain GRPO with a cooled sampler and no correction (three seeds, plus a variant that cools further whenever the KL grows) collapses in the second cycle; the same cooled sampler with Tis is stable, as is uncooled Tis . (b) 1.5B, MATH, refresh every 96: the uncorrected cooled run dips at step 200 and recovers by step 240.
Model
Task
N
Arm
Peak
Onset
KL > 1
Lead
Outcome
1.5B
GSM8K
192
Tis , Tb=1 , 3 seeds
0.81/0.80/0.83
190/190/190
179/171/181
+11/+19/+9
sustained, end 0.61/0.59/0.47
1.5B
GSM8K
192
Tis , Tb=0.8 , 4 runs
0.85–0.87
–
–
–
stable, end 0.84–0.86
1.5B
GSM8K
192
Tis , Tb=0.9
0.86
–
–
–
one transient dip at 190, end 0.84
1.5B
GSM8K
192
Tis + controller
0.84
300
–
–
sustained, end 0.19 (KL max 0.43)
1.5B
GSM8K
288
Tis , Tb=1
0.81
190
179
+11
recovers after refresh, end 0.79
1.5B
GSM8K
288
Tis , Tb=0.8
0.85
250
245
+5
recovers after refresh, end 0.85
Table 2: Degradation events under the fixed criteria of Section 2 . Onset: first validation point ≤ running max −0.10 ; KL > 1: first step with learner–sampler KL above one nat per token; lead = onset − KL > 1 (positive: KL crossed first). Peak: highest validation point of the run. Runs with no onset within 300 steps are listed as stable. Complete per-run numbers are in Appendix D .
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Adaptive and scheduled temperature. (a) Refresh every 96: a saturation-triggered controller (with and without a KL gate) and a linear schedule from 0.8 to 1.0 all end within 0.01 of fixed cooling and of uncooled Tis ; a doubled learning rate reaches 0.85 earlier but suffers one transient collapse at step 200. (b) Refresh every 192: the controller heats the sampler to 1.05 in the second cycle and collapses.
Run
Tb
KL
Entropy at start
Clip
7B, N=144 , seed 1 (full)
0.8
0.134 / 0.068
0.27 / 0.16
0.19 / 0.01
7B, N=144 , seed 1 (cut)
0.8
0.125 / 0.126
0.27 / 0.14
0.19 / 0.02
7B, N=144 , seed 43
0.8
0.080 / 0.218
0.28 / 0.11
0.19 / 0.03
7B, N=144 , seed 1
1.0
0.31 / 15.6
0.63 / 1.29
0.25 / 0.93
7B, N=144 , seed 43
1.0
0.32 / 23.1
0.63 / 0.03
0.28 / 0.95
1.5B, N=192 , 3 seeds
0.8
0.03 / 0.01
0.28–0.30 / 0.17–0.19
0.14–0.16 / 0.05–0.08
Appendix
Table 3: Cycle-aligned readings. KL: learner–sampler KL averaged over cycle ages 55–65; entropy: learner entropy averaged over the first five steps of the cycle; clip: fraction of sampled responses at the length limit at ages 55–65. Cycle 1 / cycle 2.
Figure 8: Every run at refresh interval 144: two data seeds per arm, plus a first attempt at the cooled seed-1 arm that ended at step 287 (it matches the full run through the first refresh and drops at step 280). No cooled realization degrades before the first refresh; two of three show a late-second-cycle drop (seed 1 first attempt at 280; seed 43 from 230, KL above one nat and entropy rising, recovering at the refresh at 288). Both uncooled runs degrade severely before the first refresh, and after it their sampler produces length-clipped responses with training reward below 0.2 for the whole second cycle. All are listed in Table 4 .
Model
Task
N
lr
Correction
Tb
Tℓ
Seed
Peak
Final
Onset
Severe
KL > 1
1.5B
GSM8K
64
2e-6
none
1.0
1.0
1
0.822
0.723
120
190
159
1.5B
GSM8K
96
2e-6
TIS
0.8
1.0
1
0.861
0.856
–
–
–
1.5B
GSM8K
96
2e-6
TIS
0.8
tied
1
0.864
0.864
–
–
–
1.5B
GSM8K
96
2e-6
TIS
0.8
1.0
43
0.865
0.860
–
–
–
1.5B
GSM8K
96
2e-6
TIS
0.8
1.0
44
0.867
0.865
–
–
–
1.5B
GSM8K
96
2e-6
TIS
1.0
1.0
1
0.866
0.861
–
–
–
Appendix
Table 4: Every 300-step run. Peak and final validation accuracy (GSM8K test or MATH-500), degradation onset ( ≤ running max −0.10 ), severe degradation ( ≤ max −0.15 ), first step with learner–sampler KL above 1 nat, and KL maximum. Tb : sampler temperature (* = initial value of a controller/schedule); Tℓ : learner temperature (tied = equal to Tb ). Seed: data seed. ‘cut’ = run interrupted; ‘re-run’ = sampling replicate.