Mitigating the Length-Scaling Tax with Online Distillation
Organizations: ByteDance Seed · Tongji University · Peking University
Abstract
Length scaling during reinforcement-learning (RL) post-training is often viewed as a sign of improved reasoning ability, especially on difficult problems, but may also make responses to already-solved problems unnecessarily verbose. We quantify this side effect as the length-scaling tax (LST): excess response length on already-solved queries without a commensurate accuracy gain. To mitigate LST, we propose Length Self-Distillation (LSD), which routes solved prompts to on-policy distillation and retains the original RL objective for unsolved prompts. LSD uses an exponential moving average of the online policy as its teacher, requiring no external model. We find that LSD achieves comparable or better performance than RL across multiple variants, while substantially curbing response-length growth on easy queries. LSD reduces LST from 19.0% to -3.7% on single-turn reasoning and from 31.4% to 13.7% on multi-turn agentic tasks, demonstrating that LSD effectively preserves concise response patterns on easy queries while supporting efficient exploration on difficult queries during RL post-training.
Figures & tables
| Setting | Budget | Pass@1 | ||||||
|---|---|---|---|---|---|---|---|---|
| All-4k | 4096 | 240 | pp | |||||
| All-8k | 8192 | 25.5\,{\scriptscriptstyle{\color[rgb]{1,0,0}\uparrow 4.5}} | 40.9\,{\scriptscriptstyle{\color[rgb]{1,0,0}\uparrow 22.8}} | 38.0\,{\scriptscriptstyle{\color[rgb]{1,0,0}\uparrow 21.3}} | 36.5\,{\scriptscriptstyle{\color[rgb]{1,0,0}\uparrow 22.8}} | 37.4\,{\scriptscriptstyle{\color[rgb]{1,0,0}\uparrow 21.1}} | pp | |
| All-4k | 4096 | 400 | pp | |||||
| All-4k-8k | 33.7\,{\scriptscriptstyle{\color[rgb]{0,0.5,0}\downarrow 9.2}} | 42.2\,{\scriptscriptstyle{\color[rgb]{0,0.5,0}\downarrow 2.4}} | 38.8\,{\scriptscriptstyle{\color[rgb]{0,0.5,0}\downarrow 4.2}} | 37.2\,{\scriptscriptstyle{\color[rgb]{0,0.5,0}\downarrow 2.8}} | 39.2\,{\scriptscriptstyle{\color[rgb]{0,0.5,0}\downarrow 4.2}} | pp | ||
| All-4k-8k | 31.4\,{\scriptscriptstyle{\color[rgb]{0,0.5,0}\downarrow 11.5}} | 42.6\,{\scriptscriptstyle{\color[rgb]{0,0.5,0}\downarrow 2.0}} | 39.3\,{\scriptscriptstyle{\color[rgb]{0,0.5,0}\downarrow 3.7}} | 34.8\,{\scriptscriptstyle{\color[rgb]{0,0.5,0}\downarrow 5.2}} | 39.3\,{\scriptscriptstyle{\color[rgb]{0,0.5,0}\downarrow 4.1}} | pp | ||
| Hard-4k | 4096 | 400 | \mathbf{53.9}\,{\scriptscriptstyle{\color[rgb]{1,0,0}\uparrow 11.0}} | \mathbf{58.0}\,{\scriptscriptstyle{\color[rgb]{1,0,0}\uparrow 13.4}} | \mathbf{54.7}\,{\scriptscriptstyle{\color[rgb]{1,0,0}\uparrow 11.7}} | \mathbf{49.7}\,{\scriptscriptstyle{\color[rgb]{1,0,0}\uparrow 9.7}} | \mathbf{54.7}\,{\scriptscriptstyle{\color[rgb]{1,0,0}\uparrow 11.3}} | pp |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| (a) Task performance (%) | ||||||||
|---|---|---|---|---|---|---|---|---|
| AMC | AIME25 | AIME26 | Avg. Acc. | |||||
| Method | Pass@1 | Pass@32 | Pass@1 | Pass@32 | Pass@1 | Pass@32 | Pass@1 | Pass@32 |
| RL | 60.09 | 84.34 | 22.19 | 43.33 | 16.04 | 36.67 | 42.90 | 65.73 |
| CRISP | 58.48 | 86.75 | 18.36 | 36.67 | 13.18 | 33.33 | 40.56 | 65.04 |
| Fixed SG-FKL | 45.07 | 87.95 | 11.46 | 40.00 | 8.96 | 36.67 | 30.44 | 67.13 |
| LSD (SG-FKL) | 61.30 | 86.75 | 22.92 | 43.33 | 18.15 | 36.67 | 44.20 | 67.13 |
| Method | BrowseComp-Plus | ||||||
|---|---|---|---|---|---|---|---|
| Pass@1 | Avg. Turns | Easy Len. | Hard Len. | ||||
| RL | 29.50 | 5.41 | 3531.9 | 4065.3 | 42.3 | 31.4 | 11.8 |
| RL + Length Penalty | 22.58 | 4.66 | 2708 | 3426 | 1.2 | ||
| LSD (SG-FKL) | 30.72 | 5.24 | 2706.0 | 4252.0 | 24.9 | 13.7 | 1.2 |
| LSD (SG-RKL) | 29.60 | 4.93 | 2146.5 | 4034.8 | 13.6 | 9.2 | |
| LSD (PG-RKL) | 31.65 | 5.66 | 2715.2 | 4139.5 | 31.4 | 16.1 | 2.4 |
| Avg. Pass@1 (%) | (%) | (%) | (%) | (%) | ||
|---|---|---|---|---|---|---|
| 2 | 1.00 | 40.14 | 27.79 | 14.64 | 7.29 | 15.49 |
| 4 | 1.00 | 41.85 | 29.20 | 7.63 | 2.01 | 9.06 |
| 8 | 1.00 | 40.49 | 13.01 | 11.00 | -6.21 | 2.91 |
| 4 | 0.85 | 40.91 | 27.55 | 21.24 | 10.84 | 18.42 |
| 4 | 0.75 | 39.31 | 19.15 | 10.97 | -0.68 | 8.11 |
| Pass@1 (%) | Easy seq. (%) | Easy tok. (%) | Effective OPD coef. | Param. gap ( ) | OPD loss ( ) | ||
|---|---|---|---|---|---|---|---|
| 2 | 1.00 | 40.14 | 18.87 | 10.79 | 0.252 | 0.690 | 0.614 |
| 4 | 1.00 | 41.85 | 18.54 | 10.31 | 0.247 | 1.087 | 0.693 |
| 8 | 1.00 | 40.49 | 15.24 | 7.90 | 0.192 | 1.690 | 0.841 |
| 4 | 0.85 | 40.91 | 24.67 | 14.87 | 0.354 | 1.077 | 0.715 |
| 4 | 0.75 | 39.31 | 29.50 | 19.00 | 0.450 | 1.078 | 0.689 |
| Variant | Actor entropy | Teacher–rollout gap | Easy sequences (%) |
|---|---|---|---|
| SG-FKL | 0.0743 | 0.00738 | 21.51 |
| SG-RKL | 0.0367 | 0.00481 | 20.23 |
| PG-RKL | 0.0542 | 0.00590 | 20.67 |
| Variant | Easy seq. (%) | Easy tok. (%) | Hard tok. (%) | Easy tok. median [IQR] |
|---|---|---|---|---|
| SG-FKL | 21.51 | 12.17 | 87.83 | 11.54 [8.09, 15.67] |
| SG-RKL | 20.23 | 10.58 | 89.42 | 9.99 [7.11, 13.27] |
| PG-RKL | 20.67 | 10.94 | 89.06 | 10.29 [7.31, 13.90] |
| Anchor | Method | Correct / total | Acc. (%) | Mean tokens | (pp) |
|---|---|---|---|---|---|
| RL | 318/320 | 99.38 | 783.1 | +0.00 | |
| LSD (SG-FKL) | 320/320 | 100.00 | 592.0 | +0.63 | |
| LSD (SG-RKL) | 320/320 | 100.00 | 585.3 | +0.63 | |
| LSD (PG-RKL) | 320/320 | 100.00 | 585.1 | +0.63 | |
| Fixed SG-FKL | 314/320 | 98.13 | 568.1 | -1.25 | |
| RL | 695/704 | 98.72 | 1013.5 | +0.00 |
| Method | Avg. Pass@1 (%) | Easy queries | Hard queries | (%) | ||
|---|---|---|---|---|---|---|
| Acc. (%) | Mean tokens | Acc. (%) | Mean tokens | |||
| RL | 42.90 | 98.15 | 996.9 | 32.85 | 2329.3 | 19.03 |
| LSD (SG-FKL) | 44.20 | 99.10 | 859 | 34.22 | 2511 | -3.72 |
| LSD (SG-RKL) | 42.80 | 99.43 | 757.2 | 32.50 | 2395.0 | -10.92 |
| LSD (PG-RKL) | 43.56 | 99.15 | 801.6 | 33.45 | 2392.5 | 1.45 |
| Random Routing | 22.08 | 89.74 | 895.0 | 9.78 | 2006.6 | -11.00 |
| Parameter | Single-turn reasoning | Multi-turn agentic tasks |
|---|---|---|
| Optimizer | AdamW; | |
| Learning rate / weight decay | / | |
| Training batch | 32 prompts 8 rollouts = 256 trajectories | |
| PPO epochs / advantage | 1 / GRPO | |
| PPO clip | ||
| Entropy / RL KL coef. | 0 / 0 | |
| Parameter | RL | Fixed SG-FKL | SG-FKL | SG-RKL | PG-RKL |
|---|---|---|---|---|---|
| Teacher | – | Frozen | EMA | ||
| Routing | – | Fixed map | Online rollout-group accuracy | ||
| Routing threshold | – | 1.0 | |||
| EMA half-life | – | – | 4 | ||
| Distillation objective | – | Top-32 FKL | Top-32 RKL | Sample-token RKL | |
| Distillation gradient | – | Supervised | Policy gradient | ||