cs.LGSep 30, 2026

Mitigating the Length-Scaling Tax with Online Distillation

Authors: Xu Wan, Wenyue Xu, Shengjie Zhao, Mingyang Sun

Organizations: ByteDance Seed · Tongji University · Peking University

Abstract

Length scaling during reinforcement-learning (RL) post-training is often viewed as a sign of improved reasoning ability, especially on difficult problems, but may also make responses to already-solved problems unnecessarily verbose. We quantify this side effect as the length-scaling tax (LST): excess response length on already-solved queries without a commensurate accuracy gain. To mitigate LST, we propose Length Self-Distillation (LSD), which routes solved prompts to on-policy distillation and retains the original RL objective for unsolved prompts. LSD uses an exponential moving average of the online policy as its teacher, requiring no external model. We find that LSD achieves comparable or better performance than RL across multiple variants, while substantially curbing response-length growth on easy queries. LSD reduces LST from 19.0% to -3.7% on single-turn reasoning and from 31.4% to 13.7% on multi-turn agentic tasks, demonstrating that LSD effectively preserves concise response patterns on easy queries while supporting efficient exploration on difficult queries during RL post-training.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization

    Jul 23, 2026Siwei Chen, Siqi Chen, Xupeng Miao +1Frictive Policy OptimizationResponses

  2. RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

    Jun 10, 2026Leyi Pan, Shuchang Tao, Yunpeng Zhai +5Unsupervised On-Policy Self-DistillationToken-Level Supervision

  3. Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training

    May 8, 2026Chen Wang, Hexuan Deng, Yining Zhang +5ShortReasoning Traces