cs.CLOct 8, 2026

ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery

Authors: Houcheng Jiang, Mao Zheng, Mingyang Song, Qiyong Zhong, Jie Sun, Tianyu Zhang, Junfeng Fang

Organizations: Zhongguancun Academy · Foundation Model Department, Tencent · University of Science and Technology of China · National University of Singapore

Abstract

Structured pruning reduces the deployment cost of reasoning language models, but the resulting capability degradation can hinder subsequent on-policy distillation (OPD) recovery. Because OPD relies on student-generated trajectories, pruning damage that persists after offline distillation can limit its effectiveness. We propose RECAL, Recovery-Aware Calibration, a simple plug-and-play approach that improves OPD recovery by adjusting calibration before pruning. RECAL uses forward KL between an unpruned teacher and a pruned probe to identify teacher-supported predictions disrupted by pruning, then reweights calibration statistics to guide existing pruning criteria toward preserving these predictions. Across multiple models and pruning methods, RECAL consistently improves mathematical reasoning after OPD, achieving gains of up to 16.7 percentage points on AIME, alongside improvements in most code-generation comparisons. Further analysis shows that RECAL reduces residual damage at heavily affected tokens and establishes performance advantages that persist through recovery. These results demonstrate the value of recovery-aware calibration for improving on-policy distillation recovery of pruned reasoning models.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning

    Sep 15, 2026Ha Lan Nguyen, Huy Hoang Tran, Trac-Duy Tran +1LLM PruningLarge Reasoning Models

  2. On-Policy Delta Distillation

    Jul 16, 2026Byeongho Heo, Jaehui Hwang, Sangdoo Yun +1RL for Language Model ReasoningLanguage Model Distillation

  3. ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

    Jul 14, 2026Qingyu Zhang, Qianhao Yuan, Hongyu Lin +7Open-Ended GenerationLLM Pruning