Chain-of-thought (CoT) distillation transfers reasoning behaviors from a strong teacher to a smaller student, but prior work reports a capacity gap: distillation may fail when the teacher-student capability mismatch is large. We revisit the capacity gap from a practical perspective by re-examining commonly used experimental settings. Notably, we find that CoT distillation often degrades performance compared to the student's pre-distillation baseline, and that some settings used in prior work, while suitable for establishing the capacity gap as a phenomenon, do not reflect realistic deployment scenarios. Complementing prior work that establishes the capacity gap, we evaluate its practical impact under more realistic settings and find that it does not consistently dominate; stronger teachers tend to be preferable when candidate teachers differ substantially in performance. Our results offer practical guidance for selecting teacher-student pairs in CoT distillation.
Figures & tables
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Teacher
No Filter
Filtered
Diff.
Small
67.49
65.70
−1.78
Large
67.95
64.34
−3.61
Short
68.09
64.92
−3.16
Long
69.09
66.11
−2.97
Appendix
Table 6: Average accuracy across 15 BBH tasks (averaged over four student sizes) with and without cross-teacher filtering . Diff. = Filtered − No Filter.
Task (ICL Gap)
Student
Baseline
Small
Large
Short
Long
Multistep Arithmetic Two (28.8)
0.5B
27.0
58.0
89.0
62.0
83.0
1.5B
67.0
83.0
65.0
84.0
93.0
3B
90.0
83.0
89.0
94.0
98.0
7B
95.0
73.0
88.0
96.0
82.0
Object Counting (28.9)
0.5B
48.0
53.0
40.0
43.0
26.0
1.5B
56.0
50.0
53.0
51.0
44.0
Appendix
Table 8: Accuracy on the test split for three BBH tasks whose ICL gap falls just below the 30-point threshold (in the 20–30 point band), distilled under the same protocol as the main experiments. Baseline is the pre-distillation few-shot accuracy; Small/Large (Qwen2.5-14B/72B-Instruct) and Short/Long (Qwen2.5-32B-Instruct/QwQ-32B-Preview) denote students distilled from the respective teachers. Underlined scores fall below the pre-distillation baseline.
Chain-of-Thought (CoT) distillation transfers multi-step reasoning from large reasoning models to smaller students, but verbose teacher traces inflate both training and inference cost. Existing CoT compression methods fall into two families, selective pruning and generative rewriting, yet prior studies have left key factors entangled: granularity is confounded with importance criteria in pruning, restructuring level is rarely isolated in rewriting, and compression budgets are not systematically evaluated across domains or regimes. We recast CoT compression along three dimensions: importance criterion, restructuring level, and compression budget. Sweeping these across two model families, Math and General domains, and Long-/Short-CoT regimes, we find that (i) importance criterion utility is strictly governed by granularity: step-level criteria converge on a shared reasoning backbone, while token-level pruning requires symbol-aware signals to preserve the logical core; (ii) restructuring level inverts across domains: Math degrades monotonically with structural disruption, while aggressive rewriting acts as a denoiser on General tasks; (iii) training-time compression does not necessarily translate to inference-time savings: Long-CoT students retain verbose habits despite concise supervision, making the training ratio an optimistic lower bound on deployment cost. These findings yield condition-aware guidelines for matching compression to deployment context.
Siyang Lyu, Xinghao Chen, Zhijing Sun +3
Ningbo Institute of Digital Twin, Eastern Institute of Technology, Ningbo · Viterbi School of Engineering, University of Southern California · The Hong Kong Polytechnic University +2
Reasoning distillation from powerful teacher models to smaller students faces the Gap Curse: as teachers grow more sophisticated, their complex distributions increasingly diverge from what students can approximate, causing performance degradation. Existing mitigation strategies either filter out challenging examples through data selection or introduce weaker intermediate assistant models, inherently compromising supervision coverage or quality. We propose Teacher Alignment, which directly adapts the teacher toward the student's distribution without discarding data or degrading reasoning quality. However, naive alignment through standard knowledge distillation triggers catastrophic collapse of the teacher's reasoning capabilities. To address this, we reformulate teacher alignment as reinforcement learning and introduce TeacherGRPO, built on Group Relative Policy Optimization with two key innovations: (i) Curriculum Selective Alignment applies dual token- and distribution-level curricula to focus rewards on high-signal reasoning gaps while filtering noise from trivial tokens and uncertain tail distributions, and (ii) Importance-Adaptive Length Regularization selectively penalizes verbose redundancy while preserving pedagogically critical reasoning steps. The aligned teacher then distills knowledge to students via standard pipelines. Extensive experiments show TeacherGRPO significantly outperforms baselines across diverse reasoning benchmarks and distillation methods. Our code is available at https://github.com/LzyFischer/TeacherGRPO.
Zhenyu Lei, Zihan Chen, Yaochen Zhu +5
University of Virginia · Netflix · University of Washington +2
Chain-of-Thought (CoT) reasoning has significantly improved LLMs' mathematical problem-solving capabilities, but distilling such capabilities into smaller models remains challenging due to the capacity mismatch between verbose teachers and compact students. Directly copying teachers' lengthy reasoning chains causes capacity overload, resulting in truncated outputs or repetitive failure. Existing remedies each sacrifice a critical property of CoT: implicit reasoning methods (e.g., compressing reasoning into hidden states) trade away interpretability and verifiability, while heuristic compression strategies (e.g., random step pruning) destroy logical integrity. To address this, we propose BRIDGE, a curriculum framework that first establishes structural understanding via masked reconstruction, then uses GRPO-based reinforcement learning to guide students in self-discovering the optimal balance between accuracy and brevity, and finally internalizes complex reasoning through teacher-guided rewriting on failure cases. On GSM8K, BRIDGE enables Qwen2.5-3B to achieve 11.29% accuracy improvement and 27.4% token reduction over the original model, outperforming instruction-tuned variants and distillation baselines. Zero-shot transfer experiments on SVAMP and MATH-500 further confirm the generalization of internalized reasoning. Our code and model checkpoints are publicly available at https://github.com/Applied-Machine-Learning-Lab/SDM2026_BRIDGE and https://huggingface.co/bowen0815/BRIDGE.
Bowen Yu, Sheng Zhang, Binhao Wang +8
1City University of Hong Kong · 2Mohamed Bin Zayed University of Artificial Intelligence