cs.AIMar 11, 2026

HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation

Authors: Wenjing ZhangJiangze YanJieyun HuangYi ShenShuming ShiPing ChenNing WangZhaoxiang Liu+2 more

Organizations: Data Science & AI Research Institute, China Unicom · Unicom Data Intelligence, China Unicom

Abstract

Distilling reasoning capabilities from Large Reasoning Models (LRMs) into smaller models is typically constrained by the limitations of rejection sampling. Standard methods treat the teacher as a static filter, discarding complex "corner-case" problems where the teacher fails to explore valid solutions independently, thereby creating an artificial "Teacher Ceiling" for the student. In this work, we propose Hindsight Entropy-Assisted Learning (HEAL), an RL-free framework designed to bridge this reasoning gap. Drawing on the educational theory of the Zone of Proximal Development (ZPD), HEAL synergizes three core modules: (1) Guided Entropy-Assisted Repair (GEAR), an active intervention mechanism that detects critical reasoning breakpoints via entropy dynamics and injects targeted hindsight hints to repair broken trajectories; (2) Perplexity-Uncertainty Ratio Estimator (PURE), a ratio-based filtering heuristic that reduces high-anomaly shortcut-like rationales; and (3) Progressive Answer-guided Curriculum Evolution (PACE), a three-stage distillation strategy that organizes training from foundational alignment to hard-case adaptation. Extensive experiments on multiple benchmarks demonstrate that HEAL significantly outperforms traditional SFT distillation and other baselines.

Explore similar work

CardsList
  1. Distribution Corrected Offline Data Distillation for Large Language Models

    May 13, 2026Yumeng Zhang, Zhengbang Yang, Yevin Nikhel Goonatilake +1Reasoning Traces

  2. Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning

    Aug 13, 2025Xiaojun Wu, Xiaoguang Jiang, Huiyang Li +11Scaling Laws