cs.AIFeb 3, 2026

Distilling LLM Reasoning into Graph of Concept Predictors

Authors: Ziyang Yu, Liang Zhao

Organizations: Department of Computer Science Emory University Atlanta, USA

Abstract

Deploying Large Language Models (LLMs) for discriminative workloads is often limited by inference latency, compute, and API costs at scale. Active distillation reduces these costs by querying an LLM oracle to train small discriminative students, but most pipelines distill only final labels, discarding intermediate reasoning signals and offering limited diagnostics of what reasoning is missing and where errors arise. We propose Graph of Concept Predictors (GCP), a reasoning-aware active distillation framework in which the teacher's reasoning is elicited as a directed acyclic graph of intermediate concepts and mirrored in the student. GCP enhances sample efficiency through a graph-aware acquisition strategy that weights per-concept uncertainty, gradient diversity, and coverage by node centrality. Additionally, it improves training stability and efficiency by performing targeted sub-module retraining, which attributes downstream loss to specific concept predictors and updates only the most influential modules. Experiments on eight NLP classification benchmarks demonstrate that GCP enhances performance under limited annotation budgets while yielding more interpretable and controllable training dynamics. Code is available at https://github.com/Ziyang-Yu/GCP.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Distribution Corrected Offline Data Distillation for Large Language Models

    May 13, 2026Yumeng Zhang, Zhengbang Yang, Yevin Nikhel Goonatilake +1Probe-Logit DistillationMathematical Reasoning Benchmarks

  2. Structural Rationale Distillation via Reasoning Space Compression

    May 8, 2026Jialin Yang, Jiankun Wang, Jiajun Wu +3Reasoning PathsDataset Distillation

  3. Distilled Reinforcement Learning for LLM Post-training

    Jul 19, 2026Chen Wang, Zhaochun Li, Jionghao Bai +4Post-TrainingOffline Reinforcement Learning