Distilling LLM Reasoning into Graph of Concept Predictors
Organizations: Department of Computer Science Emory University Atlanta, USA
Abstract
Deploying Large Language Models (LLMs) for discriminative workloads is often limited by inference latency, compute, and API costs at scale. Active distillation reduces these costs by querying an LLM oracle to train small discriminative students, but most pipelines distill only final labels, discarding intermediate reasoning signals and offering limited diagnostics of what reasoning is missing and where errors arise. We propose Graph of Concept Predictors (GCP), a reasoning-aware active distillation framework in which the teacher's reasoning is elicited as a directed acyclic graph of intermediate concepts and mirrored in the student. GCP enhances sample efficiency through a graph-aware acquisition strategy that weights per-concept uncertainty, gradient diversity, and coverage by node centrality. Additionally, it improves training stability and efficiency by performing targeted sub-module retraining, which attributes downstream loss to specific concept predictors and updates only the most influential modules. Experiments on eight NLP classification benchmarks demonstrate that GCP enhances performance under limited annotation budgets while yielding more interpretable and controllable training dynamics. Code is available at https://github.com/Ziyang-Yu/GCP.
Figures & tables
| Method | MNLI | GoEmotions | SemEval | MIMIC-III | AG News | Amazon | IMDB | Yelp |
|---|---|---|---|---|---|---|---|---|
| Random | 51.21 | 23.06 | 30.33 | 75.21 | 82.87 | 86.60 | 79.21 | 89.19 |
| Classic uncertainty- and diversity-based baselines | ||||||||
| Least Confidence | 57.42 | 25.01 | 35.18 | 80.32 | 83.05 | 88.94 | 82.37 | 91.84 |
| Entropy Sampling | 59.30 | 25.98 | 37.46 | 81.51 | 83.27 | 89.63 | 83.41 | 92.56 |
| Disagreement (Vote Entropy) | 58.91 | 25.44 | 36.87 | 81.03 | 83.11 | 89.21 | 83.02 | 92.19 |
| Representation- and gradient-based strong baselines | ||||||||
| Architecture | MNLI | GoEmotions | SemEval | MIMIC-III | AG News | Amazon | IMDB | Yelp |
|---|---|---|---|---|---|---|---|---|
| GCP | 68.41 | 31.92 | 47.83 | 89.27 | 87.96 | 93.42 | 91.18 | 95.63 |
| Flat CBM | 65.73 | 30.84 | 44.26 | 86.12 | 87.21 | 92.67 | 89.94 | 94.71 |
| End-to-End MLP | 63.18 | 29.97 | 41.83 | 84.03 | 86.74 | 92.11 | 88.92 | 94.02 |
| Setting | MNLI | GoEmotions | SemEval | MIMIC-III | AG News | Amazon | IMDB | Yelp |
|---|---|---|---|---|---|---|---|---|
| GCP (Full) | 68.41 | 31.92 | 47.83 | 89.27 | 87.96 | 93.42 | 91.18 | 95.63 |
| w/o SWU | 66.82 | 30.74 | 45.61 | 87.43 | 87.21 | 92.81 | 90.32 | 94.91 |
| w/o Grad | 66.41 | 30.28 | 45.02 | 87.11 | 87.08 | 92.64 | 89.97 | 94.74 |
| w/o Coverage | 67.02 | 30.91 | 46.13 | 87.86 | 87.35 | 92.93 | 90.51 | 95.02 |
| w/o Intersection (Union + tie-break) | 65.94 | 30.12 | 44.58 | 86.73 | 86.89 | 92.41 | 89.63 | 94.38 |
| w/o Sub-module Retraining | 64.87 | 29.96 | 43.21 | 85.02 | 86.71 | 92.18 | 89.11 | 94.09 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | # Instances | # Classes | Avg. Length | Task Type |
|---|---|---|---|---|
| AG News | 127,600 | 4 | 38 | Topic |
| Amazon Review Polarity | 4,000,000 | 2 | 90 | Sentiment |
| IMDB | 50,000 | 2 | 230 | Sentiment |
| Yelp Review Polarity | 598,000 | 2 | 150 | Sentiment |
| MNLI | 433,000 | 3 | 33 | NLI |
| GoEmotions | 58,000 | 27 | 30 | Emotion |
| Model | median | ||
| MNLI ( ) | |||
| MLP | 0.63 | 1.9 | negation scope |
| flat CBM | 1.10 | 3.2 | negation scope |
| GCP | 1.80 | 4.6 | negation scope |
| MIMIC-III ( ) | |||
| MLP | 0.51 | 1.5 | lab abnormality |
| Method | MNLI | SemEval | MIMIC-III | Yelp |
|---|---|---|---|---|
| Entropy | 38 | 21 | 64 | 73 |
| CoreSet | 62 | 34 | 102 | 121 |
| BADGE | 94 | 51 | 158 | 187 |
| CAL | 112 | 67 | 174 | 213 |
| GCP | 137 | 82 | 211 | 248 |
| Dataset | ||||
|---|---|---|---|---|
| Topic / sentiment | ||||
| AG News | 32 | 6 | 11 | 7 |
| Amazon | 35 | 8 | 15 | 9 |
| IMDB | 31 | 5 | 9 | 6 |
| Yelp | 33 | 7 | 13 | 9 |
| Multi-factor reasoning | ||||
| Dataset | Example concepts |
|---|---|
| AG News | geopolitical entity, market terminology, scientific jargon |
| Amazon | product defect mention, comparative phrasing, sentiment polarity |
| IMDB | cast praise, plot critique, recommendation cue |
| Yelp | service quality, food descriptor, price sentiment |
| MNLI | lexical overlap, negation scope, temporal mismatch |
| GoEmotions | intensifier, second-person address, gratitude cue |