Deploying Large Language Models (LLMs) for discriminative workloads is often limited by inference latency, compute, and API costs at scale. Active distillation reduces these costs by querying an LLM oracle to train small discriminative students, but most pipelines distill only final labels, discarding intermediate reasoning signals and offering limited diagnostics of what reasoning is missing and where errors arise. We propose Graph of Concept Predictors (GCP), a reasoning-aware active distillation framework in which the teacher's reasoning is elicited as a directed acyclic graph of intermediate concepts and mirrored in the student. GCP enhances sample efficiency through a graph-aware acquisition strategy that weights per-concept uncertainty, gradient diversity, and coverage by node centrality. Additionally, it improves training stability and efficiency by performing targeted sub-module retraining, which attributes downstream loss to specific concept predictors and updates only the most influential modules. Experiments on eight NLP classification benchmarks demonstrate that GCP enhances performance under limited annotation budgets while yielding more interpretable and controllable training dynamics. Code is available at https://github.com/Ziyang-Yu/GCP.
Figures & tables
Figure 1: Mirroring a teacher’s reasoning DAG into a Graph of Concept Predictors, shown on a Yelp review. Left: the LLM teacher judges the food ( R1 ) and the service ( R2 ), combines them into an overall impression ( R3 ), and labels the review negative. Right: GCP gives each conclusion its own concept predictor fj and wires the predictors with the same edges; every fj also reads x . The teacher’s intermediate answers serve as labels for the matching concepts cj .
Figure 2: Overview of the GCP loop: the student (lower left), graph-aware acquisition (top), and sub-module retraining by counterfactual reruns (lower right).
Figure 3: A correct but under-confident prediction, on the review from Figure 1 . The student’s service predictor f2 misses the complaint about waiting (red), and the error passes to c3 and y^ (dashed). The student still predicts negative , but with lower confidence and higher loss.
Method
MNLI
GoEmotions
SemEval
MIMIC-III
AG News
Amazon
IMDB
Yelp
Random
51.21
23.06
30.33
75.21
82.87
86.60
79.21
89.19
Classic uncertainty- and diversity-based baselines
Least Confidence
57.42
25.01
35.18
80.32
83.05
88.94
82.37
91.84
Entropy Sampling
59.30
25.98
37.46
81.51
83.27
89.63
83.41
92.56
Disagreement (Vote Entropy)
58.91
25.44
36.87
81.03
83.11
89.21
83.02
92.19
Representation- and gradient-based strong baselines
Table 1: Baseline comparison under a fixed 20% token budget. All methods are evaluated in a pool-based active learning setting and spend the same number of teacher tokens, so the label-only baselines annotate more texts than GCP. Best results are in bold .
Figure 4: Performance curves of different sample selection methods for active learning. The y-axis denotes classification accuracy, and the x-axis denotes the teacher-token budget as a percentage of the tokens needed to label the full training set with labels only. At the same budget, GCP annotates fewer texts than the baselines because each of its queries also returns concept values.
Architecture
MNLI
GoEmotions
SemEval
MIMIC-III
AG News
Amazon
IMDB
Yelp
GCP
68.41
31.92
47.83
89.27
87.96
93.42
91.18
95.63
Flat CBM
65.73
30.84
44.26
86.12
87.21
92.67
89.94
94.71
End-to-End MLP
63.18
29.97
41.83
84.03
86.74
92.11
88.92
94.02
Table 2: Architecture ablation. Comparison of GCP with a flat CBM (no dependencies) and an end-to-end MLP at 50% annotation budget.
Setting
MNLI
GoEmotions
SemEval
MIMIC-III
AG News
Amazon
IMDB
Yelp
GCP (Full)
68.41
31.92
47.83
89.27
87.96
93.42
91.18
95.63
w/o SWU
66.82
30.74
45.61
87.43
87.21
92.81
90.32
94.91
w/o Grad
66.41
30.28
45.02
87.11
87.08
92.64
89.97
94.74
w/o Coverage
67.02
30.91
46.13
87.86
87.35
92.93
90.51
95.02
w/o Intersection (Union + tie-break)
65.94
30.12
44.58
86.73
86.89
92.41
89.63
94.38
w/o Sub-module Retraining
64.87
29.96
43.21
85.02
86.71
92.18
89.11
94.09
Table 3: Ablation study on major components of GCP . Accuracy (%) at 50% annotation budget across eight datasets. Each row disables one component; the full model includes all components.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
# Instances
# Classes
Avg. Length
Task Type
AG News
127,600
4
38
Topic
Amazon Review Polarity
4,000,000
2
90
Sentiment
IMDB
50,000
2
230
Sentiment
Yelp Review Polarity
598,000
2
150
Sentiment
MNLI
433,000
3
33
NLI
GoEmotions
58,000
27
30
Emotion
Appendix
Table 4: Summary of dataset statistics. # Instances is the size of the full dataset, with all splits combined.
Model
μmin
median μv
v⋆
MNLI ( ∣V∣=16 )
MLP
0.63
1.9
negation scope
flat CBM
1.10
3.2
negation scope
GCP
1.80
4.6
negation scope
MIMIC-III ( ∣V∣=18 )
MLP
0.51
1.5
lab abnormality
Appendix
Table 5: Estimated node-level PL constants (all values ×10−2 ). μmin=minvμv is the quantity that controls the rate in Theorem 2 under uniform λv ; v⋆ is the node attaining it.
Figure 5: Ablation study of targeted sub-module retraining. The y-axis denotes classification accuracy and the x-axis denotes teacher-token usage (millions; thousands for SemEval). Budgets are set per dataset and span roughly 8–50% of the tokens needed to label its full training set with labels only (cf. the 5–50% range of Figure 4 ).
Figure 6: Compute cost versus annotation scale on a logarithmic axis. Comparison between the LLaMA-3-70B teacher and GCP.
Method
MNLI
SemEval
MIMIC-III
Yelp
Entropy
38
21
64
73
CoreSet
62
34
102
121
BADGE
94
51
158
187
CAL
112
67
174
213
GCP
137
82
211
248
Appendix
Table 6: Per-round runtime (seconds) at 20% budget.
Dataset
∣C∣
∣V∣
∣Ecand∣
∣E∣
Topic / sentiment
AG News
32
6
11
7
Amazon
35
8
15
9
IMDB
31
5
9
6
Yelp
33
7
13
9
Multi-factor reasoning
Appendix
Table 7: Concept graph statistics per dataset. ∣C∣ : candidate concepts elicited in Stage 1; ∣V∣ : nodes retained after pruning; ∣Ecand∣ : teacher-proposed dependencies whose endpoints both survive pruning, i.e. those actually submitted to the CMI test; ∣E∣ : edges retained after that test. Node count and edge density both track reasoning depth rather than input length.
Distilling reasoning traces from strong large language models into smaller ones is a promising route to improve intelligence in resource-constrained settings. Existing approaches face a fundamental trade-off: offline distillation from teacher-generated traces provides high-quality, sample-efficient supervision but suffers from distributional drift: during training, the student model conditions on teacher-generated prefixes, whereas during inference the student autoregresses on self-generated prefixes, leading to compounding errors over long reasoning trajectories. Meanwhile, on-policy or self-distillation methods better match the student's inference-time distribution, but require costly online sampling and often produce low-quality traces in early training. We propose a principled offline reasoning distillation framework that preserves the efficiency and supervision quality of offline teacher-generated data while correcting teacher-student distribution drift. It adaptively emphasizes teacher supervision that is better aligned with the student's on-policy distribution. Evaluations on mathematical reasoning benchmarks of GSM8K, MATH, MATH500, and harder held-out competition-style tasks, including AMC, AIME, and OlympiadBench, show that our method improves reasoning accuracy over prior offline distillation algorithms and yields more stable reasoning traces while preserving instruction-following capabilities. Our work shows that lightweight, distribution-correction-aware training can substantially strengthen offline reasoning distillation without online rollouts.
When distilling reasoning from large language models (LLMs) into smaller ones, teacher rationales for similar problems often vary wildly in structure and strategy. Like a chef who makes the same dish differently each time, this inconsistency burdens the student with noisy supervision that is hard to internalize. We propose Distillation through Reasoning Path Compression (D-RPC), which constrains the teacher to follow a compact, dynamically maintained bank of reusable high-level reasoning paths. For each training question, D-RPC retrieves the most relevant path and conditions the teacher to follow it, producing rationales that are consistent across similar problems yet diverse enough to cover different problem types. A PAC-Bayes analysis formalizes the resulting trade-off between bank size and coverage: smaller banks reduce supervision entropy but risk coverage gaps, and the generalization bound identifies an optimal intermediate size confirmed by our ablations. Across five math and commonsense reasoning benchmarks with two student models, D-RPC consistently outperforms chain-of-thought distillation, freeform rationale generation, direct distillation, and structured-supervision baselines, while using fewer tokens than template-heavy alternatives.
Jialin Yang, Jiankun Wang, Jiajun Wu +3
University of Calgary · Calgary, Canada · University of Michigan +1
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on coarse-grained outcome supervision, resulting in difficult credit assignment and limited capability to acquire new knowledge. OPD, meanwhile, unconditionally matches teacher logits through KL divergence, which creates a dilemma: similar teachers provide little new knowledge, while substantially different teachers often yield ineffective guidance, largely restricting OPD to within-family distillation. We propose Distilled Reinforcement Learning (Distilled RL), which integrates teacher supervision into the RL objective to provide fine-grained guidance, selectively transfer new knowledge and avoid unconditional imitation. Distilled RL contains three components: reverse importance sampling with clipping, negative sample reset, and sequence-level geometric normalization. Through a concise and interpretable case study, we demonstrate that Distilled RL can effectively transfer previously unavailable knowledge from a teacher model to a student model. Extensive experiments across both within-family and cross-family distillation settings show that Distilled RL substantially outperforms standard RL and OPD in terms of both pass@1 and pass@k. Our code is available at https://github.com/597358816/Distilled-RL.
Chen Wang, Zhaochun Li, Jionghao Bai +4
College of Elite Engineers, Nankai University · Zhongguancun Academy · Beijing Institute of Technology +4