Authors: Rui Cai, Wenhui Zhu, Xiwen Chen, Jincheng Cao, Han Yu, Shayan Mohajer Hamidi, Zelin He, Qiyao Ma, +10 more
Organizations: AI Agentic Modeling and Foundation Team, LinkedIn · University of California, Davis · Arizona State University · Clemson University · Pennsylvania State University · University of Wisconsin–Madison
Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student, making generation costly at scale. We find neither assumption holds: across students from 7B to 72B, smaller frozen models generate rejects with less inference compute yet train stronger students than self-generated rejects, before and after sequence-level knowledge distillation, on code generation and mathematical reasoning. To explain this result, we derive a finite-horizon utility bound for Direct Preference Optimization in a linearized feature model. The bound characterizes favorable reject distributions and motivates three interventions. First, mixing rejects from smaller and student-scale models improves performance as the smaller model's share increases. Second, reassigning rejects to other prompts and shuffling their code tokens still outperform length-matched gibberish, showing that task structure contributes to reject utility. Third, selecting candidates with lower likelihood under the reference policy improves net transfer when higher-likelihood candidates provide less useful contrast. Lower-likelihood selections outperform higher-likelihood ones for every source. These results suggest that effective rejects preserve task structure while limiting coupling to the reference policy, and that smaller frozen models can provide them at low cost.
Figures & tables
Figure 1 : Smaller Base models provide stronger rejects, and their advantage varies systematically with source composition. Left: KoDCode avg@4 for the original vanilla model, both Self constructions, and every smaller Base source available to each 7B–72B student. All 18 smaller Base configurations exceed both post-SeqKD Self and vanilla Self. Right: a randomized mixture intervention on a matched 14B population. Avg@4 rises from 59.73 to 63.28 as the share of Llama-1B Base rejects increases from 0% to 100% .
Figure 2 : Smaller Base models improve preference distillation across code generation and mathematical reasoning. (a, b) Absolute avg@4 for the original vanilla model, SeqKD, Continued-SFT, both Self constructions, and the best observed smaller Base source for each 7B–72B student. Labels above the green bars identify the displayed reject source. (c) Two randomized source assignment interventions vary the proportions of smaller Base and vanilla Self rejects for a 14B student. The green and gray background shows the dataset composition. Performance increases with the smaller Base share for both source families.
Figure 3 : Localizing and decomposing the utility of smaller Base rejects. (a) Absolute KoDCode avg@4 across four control settings at every student scale. Continued-SFT trains on the preference-stage chosen responses alone. Gibberish supplies length-matched artificial rejects. Prompt reassignment preserves real Base rejects while removing their original prompt correspondence, and Aligned restores that correspondence. (b) Lexical permutation preserves the generated code-token inventory of prompt-reassigned Base rejects while destroying token order, syntax, and coherent program semantics. Aligned shows the original prompt-matched Base construction.
Figure 4 : Reference likelihood tracks and controls reject utility. (a) Each point is a natural reject source scored by its mean length-normalized likelihood under the corresponding student’s SeqKD reference and standardized within prompt. Lines are fitted separately for each student scale; labels report Pearson correlations. Filled circles are smaller Base sources, squares are vanilla Self, and diamonds are post-SeqKD Self. (b) For a fixed Qwen2.5-14B-Instruct student, each source-specific candidate bank is reselected toward lower reference likelihood (Repelled), without likelihood preference (Native), or toward higher reference likelihood (Attracted).
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
SeqKD
DPO
Learning rate
5×10−6
5×10−7
DPO β
–
0.1
Epochs
1
1
Gradient accumulation
16
16
Effective batch (7B–72B)
256
256
Max sequence length
2,560
2,560
Appendix
Table 1: Training hyperparameters of the code line. All runs use full-parameter training with AdamW, a cosine learning rate schedule with 10% linear warmup, bf16 precision, gradient checkpointing, DeepSpeed ZeRO-3. Continued-SFT uses the same settings as DPO.
Student
Smaller Base source
Parameter ratio
Reject tokens
Decode cost
Total cost
Reduction
7B
0.5B
6.5%
11.317M
10.6%
8.2%
12.2×
7B
1.5B
20.3%
6.366M
18.7%
19.6%
5.1×
7B
3B
40.5%
7.074M
41.5%
40.9%
2.4×
14B
0.5B
3.3%
11.317M
5.1%
4.1%
24.3×
14B
1.5B
10.5%
6.366M
9.0%
9.8%
10.2×
14B
3B
20.9%
7.074M
20.0%
20.5%
4.9×
Appendix
Table 2 : Complete output-length-adjusted parameter-token cost comparison. Decode and total costs are percentages of same-scale vanilla Self. The reduction factor is the inverse of the total-cost ratio.
pass@1
Stochastic
Method
Reject Source
Overall
Easy
Medium
Hard
avg@4
pass@4
Student: Qwen2.5-7B-Instruct
Base
–
48.9
62.3
45.3
40.2
47.9
59.8
SeqKD
–
58.8
70.4
56.2
50.7
55.6
69.4
Continued-SFT
–
59.6
71.6
57.0
52.0
56.3
70.0
DPO
0.5B
60.5
72.0
57.5
53.1
58.5
70.5
Appendix
Table 3 : KoDCode held-out results. pass@1 uses greedy decoding; avg@4 is the mean success rate over four stochastic samples, and pass@4 is the fraction of problems solved by at least one sample. Stochastic decoding uses temperature 0.6 and top- p=1.0 . All values are percentages. Bold marks the best result within each student and metric.
Method
Reject Source
BCB-C
BCB-I
HumanEval
MBPP
Weighted Avg.
Student: Qwen2.5-7B-Instruct
Base
–
42.3
29.1
84.1
79.0
42.8
SeqKD
–
43.4
32.5
83.3
79.1
44.6
Continued-SFT
–
43.9
33.6
82.3
79.1
45.2
DPO
0.5B
45.2
35.4
84.8
78.2
46.6
1.5B
45.0
33.3
82.3
80.2
45.7
Appendix
Table 4 : External pass@1 across four code-generation benchmarks. Weighted Avg. weights each benchmark by its number of problems: 1,140 each for BCB-C and BCB-I, 164 for HumanEval, and 257 for MBPP. All values are percentages. Bold marks the best result within each student and column.
Method
Reject Source
MATH-500 avg@4
AIME avg@4
Teacher (DeepSeek-R1)
–
97.3
74.9
Student: Qwen2.5-72B-Instruct
Base
–
84.0
15.8
SeqKD
–
91.0
37.9
Continued-SFT
–
91.2
37.5
DPO
0.5B
91.6
40.0
Appendix
Table 5 : MATH-500 and AIME avg@4 over four samples.
Figure 5 : Cross-family reject sources. Bars report avg@4 on (a) KoDCode in-domain and (b) MATH-500 for the 14B, 32B, and 72B Qwen2.5 students. Self is the better of vanilla Self and post-SeqKD Self. Labels on the green bars identify the best Qwen2.5 smaller Base source; labels on the purple bars identify the Llama source.
Student
Smaller Base source
Base share
Code avg@4
Code pass@4
14B
Llama-3.2-1B-Instruct
0%
59.730
72.380
25%
61.575
73.000
50%
62.035
72.760
75%
62.670
73.380
100%
63.280
73.180
14B
Qwen2.5-3B-Instruct
0%
59.575
72.280
Appendix
Table 6 : Complete randomized source assignment results. The Base share is the proportion of prompts assigned rejects from the indicated smaller Base source. All remaining prompts use vanilla Self rejects from the student-scale model.
Student
Reject source
Continued-SFT
Gibberish
Prompt-reassigned
Aligned
7B
Qwen2.5-0.5B-Instruct
56.30
57.37
58.65
58.53
14B
Qwen2.5-3B-Instruct
60.89
60.92
61.97
62.63
32B
Qwen2.5-7B-Instruct
64.50
64.52
65.25
65.53
72B
Qwen2.5-7B-Instruct
65.00
65.24
66.36
67.53
Appendix
Table 7 : KoDCode avg@4 across the reject-control ladder. Prompt-reassigned arms use real Base rejects reassigned across prompts. Each row is one completed prompt-reassignment experiment and reports its Continued-SFT, gibberish, reassigned, and aligned settings.
Figure 6 : Implicit reward dynamics for gibberish, prompt reassigned, and prompt-matched Base rejects at each student scale. Gibberish is suppressed most strongly but does not recover the utility of real Base rejects.
Student
Gibberish
Lexical
Real reassigned
Lexical − Gibberish
Recovery
7B
57.370
57.900
58.645
+0.530
41.6%
14B
60.920
61.650
61.970
+0.730
69.5%
32B
64.520
65.195
65.250
+0.675
92.5%
72B
65.235
65.685
66.360
+0.450
40.0%
Appendix
Table 8 : Cross-scale lexical-structure interventions. Recovery is the Lexical gain over Gibberish divided by the real prompt-reassigned gain over Gibberish.
Student
Reject source
Source class
Mean log S /token
Prompt z
Mean tokens
avg@4
7B
Qwen2.5-0.5B-Instruct
Subscale Base
−0.3725
−1.048
466.7
58.525
7B
Qwen2.5-1.5B-Instruct
Subscale Base
−0.2527
−0.113
262.5
57.395
7B
Qwen2.5-3B-Instruct
Subscale Base
−0.2851
−0.435
291.8
57.705
7B
Qwen2.5-7B-Instruct
Vanilla Self
−0.1790
+0.418
284.9
54.550
7B
post-SeqKD 7B
Post-SeqKD Self
−0.0915
+1.179
243.7
52.745
14B
Qwen2.5-0.5B-Instruct
Subscale Base
−0.3848
−1.074
466.7
62.525
Appendix
Table 9 : Natural reject-source measurements underlying Figure 4 a. Each source is scored on the same 24,248 preference prompts under the corresponding student’s SeqKD reference. Mean log likelihood is length normalized. Prompt z is the mean reference likelihood after standardization within each prompt. Utility is KoDCode avg@4 on 5,000 held-out problems.
Student
n
r
ρ
p−
p2
Prompt- zr
LOO r range
7B
5
−0.967
−1.000
0.0083
0.0083
−0.968
[−0.988,−0.920]
14B
6
−0.868
−0.886
0.0083
0.0083
−0.871
[−0.909,−0.645]
32B
7
−0.752
−0.464
0.0298
0.0484
−0.770
[−0.863,−0.405]
72B
8
−0.836
−0.690
0.0057
0.0076
−0.842
[−0.931,−0.684]
Appendix
Table 10 : Cross-source association between mean reference likelihood and downstream avg@4. Exact p -values enumerate all assignments of the observed utilities to sources; p− is the one-sided test for a negative association. Prompt z uses within-prompt standardized reference likelihood. Leave-one-out (LOO) reports the range of Pearson correlations after removing each source in turn.
Coupling ( −logS /token)
avg@4
Contrast
Reject source
Repelled
Native
Attracted
Repelled
Native
Attracted
R − N
A − N
R − A
Qwen2.5-0.5B-Instruct
0.690
0.519
0.359
63.335
63.075
62.750
+0.260
−0.325
+0.585
Qwen2.5-1.5B-Instruct
0.487
0.344
0.238
63.135
62.390
61.880
+0.745
−0.510
+1.255
Qwen2.5-3B-Instruct
0.492
0.363
0.260
63.120
62.860
62.655
+0.260
−0.205
+0.465
Qwen2.5-7B-Instruct
0.320
0.243
0.183
62.100
62.070
61.470
+0.030
−0.600
+0.630
Qwen2.5-14B-Instruct
0.363
0.263
0.191
59.905
59.415
59.455
+0.490
+0.040
+0.450
Appendix
Table 11 : Reference-likelihood reselection for the Qwen2.5-14B-Instruct student. Coupling entries are the negative mean length-normalized log likelihood under the 14B SeqKD reference, so larger values indicate lower reference coupling. Utility entries are KoDCode avg@4; the final three columns give the within-source contrasts. Native arms are trained on the reselection candidate banks, so their values differ from the natural source endpoints in Table 9 .
Reject source
Mean-field norm
Δ field
Coherence
Per-token norm
Δ per-token
14B student
post-SeqKD Self (14B)
1.559
–
0.171
9.47
–
0.5B
3.146
+1.587
0.171
16.85
+7.38
1.5B
2.119
+0.561
0.153
16.96
+7.49
3B
2.829
+1.270
0.199
17.69
+8.22
vanilla 7B
2.163
+0.604
0.166
13.31
+3.84
Appendix
Table 12 : Student-local reject-score fields at the shared preference- optimization initialization. Norms are averaged over two independent sketches. Differences are source minus post-SeqKD Self. Per-token norms and differences are scaled by 103 .
Figure 7 : Empirical rejected-response fields at the shared DPO initialization for the 14B and 72B students. Post-SeqKD Self has an attenuated field relative to smaller Base sources, while vanilla Self overlaps the smaller Base range.
Figure 8 : Differences between Self and smaller Base reject classes. The panels report score-field displacement from post-SeqKD Self and overlap in failed execution tests. These measurements characterize the source classes but do not rank their downstream utility.
Reject
Chosen Δℓ
Path III (%)
Regressions
post-SeqKD Self
−6.79
5.3
288
0.5B
−6.67
11.7
262
1.5B
−10.92
7.8
324
3B
−10.19
12.5
247
7B
−8.80
12.3
215
vanilla Self
−8.23
9.6
299
Appendix
Table 13 : Endpoint fixed-pair dynamics and KoDCode outcomes. Path III is the percentage of prompts that preserve the chosen response while suppressing the source’s own reject. Regressions count problems solved by the shared preference-optimization initialization but not by the final model. The lower panel reports descriptive correlations across the six source endpoints.
Figure 9 : Training diagnostics for the 14B student. Panels report the initialization reject field, final implicit reward margin, winner preservation on fixed pairs, and the effect of RC-DPO on chosen and rejected rewards. None consistently separates smaller Base rejects from both Self controls.
Reject source
Objective
Chosen Δℓ
Own-reject Δℓ
Path III (%)
RC effect on Path III
post-SeqKD Self
DPO
−6.787
−16.337
5.3
–
RC-DPO
−11.650
−22.545
1.0
−4.30
3B Base
DPO
−10.185
−63.464
12.5
–
RC-DPO
−5.890
−53.974
18.6
+6.05
vanilla Self
DPO
−8.230
−44.353
9.6
–
RC-DPO
−5.711
−39.027
12.9
+3.32
Appendix
Table 14 : Reject source × objective factorial on the fixed 512-prompt bank. Likelihood displacements are sequence sums relative to the shared SeqKD reference. RC effects and interactions are percentage points.
Reject source
Objective
KoDCode pass@1
BCB-C avg@4
BCB-C pass@4
post-SeqKD Self
DPO
63.32
45.70
65.35
RC-DPO
62.16
44.96
64.65
3B Base
DPO
64.34
51.12
65.44
RC-DPO
64.34
50.86
64.91
vanilla Self
DPO
61.96
49.06
63.95
RC-DPO
62.52
48.84
64.12
Appendix
Table 15 : Held-out performance under standard DPO and RC-DPO for the 14B student. Values are absolute percentages.
Figure 10 : Reward dynamics under standard DPO and RC-DPO for the 14B student. Panels show post-SeqKD Self, 3B Base, and vanilla Self rejects. Solid lines track chosen implicit reward and dashed lines track rejected implicit reward. RC-DPO changes the likelihood trajectory in a source-dependent manner.
Distilling the capabilities from a large reasoning model (LRM) to a smaller student model often involves training on substantial amounts of reasoning data. However, knowledge distillation (KD) over lengthy sequences with prompt (P), chain-of-thought (CoT), and answer (A) sections makes the process computationally expensive. In this work, we investigate how the allocation of supervision across different sections (P, CoT, A) affects student performance. Our analysis shows that selective KD over only the CoT tokens can be effective when the prompt and answer information is encompassed by it. Building on this insight, we establish a truncation protocol to quantify computation-quality tradeoffs as a function of sequence length. We observe that beyond a specific length, longer training sequences provide marginal returns for downstream performance but require substantially higher memory and FLOPs. To this end, training on only the first 50% of tokens of every training sequence can retain, on average, ≈91% of full-sequence performance on math benchmarks while reducing training time, memory usage, and FLOPs by about 50% each. Codes are available at https://github.com/weiruichen01/distilling-the-essence.
Wei-Rui Chen, Vignesh Kothapalli, Ata Fatahibaarzi +5
1The University of British Columbia · 2LinkedIn · 3Canada Research Chair in NLP and ML
On-policy distillation (OPD) trains a student language model with dense feedback from a stronger teacher on student-generated trajectories. Yet standard OPD weights token-level distillation terms uniformly, implicitly treating local teacher preference as a proxy for correction utility. A decision's task value, however, depends on how the student completes the subsequent reasoning. This mismatch can cause imitation to suppress viable student strategies or reinforce paths the student cannot reliably execute. Verified trajectory outcomes provide complementary evidence about continuation quality, but do not directly identify the utility of individual decisions. We introduce Reward-Aligned Reweighting for On-Policy Distillation (R2-OPD), which uses outcome agreement and the magnitude of teacher--student disagreement to continuously reallocate teacher supervision. It gives reward-aligned corrections greater relative influence while retaining dense feedback, moving beyond uniform imitation and hard filtering. Our analysis formalizes the mismatch between local teacher preference and student continuation value and establishes sufficient conditions for reallocation to improve first-order task progress over uniform OPD. Across seven mathematical reasoning benchmarks, R2-OPD achieves the highest average accuracy among the compared training methods in both cross-size and same-size distillation. It outperforms standard OPD on all seven benchmarks, with average gains of 3.5 and 2.4 percentage points for 1.7B and 4B students, respectively. An extension to code generation yields an average gain of 1.6 percentage points over standard OPD. These results highlight outcome-guided supervision allocation as an effective way to translate dense teacher feedback into stronger student performance across model scales and task domains.
Haofeng Xu, Junwei Su, Lansong Diao +2
The University of Hong Kong · Alibaba Group · University of Science and Technology of China
Converting a pretrained Transformer into a more efficient hybrid model through distillation offers a promising approach to reducing inference costs. However, achieving high-quality generation in distilled models requires careful joint design of both the student architecture and the distillation process. Many prior distillation works evaluate downstream multiple-choice benchmarks by ranking candidate answers with log-likelihood rather than requiring autoregressive generation, which can obscure important differences in model quality. For example, on overlapping benchmarks, we show that a 7B distilled model that nearly matches its teacher to within 0.2 pp under log-likelihood scoring falls behind by 20.8 pp when it must generate answers autoregressively. We investigate this phenomenon with GenDistill, a multi-stage pipeline we designed for distilling a pretrained Transformer into an efficient Hybrid Kimi Delta Attention (Hybrid-KDA) student. Using it as a controlled testbed on Qwen3-0.6B, we systematically ablate six design axes (training objective, loss masking, training duration, dataset selection, parameter freezing, and architecture choice) and evaluate every choice under both log-likelihood and generation-based protocols. We find that log-likelihood-based evaluation consistently underestimates the gap between teacher and student, and can in some cases reverse the ranking of design choices, so conclusions drawn from perplexity-only evaluation may be misleading. Among the factors we study, dataset selection, completion-only masking, and freezing attention layers during post-training have the largest impact on generation quality. Our best distillation recipe, using a Hybrid-KDA model as the student, retains 86-90% of teacher accuracy on knowledge benchmarks while reducing KV cache memory by up to 75% and improving time-to-first-token by 2-4x at 128K-token contexts.