Authors: Rui Cai, Wenhui Zhu, Xiwen Chen, Jincheng Cao, Han Yu, Shayan Mohajer Hamidi, Zelin He, Qiyao Ma, +10 more
Organizations: AI Agentic Modeling and Foundation Team, LinkedIn · University of California, Davis · Arizona State University · Clemson University · Pennsylvania State University · University of Wisconsin–Madison
Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student, making generation costly at scale. We find neither assumption holds: across students from 7B to 72B, smaller frozen models generate rejects with less inference compute yet train stronger students than self-generated rejects, before and after sequence-level knowledge distillation, on code generation and mathematical reasoning. To explain this result, we derive a finite-horizon utility bound for Direct Preference Optimization in a linearized feature model. The bound characterizes favorable reject distributions and motivates three interventions. First, mixing rejects from smaller and student-scale models improves performance as the smaller model's share increases. Second, reassigning rejects to other prompts and shuffling their code tokens still outperform length-matched gibberish, showing that task structure contributes to reject utility. Third, selecting candidates with lower likelihood under the reference policy improves net transfer when higher-likelihood candidates provide less useful contrast. Lower-likelihood selections outperform higher-likelihood ones for every source. These results suggest that effective rejects preserve task structure while limiting coupling to the reference policy, and that smaller frozen models can provide them at low cost.
Figures & tables
Figure 1 : Smaller Base models provide stronger rejects, and their advantage varies systematically with source composition. Left: KoDCode avg@4 for the original vanilla model, both Self constructions, and every smaller Base source available to each 7B–72B student. All 18 smaller Base configurations exceed both post-SeqKD Self and vanilla Self. Right: a randomized mixture intervention on a matched 14B population. Avg@4 rises from 59.73 to 63.28 as the share of Llama-1B Base rejects increases from 0% to 100% .
Figure 2 : Smaller Base models improve preference distillation across code generation and mathematical reasoning. (a, b) Absolute avg@4 for the original vanilla model, SeqKD, Continued-SFT, both Self constructions, and the best observed smaller Base source for each 7B–72B student. Labels above the green bars identify the displayed reject source. (c) Two randomized source assignment interventions vary the proportions of smaller Base and vanilla Self rejects for a 14B student. The green and gray background shows the dataset composition. Performance increases with the smaller Base share for both source families.
Figure 3 : Localizing and decomposing the utility of smaller Base rejects. (a) Absolute KoDCode avg@4 across four control settings at every student scale. Continued-SFT trains on the preference-stage chosen responses alone. Gibberish supplies length-matched artificial rejects. Prompt reassignment preserves real Base rejects while removing their original prompt correspondence, and Aligned restores that correspondence. (b) Lexical permutation preserves the generated code-token inventory of prompt-reassigned Base rejects while destroying token order, syntax, and coherent program semantics. Aligned shows the original prompt-matched Base construction.
Figure 4 : Reference likelihood tracks and controls reject utility. (a) Each point is a natural reject source scored by its mean length-normalized likelihood under the corresponding student’s SeqKD reference and standardized within prompt. Lines are fitted separately for each student scale; labels report Pearson correlations. Filled circles are smaller Base sources, squares are vanilla Self, and diamonds are post-SeqKD Self. (b) For a fixed Qwen2.5-14B-Instruct student, each source-specific candidate bank is reselected toward lower reference likelihood (Repelled), without likelihood preference (Native), or toward higher reference likelihood (Attracted).
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
SeqKD
DPO
Learning rate
5×10−6
5×10−7
DPO β
–
0.1
Epochs
1
1
Gradient accumulation
16
16
Effective batch (7B–72B)
256
256
Max sequence length
2,560
2,560
Appendix
Table 1: Training hyperparameters of the code line. All runs use full-parameter training with AdamW, a cosine learning rate schedule with 10% linear warmup, bf16 precision, gradient checkpointing, DeepSpeed ZeRO-3. Continued-SFT uses the same settings as DPO.
Student
Smaller Base source
Parameter ratio
Reject tokens
Decode cost
Total cost
Reduction
7B
0.5B
6.5%
11.317M
10.6%
8.2%
12.2×
7B
1.5B
20.3%
6.366M
18.7%
19.6%
5.1×
7B
3B
40.5%
7.074M
41.5%
40.9%
2.4×
14B
0.5B
3.3%
11.317M
5.1%
4.1%
24.3×
14B
1.5B
10.5%
6.366M
9.0%
9.8%
10.2×
14B
3B
20.9%
7.074M
20.0%
20.5%
4.9×
Appendix
Table 2 : Complete output-length-adjusted parameter-token cost comparison. Decode and total costs are percentages of same-scale vanilla Self. The reduction factor is the inverse of the total-cost ratio.
pass@1
Stochastic
Method
Reject Source
Overall
Easy
Medium
Hard
avg@4
pass@4
Student: Qwen2.5-7B-Instruct
Base
–
48.9
62.3
45.3
40.2
47.9
59.8
SeqKD
–
58.8
70.4
56.2
50.7
55.6
69.4
Continued-SFT
–
59.6
71.6
57.0
52.0
56.3
70.0
DPO
0.5B
60.5
72.0
57.5
53.1
58.5
70.5
Appendix
Table 3 : KoDCode held-out results. pass@1 uses greedy decoding; avg@4 is the mean success rate over four stochastic samples, and pass@4 is the fraction of problems solved by at least one sample. Stochastic decoding uses temperature 0.6 and top- p=1.0 . All values are percentages. Bold marks the best result within each student and metric.
Method
Reject Source
BCB-C
BCB-I
HumanEval
MBPP
Weighted Avg.
Student: Qwen2.5-7B-Instruct
Base
–
42.3
29.1
84.1
79.0
42.8
SeqKD
–
43.4
32.5
83.3
79.1
44.6
Continued-SFT
–
43.9
33.6
82.3
79.1
45.2
DPO
0.5B
45.2
35.4
84.8
78.2
46.6
1.5B
45.0
33.3
82.3
80.2
45.7
Appendix
Table 4 : External pass@1 across four code-generation benchmarks. Weighted Avg. weights each benchmark by its number of problems: 1,140 each for BCB-C and BCB-I, 164 for HumanEval, and 257 for MBPP. All values are percentages. Bold marks the best result within each student and column.
Method
Reject Source
MATH-500 avg@4
AIME avg@4
Teacher (DeepSeek-R1)
–
97.3
74.9
Student: Qwen2.5-72B-Instruct
Base
–
84.0
15.8
SeqKD
–
91.0
37.9
Continued-SFT
–
91.2
37.5
DPO
0.5B
91.6
40.0
Appendix
Table 5 : MATH-500 and AIME avg@4 over four samples.
Figure 5 : Cross-family reject sources. Bars report avg@4 on (a) KoDCode in-domain and (b) MATH-500 for the 14B, 32B, and 72B Qwen2.5 students. Self is the better of vanilla Self and post-SeqKD Self. Labels on the green bars identify the best Qwen2.5 smaller Base source; labels on the purple bars identify the Llama source.
Student
Smaller Base source
Base share
Code avg@4
Code pass@4
14B
Llama-3.2-1B-Instruct
0%
59.730
72.380
25%
61.575
73.000
50%
62.035
72.760
75%
62.670
73.380
100%
63.280
73.180
14B
Qwen2.5-3B-Instruct
0%
59.575
72.280
Appendix
Table 6 : Complete randomized source assignment results. The Base share is the proportion of prompts assigned rejects from the indicated smaller Base source. All remaining prompts use vanilla Self rejects from the student-scale model.
Student
Reject source
Continued-SFT
Gibberish
Prompt-reassigned
Aligned
7B
Qwen2.5-0.5B-Instruct
56.30
57.37
58.65
58.53
14B
Qwen2.5-3B-Instruct
60.89
60.92
61.97
62.63
32B
Qwen2.5-7B-Instruct
64.50
64.52
65.25
65.53
72B
Qwen2.5-7B-Instruct
65.00
65.24
66.36
67.53
Appendix
Table 7 : KoDCode avg@4 across the reject-control ladder. Prompt-reassigned arms use real Base rejects reassigned across prompts. Each row is one completed prompt-reassignment experiment and reports its Continued-SFT, gibberish, reassigned, and aligned settings.
Figure 6 : Implicit reward dynamics for gibberish, prompt reassigned, and prompt-matched Base rejects at each student scale. Gibberish is suppressed most strongly but does not recover the utility of real Base rejects.
Student
Gibberish
Lexical
Real reassigned
Lexical − Gibberish
Recovery
7B
57.370
57.900
58.645
+0.530
41.6%
14B
60.920
61.650
61.970
+0.730
69.5%
32B
64.520
65.195
65.250
+0.675
92.5%
72B
65.235
65.685
66.360
+0.450
40.0%
Appendix
Table 8 : Cross-scale lexical-structure interventions. Recovery is the Lexical gain over Gibberish divided by the real prompt-reassigned gain over Gibberish.
Student
Reject source
Source class
Mean log S /token
Prompt z
Mean tokens
avg@4
7B
Qwen2.5-0.5B-Instruct
Subscale Base
−0.3725
−1.048
466.7
58.525
7B
Qwen2.5-1.5B-Instruct
Subscale Base
−0.2527
−0.113
262.5
57.395
7B
Qwen2.5-3B-Instruct
Subscale Base
−0.2851
−0.435
291.8
57.705
7B
Qwen2.5-7B-Instruct
Vanilla Self
−0.1790
+0.418
284.9
54.550
7B
post-SeqKD 7B
Post-SeqKD Self
−0.0915
+1.179
243.7
52.745
14B
Qwen2.5-0.5B-Instruct
Subscale Base
−0.3848
−1.074
466.7
62.525
Appendix
Table 9 : Natural reject-source measurements underlying Figure 4 a. Each source is scored on the same 24,248 preference prompts under the corresponding student’s SeqKD reference. Mean log likelihood is length normalized. Prompt z is the mean reference likelihood after standardization within each prompt. Utility is KoDCode avg@4 on 5,000 held-out problems.
Student
n
r
ρ
p−
p2
Prompt- zr
LOO r range
7B
5
−0.967
−1.000
0.0083
0.0083
−0.968
[−0.988,−0.920]
14B
6
−0.868
−0.886
0.0083
0.0083
−0.871
[−0.909,−0.645]
32B
7
−0.752
−0.464
0.0298
0.0484
−0.770
[−0.863,−0.405]
72B
8
−0.836
−0.690
0.0057
0.0076
−0.842
[−0.931,−0.684]
Appendix
Table 10 : Cross-source association between mean reference likelihood and downstream avg@4. Exact p -values enumerate all assignments of the observed utilities to sources; p− is the one-sided test for a negative association. Prompt z uses within-prompt standardized reference likelihood. Leave-one-out (LOO) reports the range of Pearson correlations after removing each source in turn.
Coupling ( −logS /token)
avg@4
Contrast
Reject source
Repelled
Native
Attracted
Repelled
Native
Attracted
R − N
A − N
R − A
Qwen2.5-0.5B-Instruct
0.690
0.519
0.359
63.335
63.075
62.750
+0.260
−0.325
+0.585
Qwen2.5-1.5B-Instruct
0.487
0.344
0.238
63.135
62.390
61.880
+0.745
−0.510
+1.255
Qwen2.5-3B-Instruct
0.492
0.363
0.260
63.120
62.860
62.655
+0.260
−0.205
+0.465
Qwen2.5-7B-Instruct
0.320
0.243
0.183
62.100
62.070
61.470
+0.030
−0.600
+0.630
Qwen2.5-14B-Instruct
0.363
0.263
0.191
59.905
59.415
59.455
+0.490
+0.040
+0.450
Appendix
Table 11 : Reference-likelihood reselection for the Qwen2.5-14B-Instruct student. Coupling entries are the negative mean length-normalized log likelihood under the 14B SeqKD reference, so larger values indicate lower reference coupling. Utility entries are KoDCode avg@4; the final three columns give the within-source contrasts. Native arms are trained on the reselection candidate banks, so their values differ from the natural source endpoints in Table 9 .
Reject source
Mean-field norm
Δ field
Coherence
Per-token norm
Δ per-token
14B student
post-SeqKD Self (14B)
1.559
–
0.171
9.47
–
0.5B
3.146
+1.587
0.171
16.85
+7.38
1.5B
2.119
+0.561
0.153
16.96
+7.49
3B
2.829
+1.270
0.199
17.69
+8.22
vanilla 7B
2.163
+0.604
0.166
13.31
+3.84
Appendix
Table 12 : Student-local reject-score fields at the shared preference- optimization initialization. Norms are averaged over two independent sketches. Differences are source minus post-SeqKD Self. Per-token norms and differences are scaled by 103 .
Figure 7 : Empirical rejected-response fields at the shared DPO initialization for the 14B and 72B students. Post-SeqKD Self has an attenuated field relative to smaller Base sources, while vanilla Self overlaps the smaller Base range.
Figure 8 : Differences between Self and smaller Base reject classes. The panels report score-field displacement from post-SeqKD Self and overlap in failed execution tests. These measurements characterize the source classes but do not rank their downstream utility.
Reject
Chosen Δℓ
Path III (%)
Regressions
post-SeqKD Self
−6.79
5.3
288
0.5B
−6.67
11.7
262
1.5B
−10.92
7.8
324
3B
−10.19
12.5
247
7B
−8.80
12.3
215
vanilla Self
−8.23
9.6
299
Appendix
Table 13 : Endpoint fixed-pair dynamics and KoDCode outcomes. Path III is the percentage of prompts that preserve the chosen response while suppressing the source’s own reject. Regressions count problems solved by the shared preference-optimization initialization but not by the final model. The lower panel reports descriptive correlations across the six source endpoints.
Figure 9 : Training diagnostics for the 14B student. Panels report the initialization reject field, final implicit reward margin, winner preservation on fixed pairs, and the effect of RC-DPO on chosen and rejected rewards. None consistently separates smaller Base rejects from both Self controls.
Reject source
Objective
Chosen Δℓ
Own-reject Δℓ
Path III (%)
RC effect on Path III
post-SeqKD Self
DPO
−6.787
−16.337
5.3
–
RC-DPO
−11.650
−22.545
1.0
−4.30
3B Base
DPO
−10.185
−63.464
12.5
–
RC-DPO
−5.890
−53.974
18.6
+6.05
vanilla Self
DPO
−8.230
−44.353
9.6
–
RC-DPO
−5.711
−39.027
12.9
+3.32
Appendix
Table 14 : Reject source × objective factorial on the fixed 512-prompt bank. Likelihood displacements are sequence sums relative to the shared SeqKD reference. RC effects and interactions are percentage points.
Reject source
Objective
KoDCode pass@1
BCB-C avg@4
BCB-C pass@4
post-SeqKD Self
DPO
63.32
45.70
65.35
RC-DPO
62.16
44.96
64.65
3B Base
DPO
64.34
51.12
65.44
RC-DPO
64.34
50.86
64.91
vanilla Self
DPO
61.96
49.06
63.95
RC-DPO
62.52
48.84
64.12
Appendix
Table 15 : Held-out performance under standard DPO and RC-DPO for the 14B student. Values are absolute percentages.
Figure 10 : Reward dynamics under standard DPO and RC-DPO for the 14B student. Panels show post-SeqKD Self, 3B Base, and vanilla Self rejects. Solid lines track chosen implicit reward and dashed lines track rejected implicit reward. RC-DPO changes the likelihood trajectory in a source-dependent manner.