Large language models (LLMs) typically generate text autoregressively (AR), predicting one token at a time. Block diffusion language models (dLLMs) instead generate blocks sequentially while denoising multiple tokens in parallel within each block, offering a promising way to accelerate generation. Rather than training such models from scratch, recent work adapts strong pretrained AR models into block dLLMs through distillation. On-policy distillation (OPD) has been widely used for LLM training because it supervises the student on states generated by its current policy, rather than only on fixed offline trajectories. By training on the states the student actually visits, it reduces the mismatch between training and generation and can provide more relevant supervision as the student evolves. Recent work has extended this idea to AR-to-block-diffusion conversion. However, this setting introduces a fundamental mismatch in supervision: the block-diffusion student and the causal AR teacher condition on different information at the same training state. The student predicts from the entire partially denoised block, including visible future context, whereas the standard AR teacher target is defined only from the causal prefix. As a result, the teacher distribution used for distillation is not fully aligned with the information available to the student. We therefore introduce d-OPD, a future-aware on-policy distillation method that corrects the AR teacher distribution to better align with the student-visible state by incorporating visible future information within each block, providing supervision that better matches the information used by the student. Across Qwen3 models from 0.6B to 8B, d-OPD improves the six-benchmark average by up to 4.0 points over OPDLM and reduces training time by 1.35-1.58×. The code is available at https://github.com/mit-han-lab/d-OPD.
Figures & tables
Setting
Method
MMLU
MMLU-P
GPQA-D
GSM8K
MATH500
AIME25
Avg.
Qwen3-8B: comparison across block sizes
Qwen3-8B
76.1
60.5
44.4
91.7
83.2
16.7
62.1
N=4
SFT
69.1
54.3
39.9
87.3
75.4
13.3
56.6
B-SFT
70.9
52.6
39.4
88.5
72.6
16.7
56.8
BARD
68.6
55.3
38.4
88.6
76.4
16.7
57.3
OPDLM
69.2
54.0
38.4
87.2
76.8
10.0
55.9
Table 1: Main results. Accuracy across block sizes and Qwen3 model scales on six benchmarks. d-OPD consistently outperforms OPDLM across all evaluated model scales, with improvements remaining stable from 0.6B to 8B. The gains also persist as the block size increases, reaching up to 4.0 points in the six-benchmark average. Best results are in bold. Second-best results are underlined.
Model scale
Qwen3-0.6B
Qwen3-1.7B
Qwen3-4B
Qwen3-8B
OPDLM Time (hr)
18.39
19.92
30.72
28.26
d-OPD Time (hr)
13.15
14.75
22.34
17.91
Speedup
1.40 ×
1.35 ×
1.38 ×
1.58 ×
Table 2: Training efficiency. Wall-clock time to the best six-benchmark average achieved by OPDLM at each model scale. d-OPD consistently reaches the matched OPDLM performance faster, with speedups of 1.35 – 1.58× across Qwen3 models from 0.6B to 8B.
Mass (%)
k=4
k=8
k=16
k=32
k=64
Teacher
97.87
98.88
99.45
99.67
99.79
Student
95.43
98.30
99.44
99.64
99.77
Overhead
2.44%
2.57%
2.63%
3.50%
6.02%
Table 3: Ablation study. Left: Effect of candidate-set size k , showing the probability mass covered by the selected candidates and the resulting training overhead. Even small candidate sets cover nearly all teacher and student probability mass, while k=16 provides a favorable trade-off between coverage and overhead. Right: Effect of correctness gating. The gate improves the six-benchmark average from 58.8 to 59.8 by restricting future-aware correction to verifier-correct rollouts.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Target
KL
Reduction
Causal teacher
0.4444
–
Future-aware
0.1249
71.9%
Appendix
Table 4: Mean KL divergence to the exactly enumerated future-conditioned target, averaged over 49,948 masked-token marginals across all controlled configurations. Lower is better.
Method
Gate
MMLU
MMLU-P
GPQA-D
GSM8K
MATH500
AIME25
Avg.
OPDLM
–
69.2
54.0
38.4
87.2
76.8
10.0
55.9
d-OPD
Off
74.2
56.9
37.4
87.9
76.2
20.0
58.8
d-OPD
On
74.3
55.9
38.9
89.3
77.2
23.3
59.8
Appendix
Table 5: Full results for the correctness-gating ablation on Qwen3-8B. The gate applies future-aware correction only to verifier-correct rollouts, while ordinary causal teacher supervision remains active for all rollouts. Avg. denotes the unweighted average over the six benchmarks. Higher is better.