Large language models (LLMs) typically generate text autoregressively (AR), predicting one token at a time. Block diffusion language models (dLLMs) instead generate blocks sequentially while denoising multiple tokens in parallel within each block, offering a promising way to accelerate generation. Rather than training such models from scratch, recent work adapts strong pretrained AR models into block dLLMs through distillation. On-policy distillation (OPD) has been widely used for LLM training because it supervises the student on states generated by its current policy, rather than only on fixed offline trajectories. By training on the states the student actually visits, it reduces the mismatch between training and generation and can provide more relevant supervision as the student evolves. Recent work has extended this idea to AR-to-block-diffusion conversion. However, this setting introduces a fundamental mismatch in supervision: the block-diffusion student and the causal AR teacher condition on different information at the same training state. The student predicts from the entire partially denoised block, including visible future context, whereas the standard AR teacher target is defined only from the causal prefix. As a result, the teacher distribution used for distillation is not fully aligned with the information available to the student. We therefore introduce d-OPD, a future-aware on-policy distillation method that corrects the AR teacher distribution to better align with the student-visible state by incorporating visible future information within each block, providing supervision that better matches the information used by the student. Across Qwen3 models from 0.6B to 8B, d-OPD improves the six-benchmark average by up to 4.0 points over OPDLM and reduces training time by 1.35-1.58×. The code is available at https://github.com/mit-han-lab/d-OPD.
Figures & tables
Setting
Method
MMLU
MMLU-P
GPQA-D
GSM8K
MATH500
AIME25
Avg.
Qwen3-8B: comparison across block sizes
Qwen3-8B
76.1
60.5
44.4
91.7
83.2
16.7
62.1
N=4
SFT
69.1
54.3
39.9
87.3
75.4
13.3
56.6
B-SFT
70.9
52.6
39.4
88.5
72.6
16.7
56.8
BARD
68.6
55.3
38.4
88.6
76.4
16.7
57.3
OPDLM
69.2
54.0
38.4
87.2
76.8
10.0
55.9
Table 1: Main results. Accuracy across block sizes and Qwen3 model scales on six benchmarks. d-OPD consistently outperforms OPDLM across all evaluated model scales, with improvements remaining stable from 0.6B to 8B. The gains also persist as the block size increases, reaching up to 4.0 points in the six-benchmark average. Best results are in bold. Second-best results are underlined.
Model scale
Qwen3-0.6B
Qwen3-1.7B
Qwen3-4B
Qwen3-8B
OPDLM Time (hr)
18.39
19.92
30.72
28.26
d-OPD Time (hr)
13.15
14.75
22.34
17.91
Speedup
1.40 ×
1.35 ×
1.38 ×
1.58 ×
Table 2: Training efficiency. Wall-clock time to the best six-benchmark average achieved by OPDLM at each model scale. d-OPD consistently reaches the matched OPDLM performance faster, with speedups of 1.35 – 1.58× across Qwen3 models from 0.6B to 8B.
Mass (%)
k=4
k=8
k=16
k=32
k=64
Teacher
97.87
98.88
99.45
99.67
99.79
Student
95.43
98.30
99.44
99.64
99.77
Overhead
2.44%
2.57%
2.63%
3.50%
6.02%
Table 3: Ablation study. Left: Effect of candidate-set size k , showing the probability mass covered by the selected candidates and the resulting training overhead. Even small candidate sets cover nearly all teacher and student probability mass, while k=16 provides a favorable trade-off between coverage and overhead. Right: Effect of correctness gating. The gate improves the six-benchmark average from 58.8 to 59.8 by restricting future-aware correction to verifier-correct rollouts.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Target
KL
Reduction
Causal teacher
0.4444
–
Future-aware
0.1249
71.9%
Appendix
Table 4: Mean KL divergence to the exactly enumerated future-conditioned target, averaged over 49,948 masked-token marginals across all controlled configurations. Lower is better.
Method
Gate
MMLU
MMLU-P
GPQA-D
GSM8K
MATH500
AIME25
Avg.
OPDLM
–
69.2
54.0
38.4
87.2
76.8
10.0
55.9
d-OPD
Off
74.2
56.9
37.4
87.9
76.2
20.0
58.8
d-OPD
On
74.3
55.9
38.9
89.3
77.2
23.3
59.8
Appendix
Table 5: Full results for the correctness-gating ablation on Qwen3-8B. The gate applies future-aware correction only to verifier-correct rollouts, while ordinary causal teacher supervision remains active for all rollouts. Avg. denotes the unweighted average over the six benchmarks. Higher is better.
On-policy self-distillation (OPSD) has proven effective for post-training large language models (LLMs), yet its application to diffusion LLMs (dLLMs) remains unexplored. Existing OPSD methods are inherently autoregressive-centric. They inject privileged information via left-to-right prefix conditioning with token-level divergence supervision, a design that fundamentally conflicts with the arbitraryorder generation of dLLMs. We introduce d-OPSD, the first OPSD framework tailored for dLLMs. Our approach makes two core contributions. First, we reframe self-teacher construction by using self-generated answers as suffix conditioning, enabling the student model to learn from "self future-experience" rather than privileged prefixes. Second, we shift supervision from token-level to step-level, aligning training with the iterative denoising process of dLLMs. Experiments across four reasoning benchmarks show that d-OPSD consistently outperforms RLVR and SFT baselines with superior sample efficiency, requiring only around 10% of the optimization steps by RLVR and opening a promising pathway for dLLM posttraining. The code is available at https://github.com/xingzhejun/d-OPSD.
Yifu Luo, Zeyu Chen, Haoyu Wang +4
Tsinghua University · Technical University of Munich · Nanyang Technological University +5
Diffusion large language models (dLLMs) generate text by iteratively denoising a masked sequence, offering a parallel alternative to autoregressive models, but eliciting strong reasoning through post-training remains difficult: supervised fine-tuning is off-policy and suffers from exposure bias, while reinforcement learning gives only sparse, sequence-level rewards and is hard to apply without tractable sequence likelihoods. On-policy self-distillation (OPSD) offers a promising alternative, using one model as both student and teacher to provide dense, token-level, on-policy supervision, but its effectiveness hinges on giving the teacher privileged information (PI) - typically an instance-specific ground-truth reference unavailable at inference - so the student ends up distilling a weak PI-free consensus policy that yields little improvement on dLLM reasoning. We introduce dOPSD, which instead derives the teacher's privilege directly from the student's own denoising trajectory, evaluating masked positions using later, more-decoded steps of that same trajectory rather than an external label, so the teacher's advantage emerges from the model's own decoding process; on Dream and LLaDA, dOPSD improves both in-domain math reasoning and out-of-domain code generation, outperforming supervised and on-policy baselines.
Diffusion language models (dLLMs) can predict many tokens in parallel, but accurate generation still requires many iterative denoising steps. Few-step distillation accelerates decoding by compressing multiple teacher steps into a single student transition. However, existing methods construct supervision on off-policy trajectories. At inference, the student's early parallel commitments alter the context of later predictions, so the states it actually visits drift away from the supervised ones--precisely when step compression is most aggressive. On-policy distillation is a natural remedy for this mismatch, but it leaves open how far each transition should advance: matching only the teacher's next action limits compression, while indiscriminately merging future actions can violate intermediate dependencies. To address this limitation, we propose OPTD, On-Policy Transition Distillation with consistency-guided adaptive compression. It samples partial states from the few-step student's own trajectories, uses a frozen, question-only teacher to identify outcome-aligned future candidates, and orders them by current-state confidence. The method then selects the longest prefix whose joint commitment preserves the teacher's rollout outcome. A set-bottleneck objective promotes every verified future candidate to the decoder's release threshold, while a frozen-teacher KL anchor regularizes all other active positions. Neither target construction nor training uses a gold response. Across four mathematical reasoning and code-generation benchmarks, OPTD consistently improves the quality--efficiency trade-off and attains the strongest overall quality-constrained AUP among the evaluated few-step baselines.