Diffusion large language models (dLLMs) generate text by iterative unmasking. At each denoising step, a dLLM proposes a token at every masked position, but the decoder commits only a confident subset of these proposals. Trace-based on-policy distillation (TOPD) builds on this process by matching the student to a stronger teacher, yet only at the committed positions. We argue that this discards much of the useful signal, which resides in the uncommitted proposals, where the student has made a prediction but is not yet confident enough to commit it. We call these proposals hesitations. In our pilot study on an SDAR-4B student, hesitations make up only 24% of supervisable state-position pairs but carry 66% of the teacher-student divergence. To exploit this signal, we propose Hesitation-Aware On-Policy Distillation (HOPD), which extends teacher distribution matching to every masked position of each denoising step. Because hesitations are not equally informative, we further allocate supervision using hindsight from the completed trajectory, placing more weight on positions whose proposal was later disagreed with the final token and on blocks where first-step proposals rarely survive. Since both models already produce distributions at all masked positions, HOPD requires no additional forward passes over TOPD. The only extra cost is evaluating the loss at more positions. With SDAR-1.7B and SDAR-4B students distilled from TraDo-8B-Instruct, HOPD achieves the best average score among the evaluated methods on five math and coding benchmarks, under both static and dynamic decoding and at both scales. It also speeds up decoding. On SDAR-4B, the HOPD student hesitates less and commits 11% more tokens per denoising step than TOPD, while reaching higher accuracy.
Figures & tables
Figure 1: Hesitation pilot study on the rollouts of the base SDAR-4B student under dynamic decoding, with TraDo-8B-Instruct as the teacher. Left : share of all supervisable state–position pairs (light blue) and share of the total teacher–student reverse KL (dark blue) in each commitment category. Middle : mean per-position KL, averaged over each position’s recorded steps and grouped by the retraction count ej , i.e., the number of uncommitted proposals at position j that differ from its final token. Right : mean per-block KL, averaged over masked positions within each step and then over steps, grouped by the first-step mismatch count eb(0) , i.e., the number of positions in block b whose first proposal differs from the final token; only complete four-position blocks are included. Error bars in the middle and right panels are 95% confidence intervals.
Figure 2: Overview of HOPD. Left : Dynamic decoding of a four-token block ( τ=0.9 ). Blue proposals are committed, while yellow proposals are hesitations whose positions remain masked. Right : At each visited pre-action state, the student matches the frozen teacher’s distributions at all masked positions, including hesitations. Position and block weights are computed from the completed trajectory’s retraction counts and first-step retention, respectively.
Model
MATH500
AIME2024
GSM8K
LiveCodeBench-v2
LiveBench
Avg.
Static
Dynamic
Static
Dynamic
Static
Dynamic
Static
Dynamic
Static
Dynamic
Static
Dynamic
TraDo-8B-Instruct (teacher)
78.5
75.5
13.3
11.0
92.3
91.2
25.9
22.4
22.7
20.6
46.5
44.1
SDAR-1.7B-Chat
61.7
54.2
4.0
4.7
80.9
78.4
7.2
4.2
5.2
4.2
31.8
29.1
+ SFT
49.2
46.3
1.0
3.8
77.0
74.7
7.8
6.6
6.2
8.6
28.2
28.0
+ Off-policy
63.3
56.6
6.8
3.2
81.6
75.1
11.3
9.1
11.7
7.8
34.9
30.4
+ TraceRL
62.5
56.1
6.7
3.5
81.1
78.7
9.4
4.7
10.7
5.7
34.1
29.7
Table 1: The main benchmark results across different math and coding tasks. “Static” refers to static sampling, and “Dynamic” refers to dynamic sampling.
Figure 3: Accuracy against committed tokens per denoising step on MATH500 (dynamic decoding) as the commit threshold τ goes from 0.5 to 0.99 , right to left along each curve; large markers denote the training threshold τ=0.9 . Left : SDAR-4B. Right : SDAR-1.7B.
Base
TOPD
HOPD
Proportion of pairs
Committed
75.2%
72.6%
77.1%
Deferred
11.9%
12.0%
10.5%
Retracted
12.9%
15.4%
12.4%
Predictive entropy (nats)
Committed
0.141
0.177
0.124
Table 2: Commitment categories after training.
Model
MATH500
AIME2024
GSM8K
LiveCodeBench-v2
LiveBench
Avg.
Static
Dynamic
Static
Dynamic
Static
Dynamic
Static
Dynamic
Static
Dynamic
Static
Dynamic
LLaDA-8B-Instruct
36.8
37.4
0.0
0.0
81.8
82.2
6.5
6.6
7.3
7.8
26.5
26.8
+ ESPO (teacher)
37.4
36.0
0.0
0.0
83.3
82.8
8.4
7.3
12.8
12.0
28.4
27.6
+ TOPD
38.2
38.3
0.3
0.3
83.8
82.2
6.8
6.1
10.2
9.4
27.9
27.3
+ HOPD (ours)
39.0
38.1
0.2
0.2
82.2
82.4
7.7
6.1
13.0
10.7
28.4
27.5
Table 3: Full-attention model: LLaDA-8B-Instruct with ESPO-trained teachers.
Static
Dynamic
Per-position loss
MATH500
AIME2024
GSM8K
Avg.
MATH500
AIME2024
GSM8K
Avg.
Full forward KL
75.6
9.8
91.1
58.8
71.1
8.8
90.3
56.7
Full JSD ( β=0.5 )
75.9
8.2
91.5
58.5
71.8
8.5
90.6
57.0
Sampled-token reverse KL
75.3
7.8
91.7
58.3
71.9
8.3
91.0
57.1
Full reverse KL
74.9
15.0
91.1
60.3
72.5
9.7
90.7
57.6
Table 4: Comparison of divergence objectives on math with SDAR-4B-Chat.
Static
Dynamic
Weights
MATH500
AIME2024
GSM8K
Avg.
MATH500
AIME2024
GSM8K
Avg.
Uniform ( wj=wb=1 )
75.5
11.2
91.2
59.3
72.5
9.7
90.2
57.5
wj only ( wb=1 )
75.3
8.5
91.5
58.4
71.5
8.5
90.8
56.9
wb only ( wj=1 )
75.5
7.8
92.0
58.4
72.3
10.0
90.5
57.6
wj and wb (HOPD)
74.9
15.0
91.1
60.3
72.5
9.7
90.7
57.6
Table 5: Comparison of hindsight weights on math. Bold: best per column.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Unit
Group
Units
Prompts
Mean KL
95% CI
Position
ej=0
141,254
256
0.0924
[0.0872,0.0979]
ej=1
14,588
256
0.5044
[0.4857,0.5245]
ej=2
4,545
255
0.6084
[0.5845,0.6341]
ej≥3
1,114
223
0.7196
[0.6786,0.7655]
Block
eb(0)=0
29,416
256
0.0672
[0.0629,0.0718]
eb(0)=1
5,718
256
0.3824
[0.3612,0.4068]
Appendix
Table 6: Statistics behind the middle and right panels of Figure 1 . Each unit is a position in the middle panel and a complete four-position block in the right panel.
Setting
Value
Models
Student
SDAR-4B-Chat or SDAR-1.7B-Chat ( Cheng et al., 2026 )
Teacher
TraDo-8B-Instruct ( Wang et al., 2026c ) , frozen
Precision
bfloat16 weights, bf16 mixed precision
Gradient checkpointing
enabled
Parallelism
DeepSpeed ZeRO-2, no offload; 4 GPUs (2 GPUs with accumulation for the 1.7B code run)
Appendix
Table 7: Training configuration shared by the TOPD and HOPD runs of Table 1
Method
Training states
Signal
Supervised positions
Per-position loss
SFT
fixed teacher responses under random block masks
teacher tokens
masked positions of each block
cross entropy
Off-policy
frozen teacher trajectories
teacher distribution
committed positions of the teacher’s trace
full-vocabulary reverse KL
TraceRL
student rollouts
verifiable reward
committed positions along the trace
clipped policy gradient with KL penalty
TOPD
student rollouts
teacher distribution
committed positions At
sampled-token reverse-KL estimator
HOPD
student rollouts
teacher distribution
all masked positions M(st)
full-vocabulary reverse KL with hindsight weights
Appendix
Table 8: What each baseline method of Table 1 trains on. States: inputs on which the loss is evaluated. Positions: response positions that receive gradient at each state.
Setting
SDAR-4B / 1.7B-Chat
LLaDA-8B-Instruct
Max new tokens
2000
512 (math), 1024 (code)
Block size B / steps per block K
4 / 4
32 / 32
Future horizon
–
128
Temperature
1.0
0.1
Top- p
1.0
–
Top- k (static / dynamic)
1 / –
– / –
Appendix
Table 9: Inference settings. SDAR settings follow Wang et al. (2026c) ; LLaDA settings follow Ren et al. (2026) , except the code budget, which is the dLLM-RL default. Static and dynamic decoding differ only in the commit rule and, on SDAR, in top- k . With K=B , static decoding commits one token per step, and dynamic decoding needs at most as many steps. On LLaDA, each proposal is the argmax of Gumbel-perturbed logits at the listed temperature, with no nucleus or top- k truncation. “–” : not used.
Figure 4: Accuracy against wall-clock generation throughput on MATH500 (dynamic decoding), for the threshold sweep of Figure 3 . Left : SDAR-4B. Right : SDAR-1.7B.
Method
MATH500
AIME2024
GSM8K
LCB-v2
LiveBench
Avg.
SDAR-1.7B-Chat
TOPD (sampled-token)
64.5/58.1
2.2/3.2
81.9/78.2
10.4/7.7
9.4/6.2
33.7/30.7
TOPD (full reverse KL)
64.4/57.7
4.3/3.7
81.3/76.7
10.2/8.6
7.6/7.6
33.6/30.9
HOPD
66.8/59.5
3.7/5.0
83.6/78.8
11.1/8.9
10.9/9.9
35.2 / 32.4
SDAR-4B-Chat
TOPD (sampled-token)
74.9/71.3
11.2/7.2
92.0/90.0
21.9/18.6
19.8/17.4
43.9/40.9
Appendix
Table 10: Comparison with full-vocabulary reverse-KL TOPD. Each entry reports static/dynamic scores; Avg. is the unweighted mean over the five benchmarks. Sampled-token TOPD and HOPD results are repeated from Table 1 .
Setting
Math
Code
Models and data
Student
LLaDA-8B-Instruct
Frozen teacher
ESPO-Math
ESPO-Code
Training dataset
MATH (8,523 problems)
PrimeIntellect (5,954 problems)
Max prompt length
400 tokens
512 tokens
Rollout (on-policy sampling)
Appendix
Table 11: Experimental configuration for TOPD and HOPD on LLaDA-8B-Instruct.
Run
4
8
12
16
20
24
30
TOPD, math
38.1/37.8
37.3/37.1
38.2 /37.6
38.0/37.5
37.5/ 38.3
37.1/37.2
37.0/37.3
HOPD, math
37.6/37.1
37.0/36.7
37.6/36.0
38.5/37.0
38.2/37.5
37.7/37.2
39.0 / 38.1
TOPD, code
10.2 / 9.4
9.6/8.1
10.2/9.3
8.3/8.6
6.0/7.8
7.8/9.4
6.8/6.2
HOPD, code
9.4/5.5
10.9/9.4
11.2/9.4
9.9/6.2
13.0 / 10.7
9.6/9.9
10.9/7.0
Appendix
Table 12: Selection-benchmark accuracy (static/dynamic) of the LLaDA-8B runs of Table 3 every fourth round: MATH500 for math runs, LiveBench for code runs. Bold: selected round per decoder.
Static
Dynamic
Weights
MATH500
AIME2024
GSM8K
Avg.
MATH500
AIME2024
GSM8K
Avg.
uniform ( wj=wb=1 )
64.1
7.0
81.8
51.0
58.4
4.2
77.6
46.7
wj only
62.2
2.0
82.7
49.0
57.1
4.0
77.6
46.2
wb only
65.4
4.8
81.7
50.6
58.6
3.8
77.0
46.5
wj and wb (HOPD)
66.8
3.7
83.6
51.4
59.5
5.0
78.8
47.8
Appendix
Table 13: Hindsight weights on math for SDAR-1.7B-Chat, companion of Table 5 ; same protocol. Bold: best per column.
Student
Weights
5
10
15
20
25
30
SDAR-4B
uniform
74.4/70.6
74.8/ 72.5
75.5 /71.2
73.1/70.7
75.4/71.7
74.6/71.4
wj only
73.6/70.7
75.3 /70.1
75.1/70.3
74.0/71.0
74.3/70.9
73.9/ 71.5
wb only
73.1/71.5
72.8/71.8
75.5 /71.3
74.3/71.5
74.9/ 72.3
74.3/71.8
wj and wb
73.3/71.3
74.6/70.5
74.5/71.6
74.9 /71.1
73.5/ 72.5
73.9/71.4
SDAR-1.7B
uniform
59.2/56.1
61.8/56.7
63.0/56.5
61.8/57.4
64.1 /57.3
62.9/ 58.4
wj only
56.3/56.1
57.3/56.1
57.1/56.4
62.2 / 57.1
52.7/53.1
47.3/50.5
Appendix
Table 14: MATH500 accuracy (static/dynamic, avg@3) every fifth round for the runs of Table 5 . Bold: round selected under each decoding rule (static/dynamic).
Static
Dynamic
Student
Weights
LiveCodeBench-v2
LiveBench
Avg.
LiveCodeBench-v2
LiveBench
Avg.
SDAR-4B
uniform
21.2
21.5
21.4
19.5
18.0
18.8
wj only
21.8
20.8
21.3
20.0
18.5
19.3
wb only
20.6
19.3
20.0
19.5
19.3
19.4
wj and wb
20.9
22.1
21.5
19.8
19.8
19.8
SDAR-1.7B
uniform
10.3
9.4
9.9
8.7
9.1
8.9
Appendix
Table 15: Hindsight weights on code, companion of Table 5 . Checkpoints selected on LiveCodeBench-v2 per decoding rule, LiveBench at the selected round. Bold: best per column.
Diffusion large language models (dLLMs) are a promising alternative to autoregressive generation. However, reasoning-oriented post-training for dLLMs remains challenging. Supervised fine-tuning (SFT) for dLLMs requires dense but often off-policy masked states, while reinforcement learning (RL) relies on sparse rewards or value modeling. This paper proposes \textbf{trace-based on-policy distillation (TOPD)}, a teacher-supervised framework that transfers reasoning ability to a target dLLM without reward estimation. The key idea is to supervise a dLLM on its own denoising trajectory, focusing on the trace-aligned token decisions that form the final response. Specifically, TOPD samples on-policy diffusion trajectories from the target dLLM, obtains teacher token distributions from a teacher model on the corresponding partially denoised states, and updates the target dLLM with a token-level Reverse Kullback-Leibler (Reverse-KL) objective. This design preserves dense teacher supervision while aligning training with the model's own denoising states. On mathematical reasoning benchmarks, TOPD enables SDAR-4B-Chat to match the MATH500 accuracy of its RL-trained counterpart TraDo-4B-Instruct, with gains of +5.7 under static evaluation and +4.5 under dynamic evaluation. Compared with the RL-trained counterpart, TOPD achieves this with 4× fewer rollout rounds, corresponding to an estimated 96.0× to-accuracy model-compute speedup.
Haolin Ren, Ziyang Huang, Chenhao Yuan +2
Institute of Automation, Chinese Academy of Sciences, Beijing, China · University of Chinese Academy of Sciences, Beijing, China
Diffusion large language models (dLLMs) generate text by iteratively denoising a masked sequence, offering a parallel alternative to autoregressive models, but eliciting strong reasoning through post-training remains difficult: supervised fine-tuning is off-policy and suffers from exposure bias, while reinforcement learning gives only sparse, sequence-level rewards and is hard to apply without tractable sequence likelihoods. On-policy self-distillation (OPSD) offers a promising alternative, using one model as both student and teacher to provide dense, token-level, on-policy supervision, but its effectiveness hinges on giving the teacher privileged information (PI) - typically an instance-specific ground-truth reference unavailable at inference - so the student ends up distilling a weak PI-free consensus policy that yields little improvement on dLLM reasoning. We introduce dOPSD, which instead derives the teacher's privilege directly from the student's own denoising trajectory, evaluating masked positions using later, more-decoded steps of that same trajectory rather than an external label, so the teacher's advantage emerges from the model's own decoding process; on Dream and LLaDA, dOPSD improves both in-domain math reasoning and out-of-domain code generation, outperforming supervised and on-policy baselines.
Diffusion language models (dLLMs) can predict many tokens in parallel, but accurate generation still requires many iterative denoising steps. Few-step distillation accelerates decoding by compressing multiple teacher steps into a single student transition. However, existing methods construct supervision on off-policy trajectories. At inference, the student's early parallel commitments alter the context of later predictions, so the states it actually visits drift away from the supervised ones--precisely when step compression is most aggressive. On-policy distillation is a natural remedy for this mismatch, but it leaves open how far each transition should advance: matching only the teacher's next action limits compression, while indiscriminately merging future actions can violate intermediate dependencies. To address this limitation, we propose OPTD, On-Policy Transition Distillation with consistency-guided adaptive compression. It samples partial states from the few-step student's own trajectories, uses a frozen, question-only teacher to identify outcome-aligned future candidates, and orders them by current-state confidence. The method then selects the longest prefix whose joint commitment preserves the teacher's rollout outcome. A set-bottleneck objective promotes every verified future candidate to the decoder's release threshold, while a frozen-teacher KL anchor regularizes all other active positions. Neither target construction nor training uses a gold response. Across four mathematical reasoning and code-generation benchmarks, OPTD consistently improves the quality--efficiency trade-off and attains the strongest overall quality-constrained AUP among the evaluated few-step baselines.