Diffusion drafters accelerate speculative decoding by proposing multiple tokens in parallel. Despite recent advances in speculative decoding through sequence-level drafting and verification, existing training objectives remain largely designed around token-level verification. To address this mismatch, we introduce Block Verification-aware loss (BV loss), a training objective designed to maximize the expected acceptance length of a drafted sequence. BV loss is directly derived from the block verification acceptance rule, providing a principled connection between the drafter training objective and the inference-time verification mechanism at the sequence level. Across math, code, and chat benchmarks, BV loss increases the mean number of tokens accepted per verification call under block verification by 13.0--21.0% over cross-entropy loss training for DFlash and DSpark with Qwen3-4B and Qwen3-8B without changing the inference procedure. BV loss also outperforms tokenwise acceptance objectives such as TV loss and LK loss, and its gains extend to token verification and greedy decoding. These results demonstrate the benefit of training block diffusion drafters with an objective aligned with sequence-level verification, rather than optimizing each token independently.
Figures & tables
Drafter loss
Training objective
Uses teacher probabilities?
Verification- aware?
Block- aware?
CE
Maximize log-likelihood
✗
✗
✗
KL / RKL
Minimize KL divergence
✓
✗
✗
TV
Maximize token acceptance
✓
✓
✗
LK
Maximize token acceptance
✓
✓
✗
BV (Ours)
Maximize block acceptance length
✓
✓
✓
Table 1: Comparison of drafter training objectives. CE and KL / RKL optimize token likelihood or distribution matching, whereas TV and LK loss explicitly target token-level acceptance. BV loss is instead derived from block verification and directly optimizes the expected accepted draft length, making it both verification-aware and block-aware.
T=0 (greedy)
T=1 (block verification)
DFlash
DSpark
DFlash
DSpark
Target
Loss
τ
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
Qwen3-4B
CE
4.65
3.71 ×
5.21
3.97 ×
4.09
3.21 ×
4.77
3.47 ×
KL
4.80
3.76 ×
5.49
4.15 ×
4.17
3.23 ×
4.97
3.67 ×
RKL
4.05
3.24 ×
4.87
3.73 ×
3.87
3.09 ×
4.68
3.45 ×
TV
1.98
1.60 ×
2.23
1.74 ×
1.96
1.57 ×
2.21
1.66 ×
Table 2: Average performance over seven benchmarks on Qwen3-4B and Qwen3-8B under greedy decoding ( T=0 ) and block verification ( T=1 ). Average acceptance length τ includes correction/bonus tokens. Speedup is relative to autoregressive decoding. Boldface indicates the best average in each column within each target.
Target
Drafter
τ under token verification
CE
KL
RKL
TV
LK
BV (Ours)
Qwen3-4B
DFlash
4.01
4.11
3.86
1.96
4.17
4.85
DSpark
4.64
4.84
4.64
2.20
4.98
5.20
Qwen3-8B
DFlash
4.04
4.13
3.00
1.01
4.19
4.86
DSpark
4.63
4.90
3.38
1.02
5.05
5.26
Table 3: Generalization to token verification at T=1 with the same checkpoints as Table 2 . Entries are seven-benchmark mean τ , using the same counting convention. Bold marks the best mean in each row, including ties.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Qwen3-4B
Qwen3-8B
DFlash
DSpark
DFlash
DSpark
Training loss
T=0
T=1
T=0
T=1
T=0
T=1
T=0
T=1
Baseline (CE)
4.65
4.09
5.21
4.77
4.64
4.15
5.10
4.71
BV w/o annealing
4.98
4.74
5.37
5.17
5.07
4.80
5.36
5.13
BV with annealing (Ours)
5.22
4.95
5.62
5.39
5.27
4.98
5.67
5.43
Appendix
Table 4: BV and annealing ablation. Entries are average τ over the seven benchmarks, with the counting convention of Table 2 . At T=1 , both drafters use block verification.
Drafter
Loss
Math
Code
Chat
GSM8K
MATH-500
AIME25
HumanEval
MBPP
LCB
MT-Bench
Avg.
T=0
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
DFlash
CE
4.99 ×
6.32
4.46 ×
5.64
3.63 ×
4.50
3.53 ×
4.45
3.41 ×
4.27
3.59 ×
4.44
2.34 ×
2.94
3.71 ×
4.65
KL
5.05 ×
6.51
4.54 ×
5.83
3.68 ×
4.67
3.60 ×
4.61
3.46 ×
4.40
3.61 ×
4.56
2.40 ×
3.03
3.76 ×
4.80
RKL
4.52 ×
5.69
3.84 ×
4.81
3.09 ×
3.85
3.05 ×
3.82
2.99 ×
3.71
3.22 ×
3.98
2.00 ×
2.51
3.24 ×
4.05
TV
1.95 ×
2.42
2.01 ×
2.48
1.78 ×
2.19
1.39 ×
1.71
1.44 ×
1.77
1.34 ×
1.63
1.32 ×
1.64
1.60 ×
1.98
Appendix
Table 5: Full benchmark results for DFlash and DSpark on Qwen3-4B. Here, τ is the reported verification-call length, including correction/bonus when present; the Method uses draft-only τ . Avg. retains the reported seven-benchmark mean. Speedup is relative to autoregressive decoding. Bold marks the best value for each metric within each drafter and setting..
Drafter
Loss
Math
Code
Chat
GSM8K
MATH-500
AIME25
HumanEval
MBPP
LCB
MT-Bench
Avg.
T=0
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
DFlash
CE
4.70 ×
6.11
4.28 ×
5.57
3.53 ×
4.55
3.54 ×
4.61
3.26 ×
4.20
3.58 ×
4.58
2.32 ×
2.88
3.60 ×
4.64
KL
4.86 ×
6.37
4.42 ×
5.79
3.59 ×
4.67
3.69 ×
4.80
3.37 ×
4.34
3.67 ×
4.72
2.39 ×
2.97
3.71 ×
4.81
RKL
2.95 ×
3.82
2.80 ×
3.63
2.34 ×
3.01
2.33 ×
3.03
2.34 ×
3.01
2.44 ×
3.11
1.71 ×
2.12
2.42 ×
3.10
TV
0.81 ×
1.02
0.81 ×
1.01
0.80 ×
1.01
0.80 ×
1.00
0.79 ×
1.01
0.82 ×
1.01
0.82 ×
1.01
0.81 ×
1.01
Appendix
Table 6: Full benchmark results for DFlash and DSpark on Qwen3-8B. Here, τ is the reported verification-call length, including correction/bonus when present; the Method uses draft-only τ . Avg. retains the reported seven-benchmark mean. Speedup is relative to autoregressive decoding. Bold marks the best value for each metric within each drafter and setting.