Diffusion drafters accelerate speculative decoding by proposing multiple tokens in parallel. Despite recent advances in speculative decoding through sequence-level drafting and verification, existing training objectives remain largely designed around token-level verification. To address this mismatch, we introduce Block Verification-aware loss (BV loss), a training objective designed to maximize the expected acceptance length of a drafted sequence. BV loss is directly derived from the block verification acceptance rule, providing a principled connection between the drafter training objective and the inference-time verification mechanism at the sequence level. Across math, code, and chat benchmarks, BV loss increases the mean number of tokens accepted per verification call under block verification by 13.0--21.0% over cross-entropy loss training for DFlash and DSpark with Qwen3-4B and Qwen3-8B without changing the inference procedure. BV loss also outperforms tokenwise acceptance objectives such as TV loss and LK loss, and its gains extend to token verification and greedy decoding. These results demonstrate the benefit of training block diffusion drafters with an objective aligned with sequence-level verification, rather than optimizing each token independently.
Figures & tables
Drafter loss
Training objective
Uses teacher probabilities?
Verification- aware?
Block- aware?
CE
Maximize log-likelihood
✗
✗
✗
KL / RKL
Minimize KL divergence
✓
✗
✗
TV
Maximize token acceptance
✓
✓
✗
LK
Maximize token acceptance
✓
✓
✗
BV (Ours)
Maximize block acceptance length
✓
✓
✓
Table 1: Comparison of drafter training objectives. CE and KL / RKL optimize token likelihood or distribution matching, whereas TV and LK loss explicitly target token-level acceptance. BV loss is instead derived from block verification and directly optimizes the expected accepted draft length, making it both verification-aware and block-aware.
T=0 (greedy)
T=1 (block verification)
DFlash
DSpark
DFlash
DSpark
Target
Loss
τ
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
Qwen3-4B
CE
4.65
3.71 ×
5.21
3.97 ×
4.09
3.21 ×
4.77
3.47 ×
KL
4.80
3.76 ×
5.49
4.15 ×
4.17
3.23 ×
4.97
3.67 ×
RKL
4.05
3.24 ×
4.87
3.73 ×
3.87
3.09 ×
4.68
3.45 ×
TV
1.98
1.60 ×
2.23
1.74 ×
1.96
1.57 ×
2.21
1.66 ×
Table 2: Average performance over seven benchmarks on Qwen3-4B and Qwen3-8B under greedy decoding ( T=0 ) and block verification ( T=1 ). Average acceptance length τ includes correction/bonus tokens. Speedup is relative to autoregressive decoding. Boldface indicates the best average in each column within each target.
Target
Drafter
τ under token verification
CE
KL
RKL
TV
LK
BV (Ours)
Qwen3-4B
DFlash
4.01
4.11
3.86
1.96
4.17
4.85
DSpark
4.64
4.84
4.64
2.20
4.98
5.20
Qwen3-8B
DFlash
4.04
4.13
3.00
1.01
4.19
4.86
DSpark
4.63
4.90
3.38
1.02
5.05
5.26
Table 3: Generalization to token verification at T=1 with the same checkpoints as Table 2 . Entries are seven-benchmark mean τ , using the same counting convention. Bold marks the best mean in each row, including ties.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Qwen3-4B
Qwen3-8B
DFlash
DSpark
DFlash
DSpark
Training loss
T=0
T=1
T=0
T=1
T=0
T=1
T=0
T=1
Baseline (CE)
4.65
4.09
5.21
4.77
4.64
4.15
5.10
4.71
BV w/o annealing
4.98
4.74
5.37
5.17
5.07
4.80
5.36
5.13
BV with annealing (Ours)
5.22
4.95
5.62
5.39
5.27
4.98
5.67
5.43
Appendix
Table 4: BV and annealing ablation. Entries are average τ over the seven benchmarks, with the counting convention of Table 2 . At T=1 , both drafters use block verification.
Drafter
Loss
Math
Code
Chat
GSM8K
MATH-500
AIME25
HumanEval
MBPP
LCB
MT-Bench
Avg.
T=0
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
DFlash
CE
4.99 ×
6.32
4.46 ×
5.64
3.63 ×
4.50
3.53 ×
4.45
3.41 ×
4.27
3.59 ×
4.44
2.34 ×
2.94
3.71 ×
4.65
KL
5.05 ×
6.51
4.54 ×
5.83
3.68 ×
4.67
3.60 ×
4.61
3.46 ×
4.40
3.61 ×
4.56
2.40 ×
3.03
3.76 ×
4.80
RKL
4.52 ×
5.69
3.84 ×
4.81
3.09 ×
3.85
3.05 ×
3.82
2.99 ×
3.71
3.22 ×
3.98
2.00 ×
2.51
3.24 ×
4.05
TV
1.95 ×
2.42
2.01 ×
2.48
1.78 ×
2.19
1.39 ×
1.71
1.44 ×
1.77
1.34 ×
1.63
1.32 ×
1.64
1.60 ×
1.98
Appendix
Table 5: Full benchmark results for DFlash and DSpark on Qwen3-4B. Here, τ is the reported verification-call length, including correction/bonus when present; the Method uses draft-only τ . Avg. retains the reported seven-benchmark mean. Speedup is relative to autoregressive decoding. Bold marks the best value for each metric within each drafter and setting..
Drafter
Loss
Math
Code
Chat
GSM8K
MATH-500
AIME25
HumanEval
MBPP
LCB
MT-Bench
Avg.
T=0
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
Speedup
τ
DFlash
CE
4.70 ×
6.11
4.28 ×
5.57
3.53 ×
4.55
3.54 ×
4.61
3.26 ×
4.20
3.58 ×
4.58
2.32 ×
2.88
3.60 ×
4.64
KL
4.86 ×
6.37
4.42 ×
5.79
3.59 ×
4.67
3.69 ×
4.80
3.37 ×
4.34
3.67 ×
4.72
2.39 ×
2.97
3.71 ×
4.81
RKL
2.95 ×
3.82
2.80 ×
3.63
2.34 ×
3.01
2.33 ×
3.03
2.34 ×
3.01
2.44 ×
3.11
1.71 ×
2.12
2.42 ×
3.10
TV
0.81 ×
1.02
0.81 ×
1.01
0.80 ×
1.01
0.80 ×
1.00
0.79 ×
1.01
0.82 ×
1.01
0.82 ×
1.01
0.81 ×
1.01
Appendix
Table 6: Full benchmark results for DFlash and DSpark on Qwen3-8B. Here, τ is the reported verification-call length, including correction/bonus when present; the Method uses draft-only τ . Avg. retains the reported seven-benchmark mean. Speedup is relative to autoregressive decoding. Bold marks the best value for each metric within each drafter and setting.
Speculative decoding speeds up generation with an efficient draft model (drafter) that proposes tokens for a target model to verify in one pass, preserving the target's output distribution. High-acceptance block-diffusion drafters such as DFlash and DFlare fill an entire block in one parallel pass. In many cycles, the target accepts the whole block, so the drafter exhausts its trained block horizon before verification fails. We call this unrealized acceptance stranded speed-up. A mean committed length, per prompt or per cycle, hides it, whereas the acceptance histogram exposes it as a spike in the ceiling bin, the fraction of cycles that accept the entire block. We recommend the histogram as a preflight check before spending training compute. Naively widening the block at inference does not recover the speed-up, because once the block outgrows its training size, the drafter's bidirectional attention shifts its distribution even at early positions and erodes front-of-block verification. Instead, we post-train the drafter on a longer block with a short curriculum that emphasizes the newly exposed positions, a method we call DBloom. Expanding the pretrained DFlash and DFlare drafters from block size 16 to 24 across Qwen3-8B and Qwen3-4B targets raises the per-prompt committed length on the high-ceiling benchmarks by a median of +0.8 tokens (up to +1.1). Once continuation fine-tuning precedes expansion, the increase reaches 1.37 tokens. The same expansion also lifts committed length on all seven benchmarks for Gemma-4-12B-IT, a different model family, by a median of +0.41 tokens (Arm A), and the full continuation-then-expand pipeline (Arm B) adds +0.29 to +0.98 tokens over the same B16 drafter. In a prompt-matched comparison against JetSpec, a contemporary tree-based drafter not used in our design, DBloom commits more tokens on every benchmark at tree budgets up to 64 nodes.
Block-diffusion drafters like dFlash generate an entire block of draft tokens in a single forward pass, drastically reducing the overhead of multiple-token drafting in speculative decoding. The crucial final step of the single-pass discrete denoising process involves using the logit distribution at each position to sample conditionally independent tokens. The resulting draft is thus a set of per-position marginals, rather than a joint distribution: no draft token is guaranteed to depend on its predecessors. Such independently sampled marginals tend to produce sequences with tokens that are individually likely, but jointly improbable under the target model's distribution, which verifies each token conditionally. This can cause early rejection and limits acceptance length. To address this, we propose xPress as a means to restore the missing causality in diffusion drafters. xPress is a lightweight causal refiner that reconciles the whole diffusion block at once through parallel refinement, restoring and propagating causal dependencies across the draft without a token-by-token loop. On Qwen3-8B, across seven math, code, and chat benchmarks, xPress raises acceptance length by about 30% on average (up to +56%) and its end-to-end decoding throughput by about 1.3 on average (up to 1.7) compared to the original dFlash diffusion drafter.
Zheng Wang, Davis Wertheimer, Yu Chin Fabian Lim +4
Speculative decoding speeds up LLM inference by using a draft model to generate tokens, with an acceptance-rejection scheme that ensures that the output matches the target distribution. Adapting this to continuous diffusions is difficult because speculative sampling requires drawing from a residual distribution. While straightforward in discrete spaces, efficiently sampling this residual in continuous space is non-trivial. Consequently, existing diffusion adaptations either use computationally inefficient sampling techniques or rely on an alternative scheme. In this work, we introduce a novel scheme that efficiently implements the original speculative sampling mechanism for diffusion models. Our approach offers a critical advantage over current methods: it enables us to adapt block verification from LLMs to diffusions -- which provably improves the acceptance rate of drafts. Furthermore, we formalize and analyze the Free Drafter, a heuristic self-speculative drafter for diffusions that requires no training. By enabling block verification, our Free Drafter yields up to a 6.3% speedup over existing speculative methods with no additional training and negligible overhead beyond the existing parallel verification pass.
Alexander Soen, Hisham Husain, Valentin De Bortoli +1