Block diffusion accelerates speculative decoding by drafting multiple tokens in one forward pass. However, each position predicts a marginal distribution without observing earlier proposed tokens, limiting draft quality and acceptance length. We identify a concrete failure, the \emph{repetition trap}, in which neighboring positions produce redundant copies of the same token. We explain this tendency theoretically and empirically examine its association with shorter accepted drafts. Recent methods refine marginal predictions with an additional causal head or a separately trained drafter, increasing parameter storage and introducing separate training objectives. We instead propose D-Loop, which introduces \emph{intra-block causal conditioning} within the original diffusion drafter without additional model components. Inspired by semi-autoregressive generation and parameter sharing, D-Loop reuses the same backbone across looped passes. The first pass proposes a block, and the second conditions on a selected prefix to regenerate the suffix in parallel. A complementary prefix--suffix objective trains the shared drafter for both anchor-only prefix prediction and prefix-conditioned suffix prediction. Across eight math, code, and chat benchmarks, D-Loop can beat DFlash and DSpark on Qwen3-4B and Qwen3-8B with obvious gains.
Figures & tables
Dataset
Adjacent repeat
Mean longest run
Content-repeat blocks
Acc. Len.
GSM8K
17.4%
2.5
72.5%
5.28
LCB
29.4%
3.6
84.9%
3.26
MT-Bench
42.6%
4.7
94.5%
3.10
Table 1: DFlash repetition statistics and illustrative examples. Acc. Len. uses Qwen3-8B at T=1 (Table 4 ). LCB denotes LiveCodeBench.
Figure 1: Overview of DFlash and D-Loop. (a) DFlash predicts all proposal tokens in one parallel pass. (b) D-Loop retains a selected prefix from the first pass, remasks the suffix, and invokes the same drafter to regenerate the suffix conditioned on the retained tokens. Both passes share all draft parameters and use target features from the verified context. Target verification follows the second pass. One selected prefix is shown for clarity.
Dataset
Pass 1
Pass 2
GSM8K
19.4%
5.1%
LCB
28.8%
4.9%
MT-Bench
44.5%
9.0%
Table 2: Adjacent-repeat rates.
Greedy ( T=0 )
Sampling ( T=1 )
Acc. Len.
Speedup ( × )
Acc. Len.
Speedup ( × )
Dataset
DFlash
EAGLE-3
DSpark
D-Loop
DFlash
EAGLE-3
DSpark
D-Loop
DFlash
EAGLE-3
DSpark
D-Loop
DFlash
EAGLE-3
DSpark
D-Loop
GSM8K
8.51
3.77
8.83
10.20
6.41
2.27
6.62
6.77
6.39
3.68
6.83
8.96
4.82
2.12
5.12
5.94
MATH-500
8.39
3.52
8.59
9.84
6.32
2.10
6.44
6.53
5.86
3.44
6.26
7.86
4.42
1.97
4.69
5.21
AIME25
7.93
3.51
7.63
9.09
5.98
2.13
5.72
6.03
3.87
3.20
4.24
5.46
2.92
1.83
3.18
3.62
HumanEval
6.60
3.47
6.88
7.87
4.97
2.12
5.16
5.22
5.05
3.39
5.33
7.03
3.81
1.94
3.99
4.66
Table 3: Average acceptance length ( τ ) and wall-clock speedup over autoregressive decoding on Qwen3-4B under greedy ( T=0 ) and sampling ( T=1 ) decoding. DFlash, DSpark, and D-Loop results are from our runs. EAGLE-3 results are quoted from Table 1 of Chen et al. (2026) with tree size 60. Bold indicates the best result among the methods measured in our runs.
Greedy ( T=0 )
Sampling ( T=1 )
Acc. Len.
Speedup ( × )
Acc. Len.
Speedup ( × )
Dataset
DFlash
EAGLE-3
DSpark
D-Loop
DFlash
EAGLE-3
DSpark
D-Loop
DFlash
EAGLE-3
DSpark
D-Loop
DFlash
EAGLE-3
DSpark
D-Loop
GSM8K
7.35
3.71
7.42
9.46
5.58
2.23
5.63
6.61
5.28
3.59
5.74
8.24
4.01
2.07
4.36
5.76
MATH-500
7.62
3.49
8.04
9.97
5.79
2.05
6.10
6.96
5.21
3.38
5.71
7.79
3.96
1.94
4.34
5.44
AIME25
7.07
3.44
7.32
8.69
5.37
2.05
5.56
6.07
3.62
3.18
3.93
5.41
2.75
1.84
2.98
3.78
HumanEval
5.99
3.65
6.36
7.41
4.55
2.17
4.83
5.18
4.14
3.54
4.59
6.36
3.15
2.05
3.48
4.44
Table 4: Average acceptance length ( τ ) and wall-clock speedup over autoregressive decoding on Qwen3-8B under greedy ( T=0 ) and sampling ( T=1 ) decoding. DFlash, DSpark, and D-Loop results are from our runs. EAGLE-3 results are quoted from Table 1 of Chen et al. (2026) with tree size 60. Bold indicates the best result among the methods measured in our runs.
Figure 2: Qwen3-4B on GSM8K: (a) acceptance length and speedup versus K ; (b) acceptance prob. over position j .
Decoding
Method
# Draft Layers
GSM8K
MATH
AIME25
HumanEval
MBPP
LCB
Average
D 2 SD
10
9.21 / 6.54
8.56 / 6.03
7.26 / 5.10
7.47 / 5.37
6.96 / 5.03
6.44 / 4.62
7.65 / 5.45
D-Loop-suffix-loss
5
9.12 / 6.43
9.62 / 6.78
8.57 / 6.04
7.34 / 5.18
6.36 / 4.49
6.68 / 4.67
7.95 / 5.60
Greedy ( T=0 )
D-Loop-full-loss
5
9.46 / 6.61
9.97 / 6.96
8.69 / 6.07
7.41 / 5.18
6.39 / 4.46
6.75 / 4.71
8.11 / 5.67
D 2 SD
10
7.38 / 5.32
6.81 / 4.90
5.75 / 4.11
5.50 / 4.05
4.88 / 3.24
4.90 / 3.53
5.87 / 4.19
D-Loop-suffix-loss
5
8.23 / 5.80
7.60 / 5.30
5.48 / 3.87
6.07 / 4.28
5.59 / 3.94
4.58 / 3.23
6.26 / 4.40
Sampling ( T=1 )
D-Loop-full-loss
5
8.24 / 5.76
7.79 / 5.44
5.41 / 3.78
6.36 / 4.44
5.67 / 3.96
4.73 / 3.30
6.37 / 4.45
Table 5: Qwen3-8B comparison and loss ablation (acceptance length / speedup in × ). Bold marks metric maxima. D-Loop-suffix-loss uses suffix supervision only. D-Loop uses the full complementary prefix–suffix loss. D 2 SD results are from Zhang et al. (2026) , Table 3 ( γ=16 , K=4 ), using two five-layer drafters ( 5+5=10 ). Our two variants each share one five-layer drafter.
Method
tar
tdraft
tverify
AL
Speedup
ms/tok
ms/iter
ms/iter
DFlash
22.29
4.208
25.363
5.28
3.98×
DSpark
22.29
4.208
25.531
5.74
4.30×
D-Loop
22.29
9.428
25.604
8.24
5.24×
Table 6: Latency on GSM8K with Qwen3-8B. tar , tdraft , and tverify mean the time consumption for autoregressive token-by-token decoding, drafting, and target verification.
Block-diffusion drafters like dFlash generate an entire block of draft tokens in a single forward pass, drastically reducing the overhead of multiple-token drafting in speculative decoding. The crucial final step of the single-pass discrete denoising process involves using the logit distribution at each position to sample conditionally independent tokens. The resulting draft is thus a set of per-position marginals, rather than a joint distribution: no draft token is guaranteed to depend on its predecessors. Such independently sampled marginals tend to produce sequences with tokens that are individually likely, but jointly improbable under the target model's distribution, which verifies each token conditionally. This can cause early rejection and limits acceptance length. To address this, we propose xPress as a means to restore the missing causality in diffusion drafters. xPress is a lightweight causal refiner that reconciles the whole diffusion block at once through parallel refinement, restoring and propagating causal dependencies across the draft without a token-by-token loop. On Qwen3-8B, across seven math, code, and chat benchmarks, xPress raises acceptance length by about 30% on average (up to +56%) and its end-to-end decoding throughput by about 1.3 on average (up to 1.7) compared to the original dFlash diffusion drafter.
Zheng Wang, Davis Wertheimer, Yu Chin Fabian Lim +4
Speculative decoding accelerates autoregressive large language model inference by drafting multiple tokens and verifying them in a single target-model forward pass. Recent diffusion-based drafters generate an entire block of tokens in parallel but usually commit to a single draft sequence per verification: once the first mismatch occurs, all subsequent draft tokens are discarded, resulting in a limited acceptance rate. Naively batching more draft candidate sequences only introduces a marginal improvement, as redundant or poorly placed branches increase the cost of drafting and verification without proportionally increasing the number of accepted tokens. We propose D^2SD, a dual diffusion draft speculative decoding framework that organizes candidates into a confidence-guided prefix tree, where the first diffusion drafter generates a block along with per-position confidence scores that are used to identify the most likely rejection boundary and select the top-K prefix ranges for recovery; the second variable-prefix diffusion drafter re-anchors at each selected prefix and proposes alternative continuations in one batched pass; the resulting shared-prefix candidates are jointly verified via cascade attention. Empirically, D^2SD shows clear improvements over both the underlying diffusion approach and strong autoregressive speculative decoding baselines.
Liyuan Zhang, Jiarui Zhang, Jinwei Yao +6
Peking University · Tsinghua University · HKUST +2
Block diffusion speculative decoding accelerates LLM inference by predicting all tokens within a block simultaneously for the target model to verify in parallel. Predicting an entire block at once requires a sufficiently capable draft model and effective utilization of the target model's internal knowledge. However, the state-of-the-art method DFlash constrains all draft layers to share a single fused representation derived from only a few target layers, limiting per-layer expressiveness and hindering further scaling of draft capacity. In this paper, we present \modelname, which flares out the narrow conditioning bottleneck of DFlash through a lightweight layer-wise fusion mechanism: each draft layer attends to its own learnable combination of a broad set of target layers at negligible overhead, simultaneously injecting richer target knowledge and providing every draft layer with a distinct input. This enhanced per-layer expressiveness enables scaling the draft model to deeper architectures with consistent gains. We further scale training data from 800K to 2.4M samples to fully exploit the enlarged capacity. On six benchmarks spanning mathematical reasoning, code generation, and conversation, \modelname attains average wall-clock speedups of 5.52x on Qwen3-4B, 5.46x on Qwen3-8B, and 3.91x on GPT-OSS-20B, improving over DFlash by roughly 11%, 8%, and 5% respectively. Our code is available at https://github.com/Tencent/AngelSlim.
Jiebin Zhang, Zhenghan Yu, Song Liu +9
School of Computer Science, Peking University · Tencent