Block diffusion accelerates speculative decoding by drafting multiple tokens in one forward pass. However, each position predicts a marginal distribution without observing earlier proposed tokens, limiting draft quality and acceptance length. We identify a concrete failure, the \emph{repetition trap}, in which neighboring positions produce redundant copies of the same token. We explain this tendency theoretically and empirically examine its association with shorter accepted drafts. Recent methods refine marginal predictions with an additional causal head or a separately trained drafter, increasing parameter storage and introducing separate training objectives. We instead propose D-Loop, which introduces \emph{intra-block causal conditioning} within the original diffusion drafter without additional model components. Inspired by semi-autoregressive generation and parameter sharing, D-Loop reuses the same backbone across looped passes. The first pass proposes a block, and the second conditions on a selected prefix to regenerate the suffix in parallel. A complementary prefix--suffix objective trains the shared drafter for both anchor-only prefix prediction and prefix-conditioned suffix prediction. Across eight math, code, and chat benchmarks, D-Loop can beat DFlash and DSpark on Qwen3-4B and Qwen3-8B with obvious gains.
Figures & tables
Dataset
Adjacent repeat
Mean longest run
Content-repeat blocks
Acc. Len.
GSM8K
17.4%
2.5
72.5%
5.28
LCB
29.4%
3.6
84.9%
3.26
MT-Bench
42.6%
4.7
94.5%
3.10
Table 1: DFlash repetition statistics and illustrative examples. Acc. Len. uses Qwen3-8B at T=1 (Table 4 ). LCB denotes LiveCodeBench.
Figure 1: Overview of DFlash and D-Loop. (a) DFlash predicts all proposal tokens in one parallel pass. (b) D-Loop retains a selected prefix from the first pass, remasks the suffix, and invokes the same drafter to regenerate the suffix conditioned on the retained tokens. Both passes share all draft parameters and use target features from the verified context. Target verification follows the second pass. One selected prefix is shown for clarity.
Dataset
Pass 1
Pass 2
GSM8K
19.4%
5.1%
LCB
28.8%
4.9%
MT-Bench
44.5%
9.0%
Table 2: Adjacent-repeat rates.
Greedy ( T=0 )
Sampling ( T=1 )
Acc. Len.
Speedup ( × )
Acc. Len.
Speedup ( × )
Dataset
DFlash
EAGLE-3
DSpark
D-Loop
DFlash
EAGLE-3
DSpark
D-Loop
DFlash
EAGLE-3
DSpark
D-Loop
DFlash
EAGLE-3
DSpark
D-Loop
GSM8K
8.51
3.77
8.83
10.20
6.41
2.27
6.62
6.77
6.39
3.68
6.83
8.96
4.82
2.12
5.12
5.94
MATH-500
8.39
3.52
8.59
9.84
6.32
2.10
6.44
6.53
5.86
3.44
6.26
7.86
4.42
1.97
4.69
5.21
AIME25
7.93
3.51
7.63
9.09
5.98
2.13
5.72
6.03
3.87
3.20
4.24
5.46
2.92
1.83
3.18
3.62
HumanEval
6.60
3.47
6.88
7.87
4.97
2.12
5.16
5.22
5.05
3.39
5.33
7.03
3.81
1.94
3.99
4.66
Table 3: Average acceptance length ( τ ) and wall-clock speedup over autoregressive decoding on Qwen3-4B under greedy ( T=0 ) and sampling ( T=1 ) decoding. DFlash, DSpark, and D-Loop results are from our runs. EAGLE-3 results are quoted from Table 1 of Chen et al. (2026) with tree size 60. Bold indicates the best result among the methods measured in our runs.
Greedy ( T=0 )
Sampling ( T=1 )
Acc. Len.
Speedup ( × )
Acc. Len.
Speedup ( × )
Dataset
DFlash
EAGLE-3
DSpark
D-Loop
DFlash
EAGLE-3
DSpark
D-Loop
DFlash
EAGLE-3
DSpark
D-Loop
DFlash
EAGLE-3
DSpark
D-Loop
GSM8K
7.35
3.71
7.42
9.46
5.58
2.23
5.63
6.61
5.28
3.59
5.74
8.24
4.01
2.07
4.36
5.76
MATH-500
7.62
3.49
8.04
9.97
5.79
2.05
6.10
6.96
5.21
3.38
5.71
7.79
3.96
1.94
4.34
5.44
AIME25
7.07
3.44
7.32
8.69
5.37
2.05
5.56
6.07
3.62
3.18
3.93
5.41
2.75
1.84
2.98
3.78
HumanEval
5.99
3.65
6.36
7.41
4.55
2.17
4.83
5.18
4.14
3.54
4.59
6.36
3.15
2.05
3.48
4.44
Table 4: Average acceptance length ( τ ) and wall-clock speedup over autoregressive decoding on Qwen3-8B under greedy ( T=0 ) and sampling ( T=1 ) decoding. DFlash, DSpark, and D-Loop results are from our runs. EAGLE-3 results are quoted from Table 1 of Chen et al. (2026) with tree size 60. Bold indicates the best result among the methods measured in our runs.
Figure 2: Qwen3-4B on GSM8K: (a) acceptance length and speedup versus K ; (b) acceptance prob. over position j .
Decoding
Method
# Draft Layers
GSM8K
MATH
AIME25
HumanEval
MBPP
LCB
Average
D 2 SD
10
9.21 / 6.54
8.56 / 6.03
7.26 / 5.10
7.47 / 5.37
6.96 / 5.03
6.44 / 4.62
7.65 / 5.45
D-Loop-suffix-loss
5
9.12 / 6.43
9.62 / 6.78
8.57 / 6.04
7.34 / 5.18
6.36 / 4.49
6.68 / 4.67
7.95 / 5.60
Greedy ( T=0 )
D-Loop-full-loss
5
9.46 / 6.61
9.97 / 6.96
8.69 / 6.07
7.41 / 5.18
6.39 / 4.46
6.75 / 4.71
8.11 / 5.67
D 2 SD
10
7.38 / 5.32
6.81 / 4.90
5.75 / 4.11
5.50 / 4.05
4.88 / 3.24
4.90 / 3.53
5.87 / 4.19
D-Loop-suffix-loss
5
8.23 / 5.80
7.60 / 5.30
5.48 / 3.87
6.07 / 4.28
5.59 / 3.94
4.58 / 3.23
6.26 / 4.40
Sampling ( T=1 )
D-Loop-full-loss
5
8.24 / 5.76
7.79 / 5.44
5.41 / 3.78
6.36 / 4.44
5.67 / 3.96
4.73 / 3.30
6.37 / 4.45
Table 5: Qwen3-8B comparison and loss ablation (acceptance length / speedup in × ). Bold marks metric maxima. D-Loop-suffix-loss uses suffix supervision only. D-Loop uses the full complementary prefix–suffix loss. D 2 SD results are from Zhang et al. (2026) , Table 3 ( γ=16 , K=4 ), using two five-layer drafters ( 5+5=10 ). Our two variants each share one five-layer drafter.
Method
tar
tdraft
tverify
AL
Speedup
ms/tok
ms/iter
ms/iter
DFlash
22.29
4.208
25.363
5.28
3.98×
DSpark
22.29
4.208
25.531
5.74
4.30×
D-Loop
22.29
9.428
25.604
8.24
5.24×
Table 6: Latency on GSM8K with Qwen3-8B. tar , tdraft , and tverify mean the time consumption for autoregressive token-by-token decoding, drafting, and target verification.