Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass. While prior DLLM acceleration methods primarily exploit temporal redundancy across denoising steps, we identify a complementary redundancy axis within each speculative verification step: multi-branch computational redundancy. During speculative verification, draft branches inherit most tokens from their parents while unmasking a small set of additional positions, causing large portions of hidden states to remain highly similar across branches. We propose SpecFold, an algorithm-system co-design that exploits this multi-branch redundancy to reduce the cost of multi-branch speculative verification. Algorithmically, SpecFold performs token-level residual gating and selectively reuses parent computation through folded attention and FFN while preserving residual hidden states. Systemically, a Triton kernel implementation translates this fine-grained reuse into end-to-end throughput gains through efficient sparse multi-branch execution. SpecFold is orthogonal to temporal caching and compatible with existing DLLM speculation strategies. Across two DLLM families, five models, and five standard benchmarks, SpecFold achieves up to 1.64x throughput over Spiffy and up to 1.99x over vanilla decoding, while maintaining comparable task performance.
Figures & tables
Figure 1: Overview of SpecFold . (a) Standard DLLM decoding iteratively unmasks tokens through repeated denoising forward passes. (b) Multi-branch speculative DLLM decoding verifies a main branch together with several draft branches in one batched forward pass. If a draft matches a subsequent main-branch state, its precomputed logits are reused to skip a future denoising forward pass. (c) SpecFold exploits cross-branch redundancy within each verification step. Before every Transformer layer, a parent–child residual gate identifies folded positions, which reuse parent computation, while divergent positions are recomputed through the folded attention and FFN execution, improving decoding throughput.
Figure 2: Per-token relative residual ρ of the layer-input hidden state between each draft branch and its immediate parent branch in the draft graph, recorded during the verification steps of Fast-dLLM-v2-7B decoding a GSM8K sample. Columns are the 8 draft branches ( D denotes a branch’s depth in the draft graph), and rows are layers spanning the model depth. Within a panel, each row is one denoising step and each column one block position. The color indicates the level of divergence.
Figure 3: Mean latency per forward pass of dense verification versus draft budget B on Nemotron-Labs-Diffusion-3B on H100.
Benchmark
Method
Fast-dLLM-v2-1.5B
Fast-dLLM-v2-7B
NLD-3B
NLD-8B
NLD-14B
NFE ↓
TPS ↑
Acc. ↑
NFE ↓
TPS ↑
Acc. ↑
NFE ↓
TPS ↑
Acc. ↑
NFE ↓
TPS ↑
Acc. ↑
NFE ↓
TPS ↑
Acc. ↑
GSM8K
Vanilla
113
145
63.08%
113
170
82.41%
112
144
88.02%
128
102
84.84%
113
84
83.02%
Spiffy
90
150
62.77%
95
173
83.09%
86
161
87.95%
93
117
84.84%
82
98
83.47%
SpecFold
90
229 ( ↑ 58%)
62.55%
96
188 ( ↑ 11%)
81.96%
86
204 ( ↑ 42%)
88.17%
93
131 ( ↑ 29%)
84.61%
82
102 ( ↑ 21%)
82.87%
MATH
Vanilla
263
173
38.60%
229
160
57.80%
225
156
72.60%
297
113
83.80%
246
104
84.00%
Spiffy
219
189
37.40%
189
163
57.40%
172
164
72.60%
219
126
82.20%
183
110
84.60%
Table 1: SpecFold evaluation, reporting per-sample NFE, decoding throughput (tokens/s), and accuracy across five benchmarks and five model scales spanning the Fast-dLLM-v2 and Nemotron-Labs-Diffusion families. SpecFold uses recompute threshold δ=0.1 . Nemotron-Labs-Diffusion models run in diffusion mode. The MATH row uses the MATH-500 test set.
Figure 4: Recompute-threshold sweep of SpecFold on Nemotron-Labs-Diffusion-8B (H100, B=7 ). Each panel plots task accuracy against the throughput (TPS) gain over Vanilla as δ increases from 0.01 to 0.7 (darker markers = larger δ , i.e., more folding); the dotted line marks the Vanilla accuracy. Upper-right is better.
Figure 5: Draft-budget sweep on Nemotron-Labs-Diffusion-3B, fixing δ=0.1 . Each panel plots the throughput gain (TPS) over Vanilla against the draft budget B . The gap between the two curves demonstrates the effect of folding.
Figure 6: Per-position relative residual between a parent draft and a child draft that unmasks a single token. Each model decodes a GSM8K prompt to the middle of a block: the parent draft is the block state with 17 of the 32 positions committed, and the child draft adds one more unmasked token being the target’s next most-confident masked position (the “edited position”, white dotted line). Heatmaps show the per-position parent–child relative residual ρnℓ (Eq. 4 ) across layers.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Kernel
Grid
Role
residual_delta_gate
(B+1)N
gate, source map, compaction, RMSNorm stats
src_resolve
(B+1)N
pointer-jump fold chains
ln_qkv_rope_scatter
∣R∣
norm + packed QKV + RoPE for recomputed positions
flash_attn (+state)
∣R∣×H
recomputed-position attention; emits (m,Z,O)
chain_correction
(B+1)N×H
correcting attention for folded positions
ln_gateup_silu /
∣R∣
fused FFN + residual add for recomputed positions
Appendix
Table 2: Kernel pipeline for one SpecFoldLayer . R denotes the union of the per-branch recompute sets Riℓ . The grid sizes are per-layer launch dimensions.
Figure 7: Draft-budget sweep on the four remaining models (rows) across the five benchmarks (columns), presented as in Fig. 5 : each panel plots the throughput gain (TPS) of Spiffy and SpecFold over Vanilla against the draft budget B , with the dotted line marking Vanilla. At every B the two methods share identical drafts and NFE, so the gap between the curves is the effect of folding.
Figure 8: Drafting-strategy ablation on Nemotron-Labs-Diffusion-3B: throughput gain (TPS) over Vanilla versus the draft-graph depth cap D at fixed budget B=7 , presented as in Fig. 5 . Spiffy and SpecFold share identical drafts and NFE at every depth.
Benchmark
Method
A100
H100
TPS ↑
TPS gain
Acc. ↑
TPS ↑
TPS gain
Acc. ↑
GSM8K
Vanilla
81
1.00 ×
88.32%
144
1.00 ×
88.02%
Spiffy
92
1.14 ×
88.02%
162
1.12 ×
87.95%
SpecFold
121
1.49 ×
88.78%
197
1.36 ×
88.63%
MATH
Vanilla
92
1.00 ×
72.80%
156
1.00 ×
72.60%
Spiffy
103
1.12 ×
74.20%
167
1.07 ×
72.60%
Appendix
Table 3: GPU-type ablation on Nemotron-Labs-Diffusion-3B: decoding throughput (tokens/s), TPS gain over the same GPU’s vanilla decoding, and task accuracy.
Diffusion large language models (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs, offering faster inference through parallel or blockwise decoding. However, their masked language modeling formulation remains incompatible with standard token-level speculative decoding, one of the most effective acceleration techniques for AR models. In AR decoding, the causal mask preserves temporally valid token-level contexts, enabling a target model to verify multiple drafted tokens in a single forward pass. In contrast, dLLMs rely on mask tokens and bidirectional attention, causing the effective context to change across denoising steps and preventing direct token-level speculative verification. To bridge this gap, we propose a simple but effective speculative decoding algorithm for diffusion language models, named SimSD, which mainly adopts a plug-and-play masking strategy that equips dLLMs with temporally valid token-level contexts for speculative decoding. Our method explicitly introduces reference tokens from draft-model predictions and designs an attention mask that regulates their interaction with current-step tokens, allowing dLLMs to compute valid logits for drafted tokens in a single forward pass. This restores the key verification ability provided by causal masking in AR models while preserving the parallel decoding advantages of dLLMs. The proposed method is training-free and can be flexibly integrated with other acceleration techniques such as KV cache and blockwise decoding. Experiments on SDAR-family dLLMs across four benchmarks show that our method achieves up to 7.46x higher decoding throughput while maintaining and even improving average generation quality.
Junxia Cui, Haotian Ye, Runchu Tian +9
University of California San Diego · University of Illinois Urbana-Champaign · 3Google +1
Speculative decoding accelerates large language model inference by drafting multiple tokens for parallel verification, with efficiency critically determined by the speculative length selected at each decoding round. Existing dynamic speculation methods select the speculation length by estimating how many tokens will be accepted, which is reasonable for autoregressive drafters that generates tokens sequentially. The recent wave of diffusion-based drafters, however, generates candidate blocks in parallel at substantially lower drafting cost, shifting the key question from how many tokens to generate to how many generated tokens are worth verifying. We therefore reformulate dynamic speculative-length selection as expected-speedup optimization and derive a marginal criterion that extends the speculative sequence only when its acceptance gain outweighs the additional verification cost. Building on this criterion, we develop \textit{LibraSpec}, a training-free and plug-and-play algorithm that iteratively determines the speculative length using drafter confidence scores. Theoretically, we prove that LibraSpec monotonically converges toward the optimal speculative length. Experiments across six target models, three diffusion-based speculative decoding methods, and math, coding, and chat benchmarks show consistent improvements under both greedy and sampling settings, achieving a further 0.5∼1.5× improvement over baselines and up to 8.49× speedup over autoregressive decoding.
Zexun Lin, Yuan Feng, Junlin Lv +2
Suzhou Institute for Advanced Research, University of Science and Technology of China Suzhou, Jiangsu, China
Speculative decoding accelerates autoregressive large language model inference by drafting multiple tokens and verifying them in a single target-model forward pass. Recent diffusion-based drafters generate an entire block of tokens in parallel but usually commit to a single draft sequence per verification: once the first mismatch occurs, all subsequent draft tokens are discarded, resulting in a limited acceptance rate. Naively batching more draft candidate sequences only introduces a marginal improvement, as redundant or poorly placed branches increase the cost of drafting and verification without proportionally increasing the number of accepted tokens. We propose D^2SD, a dual diffusion draft speculative decoding framework that organizes candidates into a confidence-guided prefix tree, where the first diffusion drafter generates a block along with per-position confidence scores that are used to identify the most likely rejection boundary and select the top-K prefix ranges for recovery; the second variable-prefix diffusion drafter re-anchors at each selected prefix and proposes alternative continuations in one batched pass; the resulting shared-prefix candidates are jointly verified via cascade attention. Empirically, D^2SD shows clear improvements over both the underlying diffusion approach and strong autoregressive speculative decoding baselines.
Liyuan Zhang, Jiarui Zhang, Jinwei Yao +6
Peking University · Tsinghua University · HKUST +2