Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context grows, existing state-of-the-art drafters become increasingly expensive, eroding the very efficiency advantage they are designed to provide. We argue that this scaling is unnecessary. A standalone language model must grow with its prefix because it is solely responsible for every token it produces. A drafter, by contrast, only proposes candidates; the target catches and corrects every error before any token is committed. The drafter's decoding cost can therefore be made entirely independent of the prefix length. We introduce LongSpark, a block-diffusion drafter that achieves this by extracting fixed-size, multiscale views from the target's verification pass, thereby eliminating the need for a growing persistent state. Extensive evaluations demonstrate that LongSpark achieves state-of-the-art end-to-end efficiency across multiple model scales and realistic serving conditions. Notably, it delivers the lowest time-per-output-token on long-context tasks while reducing the drafter's context state by several orders of magnitude.
Figures & tables
Figure 1: LongSpark : Scalable speculative decoding via fixed-cost drafting. Its O(1) drafting cost keeps drafting time nearly constant as the prefix grows, with 406× smaller drafter context state than DSpark at 128K tokens (left). Drafting time includes proposal and sampling. Lower drafting overhead improves end-to-end throughput on LongSpec (right). Both panels use Qwen3-8B with concurrency 16.
Figure 2: One decoding round of LongSpark . (a) The target verifies the proposed tokens. (b) Three fixed-size context views are extracted from the target’s state: a boundary state, a recent KV window, and a global context summary. (c) The drafter uses these views to propose the next token block at a cost independent of prefix length.
Figure 3: Incremental global context summary during inference. The summary is initialized over the full prefix during target prefill, then updated using only newly retained KV rows N ( Δ=∣N∣≤B+1 ).
Target
Drafter
Math
Code
Chat
Overall
GSM8K
MATH-500
AIME25
MBPP
HumanEval
LCB
MT-Bench
Alpaca
Avg.
Q3-4B
EAGLE-3
5.15 (1.36)
4.60 (1.35)
3.84 (1.24)
3.69 (1.03)
4.15 (1.17)
3.77 (1.14)
2.42 (0.72)
2.26 (0.69)
3.74 (1.09)
DFlash
5.37 (1.81)
4.89 (1.89)
4.01 (1.72)
4.40 (1.55)
4.74 (1.72)
4.23 (1.61)
3.05 (1.18)
2.94 (1.14)
4.20 (1.58)
DSpark
6.10 (1.99)
5.72 (2.14)
4.92 (2.05)
5.14 (1.73)
5.44 (1.91)
4.92 (1.79)
3.65 (1.36)
3.55 (1.32)
4.93 (1.79)
LongSpark
5.90 ( 2.08 )
5.52 ( 2.28 )
4.67 ( 2.17 )
5.01 ( 1.84 )
5.12 ( 1.97 )
4.63 ( 1.86 )
3.52 ( 1.46 )
3.46 ( 1.42 )
4.73 ( 1.88 )
Q3-8B
EAGLE-3
5.26 (1.49)
4.75 (1.50)
4.02 (1.37)
3.92 (1.16)
4.32 (1.30)
4.17 (1.28)
2.67 (0.85)
2.54 (0.82)
3.96 (1.22)
Table 1: Speculative decoding across model scales and tasks. Entries report accepted length τ , with throughput speedup over autoregressive decoding in parentheses. Avg. denotes the mean across eight benchmarks; bold marks the highest speedup.
Figure 4: Higher concurrency increases serving load. As this load rises, LongSpark ’s throughput advantage over DSpark grows.
Method
LongSpec 32K
CodeSpan 64K
LSWE 128K
TPS ↑
TPOT ↓
TPS ↑
TPOT ↓
TPS ↑
TPOT ↓
Vanilla
1128.9
14.0
884.4
17.8
160.5
97.2
EAGLE-3
1258.5
12.5
1071.8
14.7
158.7
97.4
DFlash
1313.7
12.0
919.1
17.2
163.6
94.6
DSpark
1736.6
9.0
1217.3
13.0
190.9
81.4
LongSpark
1917.6
8.2
1571.8
10.0
212.7
74.3
Table 2: Long-context decoding on Qwen3-8B. We report TPS (tokens/s) and TPOT (ms/token), averaged over ten seeds. Bold marks the best mean.
Figure 5: Ablation on Qwen3-8B. (a) Mean accepted length by domain; (b) Training loss curve.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: One block proposal by the LongSpark drafter.
Target
Drafter
Math
Code
Chat
Overall
GSM8K
MATH-500
AIME25
MBPP
HumanEval
LCB
MT-Bench
Alpaca
Avg.
Q3-4B
EAGLE-3
5.46 (1.55)
5.23 (1.68)
4.72 (1.68)
4.12 (1.24)
4.56 (1.40)
4.92 (1.60)
2.70 (0.89)
2.48 (0.82)
4.27 (1.36)
DFlash
5.93 (2.21)
5.87 (2.54)
5.28 (2.57)
4.88 (1.90)
5.21 (2.12)
5.34 (2.25)
3.44 (1.51)
3.30 (1.43)
4.90 (2.07)
DSpark
6.33 (2.26)
6.30 (2.64)
5.77 (2.73)
5.38 (2.01)
5.68 (2.24)
5.71 (2.33)
3.84 (1.62)
3.67 (1.53)
5.34 (2.17)
LongSpark
6.14 ( 2.41 )
6.07 ( 2.85 )
5.44 ( 2.94 )
5.27 ( 2.17 )
5.37 ( 2.37 )
5.44 ( 2.47 )
3.70 ( 1.76 )
3.60 ( 1.68 )
5.13 ( 2.33 )
Q3-8B
EAGLE-3
5.58 (1.68)
5.34 (1.79)
4.88 (1.78)
4.30 (1.35)
4.68 (1.51)
5.11 (1.66)
2.92 (1.00)
2.75 (0.94)
4.44 (1.46)
Appendix
Table 3: Greedy decoding across model scales ( T=0 ). Entries report accepted length and throughput speedup using the conventions of Table 1 .
Target
Method
LongSpec 32K
CodeSpan 64K
LSWE 128K
TPS ↑
TPOT ↓
TPS ↑
TPOT ↓
TPS ↑
TPOT ↓
Q3-4B
Vanilla
1250.1
12.7
864.4
18.2
185.1
83.0
EAGLE-3
1810.5
8.7
1197.2
13.1
202.0
77.4
DFlash
1453.1
10.9
748.9
21.1
175.6
88.7
DSpark
1761.8
8.9
884.0
17.9
184.0
83.8
LongSpark
2274.6
6.9
1497.0
10.4
233.2
66.2
Appendix
Table 4: Long-context decoding on Qwen3-4B and 14B, using the evaluation protocol and reporting conventions of Table 2 .
(a) Accepted length τ↑
Variant
LongSpec 32K
CodeSpan 64K
LSWE 128K
LongSpark
2.88 ± 0.11
3.40 ± 0.11
2.53 ± 0.22
w/o global summary
2.70 ± 0.09
3.18 ± 0.11
2.25 ± 0.06
w/o boundary state
2.44 ± 0.06
2.82 ± 0.10
2.08 ± 0.10
w/o recent KV window
1.90 ± 0.07
1.77 ± 0.04
1.76 ± 0.08
(b) Serving performance
Appendix
Table 5: Long-context component ablations. All other settings follow Table 2 .
C
Drafter
Math
Code
Chat
Overall
GSM8K
MATH-500
AIME25
MBPP
HumanEval
LCB
MT-Bench
Alpaca
Avg.
8
EAGLE-3
5.23 (1.92)
4.70 (1.82)
4.01 (1.61)
3.91 (1.48)
4.34 (1.65)
4.18 (1.58)
2.72 (1.06)
2.54 (0.99)
3.95 (1.51)
DFlash
5.38 (2.68)
4.92 (2.64)
4.07 (2.27)
4.37 (2.25)
4.68 (2.44)
4.50 (2.27)
3.09 (1.66)
2.99 (1.60)
4.25 (2.23)
DSpark
6.16 ( 2.98 )
5.78 ( 3.03 )
5.00 ( 2.72 )
5.19 ( 2.59 )
5.47 ( 2.78 )
5.12 ( 2.49 )
3.67 ( 1.92 )
3.60 ( 1.87 )
5.00 ( 2.55 )
LongSpark
5.96 (2.96)
5.57 (2.98)
4.69 (2.63)
5.03 (2.54)
5.14 (2.69)
4.86 (2.44)
3.51 (1.89)
3.46 (1.82)
4.78 (2.49)
32
EAGLE-3
5.26 (1.49)
4.75 (1.50)
4.02 (1.37)
3.92 (1.16)
4.32 (1.30)
4.17 (1.28)
2.67 (0.85)
2.54 (0.82)
3.96 (1.22)
Appendix
Table 6: Concurrency scaling on Qwen3-8B ( C : concurrency). Entries follow Table 1 , with speedups relative to autoregressive decoding at the same concurrency.
Method
LongSpec 32K
CodeSpan 64K
LSWE 128K
EAGLE-3
82.61
175.67
127.39
DFlash
71.11
202.01
138.75
DSpark
63.66
153.60
116.23
LongSpark
56.14
106.77
99.71
Appendix
Table 7: Mean end-to-end completion time for 32 requests at concurrency 16, averaged over three seeds (seconds; lower is better).
Speculative decoding accelerates LLM inference by drafting multiple tokens and verifying them in parallel. Block-parallel drafters such as DFlash further improve drafting efficiency by predicting an entire block in one pass, but their position-wise predictions lack explicit intra-block causal conditioning. Recent methods such as Domino and DSpark attempt to introduce such causality into block-parallel drafting, but they require training the draft model from scratch, which limits their flexibility and increases training cost. We propose DeLS-Spec, a decoupled long-short context speculative decoding method. DeLS-Spec treats the fixed DFlash model as a long-context expert and introduces a lightweight local head as a short-context expert. The local head can be trained independently with a standard next-token prediction objective, without joint training with the target model or the DFlash backbone, leading to extremely low training cost. At inference time, DeLS-Spec combines long-context and short-context logits, and the local head is not tied to a specific DFlash checkpoint, making the method more modular and flexible. Experiments on Qwen3 models show that DeLS-Spec consistently improves speedup and average acceptance length over DFlash across math, code, and dialogue benchmarks.
Hong-Kai Zheng, Piji Li
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics, China · MIIT Key Laboratory of Pattern Analysis and Machine Intelligence, Nanjing, China · The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, Nanjing, China
Speculative decoding accelerates autoregressive large language model inference by drafting multiple tokens and verifying them in a single target-model forward pass. Recent diffusion-based drafters generate an entire block of tokens in parallel but usually commit to a single draft sequence per verification: once the first mismatch occurs, all subsequent draft tokens are discarded, resulting in a limited acceptance rate. Naively batching more draft candidate sequences only introduces a marginal improvement, as redundant or poorly placed branches increase the cost of drafting and verification without proportionally increasing the number of accepted tokens. We propose D^2SD, a dual diffusion draft speculative decoding framework that organizes candidates into a confidence-guided prefix tree, where the first diffusion drafter generates a block along with per-position confidence scores that are used to identify the most likely rejection boundary and select the top-K prefix ranges for recovery; the second variable-prefix diffusion drafter re-anchors at each selected prefix and proposes alternative continuations in one batched pass; the resulting shared-prefix candidates are jointly verified via cascade attention. Empirically, D^2SD shows clear improvements over both the underlying diffusion approach and strong autoregressive speculative decoding baselines.
Liyuan Zhang, Jiarui Zhang, Jinwei Yao +6
Peking University · Tsinghua University · HKUST +2
Speculative decoding alleviates the memory-bandwidth bottleneck in large language model inference, but its acceleration is jointly constrained by drafting overhead, token acceptance, and speculation length. We present a unified efficiency analysis showing that extending the speculation horizon can reduce rather than improve speedup when the marginal acceptance probability falls below the relative drafting cost. Guided by this analysis, we introduce SparseSpec-L, a training-free self-speculative decoding framework for long-context inference. SparseSpec-L generates lightweight drafts directly from the target model using a dynamically sparsified and recallable KV cache. It recycles per-head attention statistics produced during full-context verification as a no-extra-forward importance signal, allowing critical historical tokens to be recalled without permanently discarding the dense KV cache. An online entropy-based controller further selects the speculation length according to expected step-wise efficiency. Experiments across multiple long-context tasks and model scales show consistent end-to-end acceleration, with up to speedup over autoregressive decoding while preserving the target model's output distribution.