Parallel speculative drafting generates multiple candidates in one backbone pass, but independent token selection can produce inconsistent continuations that shorten the accepted prefix. Existing methods mostly leave conditional decoding to a lightweight module after the backbone, which limits the flow of predecessor information to successors. Our analysis of DFlash shows that early positions already form recoverable predictions in shallow layers, and that accurate adjacent predecessors help successors more when they enter earlier. We therefore propose DSpine, a drafter with causal conditioning injection throughout the backbone: at every layer, gated adjacent injection writes each predecessor's predicted feature into its successor, so the causal conditioning chain unfolds over network depth while all positions update in parallel. A unified transfer space built from the target model's output embeddings unifies layer-wise injection with predecessor-conditioned decoding, and layer-wise output-embedding supervision promotes the formation of predicted features in shallow layers. Fused kernels and a transition cache execute both efficiently in parallel within SGLang. Across seven math, code, and chat benchmarks, DSpine achieves the longest acceptance length at both temperatures on Qwen3-4B and Qwen3-8B. At temperature zero on Qwen3-8B, it raises the seven-benchmark mean from DFlash's 3.77 to 4.82 (+27.8%); in SGLang serving tests, it delivers 23.3% higher throughput than DFlash on average.
Figures & tables
Figure 2: (a) Readout accuracy of each layer, grouped by block position. (b) Gain in consecutive correct tokens when the drafted prefix is replaced by the correct prefix at the filled (orange) layers; top: from one layer onward; bottom: at two layers, moving the earlier one. (c) Gain in successor accuracy when, at the specified layers, the adjacent predecessor’s features come from the correct token rather than its draft prediction (L1–L5: all layers); line color marks whether the history before the predecessor is correct (orange) or drafted (blue).
Figure 3: Overview of DSpine. (a) Injection passes predecessor features to successors at every layer (green); after the final layer, transition-cached decoding selects each token conditioned on its selected predecessor (orange). (b) One injection layer: predecessor and successor features jointly gate how much of the predecessor’s message enters the successor’s residual stream; training adds cosine alignment to C and, with probability p , substitutes Cyt−1 for the predecessor feature.
Figure 4: Transition cache. (a) Serial conditional scoring. (b) Parallel transition precomputation followed by sequential selection; greedy implementation shown.
Model
Method
Math
Code
Chat
Overall
GSM8K
MATH-500
HumanEval
MBPP
LCB
MT-Bench
Arena-Hard
Avg.
DFlash
4.25 / 3.96
4.23 / 3.86
3.92 / 3.59
3.93 / 3.56
3.53 / 2.99
3.27 / 2.97
3.30 / 2.80
3.77 / 3.39
Domino
4.83 / 4.52
4.81 / 4.30
4.21 / 3.79
4.31 / 3.85
3.90 / 3.23
3.58 / 3.18
3.61 / 2.99
4.18 / 3.69
Qwen3-8B
DSpark
4.99 / 4.61
4.90 / 4.48
4.43 / 4.07
4.50 / 4.12
4.01 / 3.49
3.69 / 3.37
3.69 / 3.26
4.32 / 3.92
DFlash2
5.03 / 4.56
5.10 / 4.47
4.43 / 4.04
4.52 / 4.06
4.09 / 3.50
3.71 / 3.35
3.82 / 3.29
4.39 / 3.90
DSpine (ours)
5.67 / 5.31
5.49 / 5.05
4.94 / 4.63
4.98 / 4.61
4.46 / 3.91
4.11 / 3.79
4.12 / 3.62
4.82 / 4.42
Table 1: Average acceptance length τ on Qwen3 models. Each cell gives temperature 0 / 1; bold marks the best result for each model and temperature, and underline marks the second best.
Figure 5: Serving throughput gain over DFlash on SGLang, in thousands of tokens per second (temperature zero). The dashed zero line is DFlash; shading and the strip above each panel give the lead of DSpine over DFlash2, the strongest baseline.
Figure 6: Module-level breakdown of Qwen3-8B decoding time on SGLang.
Table 7
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Draft layers L
5
Hidden size
4096
Target feature layers
{1,9,17,25,33}
Transfer-space dimension r
1024
Message dimension a
512
Candidates K
16
Appendix
Table 4: Remaining DSpine hyperparameters for Qwen3-8B.
Schedule
GSM8K gain (%)
HumanEval gain (%)
P P P P P
+0.00
+0.00
P G G G G
+6.23
+7.12
P P G G G
+1.87
+3.93
P P P P G
− 0.18
+0.26
P G P P G
+4.18
+4.50
P P G P G
+1.06
+3.09
Appendix
Table 5: Prefix-timing conditions. Schedules list layers 1–5; P/G denotes predicted/correct prefix features. Gains in K8 are relative to the all-P baseline.
Dataset
History
Position 4
Position 8
Position 12
GSM8K
Predicted
14.88
19.63
23.38
GSM8K
Correct
17.62
31.87
42.50
HumanEval
Predicted
14.02
23.63
19.51
HumanEval
Correct
17.53
38.72
42.53
Appendix
Table 6: Accuracy gain (percentage points) from correcting only the predecessor. History refers to tokens before the predecessor; positions refer to the original block.
Dataset
Position
Accuracy difference
K4 difference
GSM8K
4
7.38
0.226
GSM8K
8
17.00
0.565
GSM8K
12
18.63
0.755
HumanEval
4
11.28
0.280
HumanEval
8
15.85
0.694
HumanEval
12
18.60
0.745
Appendix
Table 7: L2+L5 minus L4+L5 with correct earlier history. Accuracy differences are in percentage points; K4 counts at most four consecutive correct candidates, excluding the anchor.
Task
Method
Qwen3-8B (concurrency)
Qwen3-4B (concurrency)
2
4
8
16
32
2
4
8
16
32
Baseline
381
747
1405
2729
4922
552
1072
2059
3794
6768
DFlash
1163 3.05 ×
2191 2.93 ×
3799 2.70 ×
5349 1.96 ×
6327 1.29 ×
1615 2.92 ×
2991 2.79 ×
5250 2.55 ×
7873 2.08 ×
9534 1.41 ×
GSM8K
Domino
1193 3.13 ×
2229 2.98 ×
3884 2.76 ×
5576 2.04 ×
6675 1.36 ×
1626 2.94 ×
3001 2.80 ×
5254 2.55 ×
7909 2.08 ×
10000 1.48 ×
DSpark
1223 3.21 ×
2280 3.05 ×
3931 2.80 ×
5656 2.07 ×
6789 1.38 ×
1663 3.01 ×
3057 2.85 ×
5296 2.57 ×
7950 2.10 ×
10002 1.48 ×
DFlash2
1327 3.49 ×
2473 3.31 ×
4234 3.01 ×
5899 2.16 ×
6975 1.42 ×
1876 3.40 ×
3460 3.23 ×
6015 2.92 ×
8807 2.32 ×
10683 1.58 ×
Appendix
Table 8: SGLang throughput (tokens/s; Figure 5 ). Green: speedup over AR; bold: best speculative method.
Speculative decoding accelerates LLM inference by drafting multiple tokens and verifying them in parallel. Block-parallel drafters such as DFlash further improve drafting efficiency by predicting an entire block in one pass, but their position-wise predictions lack explicit intra-block causal conditioning. Recent methods such as Domino and DSpark attempt to introduce such causality into block-parallel drafting, but they require training the draft model from scratch, which limits their flexibility and increases training cost. We propose DeLS-Spec, a decoupled long-short context speculative decoding method. DeLS-Spec treats the fixed DFlash model as a long-context expert and introduces a lightweight local head as a short-context expert. The local head can be trained independently with a standard next-token prediction objective, without joint training with the target model or the DFlash backbone, leading to extremely low training cost. At inference time, DeLS-Spec combines long-context and short-context logits, and the local head is not tied to a specific DFlash checkpoint, making the method more modular and flexible. Experiments on Qwen3 models show that DeLS-Spec consistently improves speedup and average acceptance length over DFlash across math, code, and dialogue benchmarks.
Hong-Kai Zheng, Piji Li
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics, China · MIIT Key Laboratory of Pattern Analysis and Machine Intelligence, Nanjing, China · The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, Nanjing, China
DSpark-style parallel drafters have made speculative decoding highly effective, yet their draft phase remains serialized on the critical path of every round. Parallel speculative decoding (PSD) overlaps drafting with verification, yet existing methods must guess the accepted prefix and bonus token in advance: a wrong guess reverts the whole batch to serial drafting. We present DPara, a PSD framework that reuses effective parallel drafters yet guarantees backbone--verification overlap in every round, thereby eliminating this probabilistic fallback altogether. While the target verifies, DPara's diffusion backbone precomputes draft representations for every acceptance boundary with the bonus left unspecified; a lightweight autoregressive head then combines the revealed verification outcome with the matching precomputed representation to emit the next round's draft tokens almost instantly---fully parallelizing the dominant backbone forward with verification and leaving only the negligible head cost serial. Experiments on Qwen3-8B and Qwen3-14B across seven math, coding, and chat benchmarks show that DPara achieves average speedups of 3.21× and 3.52× over autoregressive decoding, surpassing the strongest serial and parallel speculative decoding baselines alike.
Fuliang Liu, Xue Li, Kun Qian +4
State Key Laboratory of Novel Software Technology, Nanjing University · Alibaba Group
Speculative decoding accelerates LLM inference by drafting a tree of candidate continuations and verifying it in one target forward. Existing drafters fall into two camps with opposite weaknesses. Autoregressive drafters such as EAGLE-3 preserve dependence along each draft path but call the drafter once per tree depth, making drafting a non-trivial share of per-iteration latency. Parallel drafters cut drafter calls by predicting multiple future positions in one forward, but each position is predicted without seeing the others, producing paths the verifier rejects. In this paper, we propose SpecBlock, a block-iterative drafter that combines path dependence with cheap drafting. Each drafter forward produces K dependent positions and we call this a block. The draft tree grows through repeated block expansions. Two mechanisms explicitly carry path dependence to keep later draft positions accurate. Within each block, a layer-wise shift carries the previous position's hidden state into every decoder layer. Across blocks, each new block can start from any position of the previous block, inheriting its hidden state to extend the path. To spend verifier budget where acceptance is likely, a co-trained rank head replaces the fixed top-k tree by allocating per-position branching during drafting. To avoid training the drafter on prefixes it never produces at inference, a valid-prefix mask drops the loss at later positions once an earlier one is wrong. Beyond static drafting, a cost-aware bandit at deployment uses free verifier feedback to update the drafter selectively, only when the expected throughput gain exceeds the update cost. Experiments show that SpecBlock improves mean speedup by 8-13% over EAGLE-3 at 44-52% of its drafting cost, and cost-aware adaptation extends this lead to 11-19%.
Weijie Shi, Qiang Xu, Fan Deng +9
Hong Kong University of Science and Technology · MetaX · Zhejiang Normal University +1