Long-form reasoning makes inference expensive, and speculative decoding mitigates this cost by verifying multiple draft tokens in parallel. Its speedup, however, can fade as context grows and draft acceptance declines. We focus on attention-mass dilution: as softmax normalizes over more visible Keys, the mass concentrated on the highest-scoring Keys can decrease. We introduce SharpDraft, a training-free method that counteracts this effect through cardinality-aware Query scaling, without the computational overhead of online adaptation. Under explicit assumptions, we derive an exact top-k mass correction and deploy a closed-form fixed-slope approximation. Across AIME-26, GPQA-Diamond, and LongGenBench Diary, SharpDraft achieves 2.59-3.19× geometric-mean end-to-end speedups over target-only autoregressive decoding when applied to DFlash, PARD, and EAGLE 3.1. With DFlash, it improves decoding speed and outperforms full-parameter and LoRA-based online adaptation in end-to-end speedup, while matching the unmodified drafter's reported peak allocated GPU memory.
Figures & tables
Figure 1: Long-generation behavior of Qwen3-8B on LongGenBench Diary. (a) Decode throughput of different methods in tokens per second. (b) Top-32 attention mass of different methods.
Model
Method
Trainable Params ↓
Peak Mem. ↓
Online Update Cost (ms/update) ↓
Speedup ↑
AIME-26
GPQA-D
LGB(Diary)
Overall
Qwen3-4B
Target-only AR
0
12.54GiB
0
1.000×
1.000×
1.000×
1.000×
DFlash (Frozen)
0
16.68GiB
0
1.267×
1.598×
1.637×
1.491×
+ Full Adaptation
537.43M
24.46GiB
103.14
1.368×
1.981×
2.478×
1.886×
+ LoRA ( r=32 )
9.67M
19.95GiB
51.26
1.434×
1.952×
2.568×
1.930×
+ LoRA ( r=16 )
4.83M
19.84GiB
49.91
1.589×
2.183×
2.865×
2.150×
Table 1: End-to-end speedup and adaptation overhead on long-generation workloads. Speedups are relative to target-only AR; Overall is the geometric mean across the three benchmarks. Peak memory and update cost are measured on AIME-26 and GPQA-Diamond only.
Figure 2: Mean end-to-end latency and recorded GPU training time per AIME-26 request.
Method
MAL ↑
Speedup ↑
AIME-26
GPQA-D
LGB(Diary)
AIME-26
GPQA-D
LGB(Diary)
Overall
Target-only AR
–
–
–
1.000×
1.000×
1.000×
1.000×
PARD
2.368
2.728
2.252
2.691×
3.389×
2.729×
2.919×
+ LoRA
1.637
2.224
1.390
1.015×
1.567×
0.923×
1.137×
+ SharpDraft
2.441
2.906
2.424
2.773×
3.939×
2.975×
3.191×
EAGLE 3.1
3.152
3.659
5.668
1.922×
2.529×
3.515×
2.575×
Table 2: MAL and end-to-end speedup across draft model families with Qwen3-8B. Speedups are relative to target-only AR; Overall is their geometric mean across the three benchmarks.
Figure 3: Draft acceptance during multi-turn MATH-500 generation, up to 32K tokens in (a)–(c) and 100K in (d). Lines show mean session MAL; bands indicate pointwise 95% bootstrap confidence intervals.
Method
Final output length
All problems
0–10K
10–20K
20–32K
Overall MAL ↑
Frozen ( β=1 )
4.092
3.113
2.596
3.182
Fixed β=e0.1
4.273 (+0.181)
3.524 (+0.411)
3.318 (+0.723)
3.660 (+0.477)
Fixed β=e0.2
4.305 (+0.214)
3.726 (+0.613)
3.785 (+1.189)
3.925 (+0.743)
Fixed β=e0.3
4.284 (+0.193)
3.784 (+0.671)
3.989 (+1.393)
4.023 (+0.841)
Fixed β=e0.4
4.235 (+0.143)
3.756 (+0.643)
4.028 (+1.433)
4.018 (+0.835)
Table 3: Fixed versus dynamic Query scaling on Qwen3-8B/DFlash using 30 output-matched AIME-26 problems. Values are mean whole-request MAL, grouped by final output length. Parentheses show gains over Frozen.
Figure 4: Fixed and dynamic scaling during AIME-26 generation with Qwen3-8B. (a) Top-32 attention mass on matched prefixes. (b) Mean cumulative MAL relative to SharpDraft .
Figure 5: Fixed-slope sensitivity and layerwise calibration on Qwen3-8B/AIME-26. (a) MAL across slope settings. (b,c) Top-32 mass error and MAL. Shading shows pointwise 95% bootstrap confidence intervals.
Figure 6: Sensitivity to the reference cardinality across models and datasets. Verifier-call-rate savings are relative to 2K for Qwen/Llama and 4K for GLM.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Effective slope inferred from the measured reference contrast for Qwen3-8B/DFlash. The dashed line marks the deployed setting s=5 .
Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context grows, existing state-of-the-art drafters become increasingly expensive, eroding the very efficiency advantage they are designed to provide. We argue that this scaling is unnecessary. A standalone language model must grow with its prefix because it is solely responsible for every token it produces. A drafter, by contrast, only proposes candidates; the target catches and corrects every error before any token is committed. The drafter's decoding cost can therefore be made entirely independent of the prefix length. We introduce LongSpark, a block-diffusion drafter that achieves this by extracting fixed-size, multiscale views from the target's verification pass, thereby eliminating the need for a growing persistent state. Extensive evaluations demonstrate that LongSpark achieves state-of-the-art end-to-end efficiency across multiple model scales and realistic serving conditions. Notably, it delivers the lowest time-per-output-token on long-context tasks while reducing the drafter's context state by several orders of magnitude.
Speculative decoding speeds up autoregressive decoding by using a drafter to propose multiple tokens that a verifier validates in parallel. In resource-constrained deployments, the drafter uses a sparse KV cache to limit peak GPU memory and end-to-end latency under a fixed KV budget, while the verifier keeps a full KV cache. Mid-to-long context inference (4K--16K context length) is common in real applications. However, naive sparse/full speculative decoding suffers from the sparse/full mismatch as context length grows, causing the acceptance rate to drop quickly. We propose BudgetDraft, a multi-view sparse training method for sparse drafting in mid-to-long inference. The drafter is exposed to multiple sampled KV budgets during training and learns to align each sparse view with one shared full-cache teacher target. BudgetDraft combines an acceptance-aware loss on a full-cache branch with a multi-view loss on a sparse-cache branch, producing a single budget-robust drafter that recovers acceptance across sparsity levels without extra inference-time components. Experimental results on PG-19, LongBench, and LWM show that BudgetDraft achieves up to 6.55x, 4.46x, and 2.10x end-to-end speedup vs AR at 4K, 8K, and 16K context lengths, while keeping the inference pipeline memory-friendly.
Liang He, Jingbo Wen, Qishi Zhan +4
Shanghai Institute of Optics and Fine Mechanics · The University of Sydney · Marquette University +4
Speculative decoding accelerates LLM inference by drafting multiple tokens and verifying them in parallel. Block-parallel drafters such as DFlash further improve drafting efficiency by predicting an entire block in one pass, but their position-wise predictions lack explicit intra-block causal conditioning. Recent methods such as Domino and DSpark attempt to introduce such causality into block-parallel drafting, but they require training the draft model from scratch, which limits their flexibility and increases training cost. We propose DeLS-Spec, a decoupled long-short context speculative decoding method. DeLS-Spec treats the fixed DFlash model as a long-context expert and introduces a lightweight local head as a short-context expert. The local head can be trained independently with a standard next-token prediction objective, without joint training with the target model or the DFlash backbone, leading to extremely low training cost. At inference time, DeLS-Spec combines long-context and short-context logits, and the local head is not tied to a specific DFlash checkpoint, making the method more modular and flexible. Experiments on Qwen3 models show that DeLS-Spec consistently improves speedup and average acceptance length over DFlash across math, code, and dialogue benchmarks.
Hong-Kai Zheng, Piji Li
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics, China · MIIT Key Laboratory of Pattern Analysis and Machine Intelligence, Nanjing, China · The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, Nanjing, China