Parallel drafting reduces the drafting overhead of speculative decoding for large language models (LLMs), but its gains remain limited by the accepted prefix length. Even when the correct token is present in the candidate pool, a single early selection error prevents subsequent predictions from being used. We propose DRelay, which uses global information from the entire draft block to perform prefix-aware selective repair of candidate selections before target-model verification. DRelay bases its decisions on candidate correlations and the selected path: a global reader extracts predictive information across positions for each candidate. While a causal selector combines candidate-level information extracted by the global read with the tokens selected at preceding positions to determine whether the native choice at the current position is consistent with the global evidence and the selected prefix. It then decides whether to retain or replace the token, thereby repairing early errors and extending the accepted prefix. We further jointly train the draft backbone and the selector, combining candidate-support learning with a repair objective, while weighting the repair loss according to each block position's potential contribution to the consecutive accepted prefix. Across eight diverse benchmarks on an H800 GPU, DRelay consistently improves both average acceptance length and end-to-end decoding performance over DFlash, Domino, and DSpark. Under SGLang serving, DRelay improves average end-to-end speedup over DFlash, Domino, and DSpark by 14.7%-16.8%, 8.7%-9.3%, and 8.1%-9.3%, respectively.
Figures & tables
Figure 1: End-to-end decoding speedup on Qwen3-4B. DFlash ( Chen et al., 2026 ) , Domino ( Huang et al., 2026 ) , DSpark ( Cheng et al., 2026 ) , and DRelay are compared against autoregressive decoding.
Figure 2: Candidate availability, continuation evidence, and repair value. Left: original top-16 coverage at first base/served errors ( n=7,787/5,348/4,825 ). Middle: real-minus-correctness-matched best-alternative accuracy at covered early served errors ( n=86/69 ), using served pools (post-head for Domino). Receiver and donor each have four correct subsequent predictions, an oracle condition. Intervals are 97.5% per model for two primary comparisons. Right: original-pool one-edit gain over 8,057/5,896 complete blocks, separating the repaired token from the recovered consecutive run. Left/right intervals are 95%; all intervals resample prompts within task. These panels measure repair opportunities, not trained DRelay performance.
Figure 3: DRelay reads globally and decides causally. The parallel draft predicts candidate pools from anchor A and masked slots M. Global Read combines a top-16 slot mixture with candidate-specific reads; Fusion λi supplies contextual features to the Causal selector. Prefix read accesses Pi=[x0,…,xi−1] , the selected prefix. Gate and Rank jointly choose KEEP or REPAIR: replacing C with D extends the prefix to [A,B,D] , so the next decision uses D. The target verifies the proposed path and supplies the next anchor. Anchor-feature and verified-history projections are omitted for clarity.
Math
Code
Chat
Overall
Method
GSM8K
MATH-500
AIME25
HumanEval
MBPP
LCB
MT-Bench
Alpaca
Avg.
S
τ
S
τ
S
τ
S
τ
S
τ
S
τ
S
τ
S
τ
S
τ
Temperature = 0 (greedy)
Qwen3-4B
DFlash
3.05 ×
6.40
3.99 ×
7.89
3.60 ×
7.32
3.24 ×
6.61
3.47 ×
6.01
3.76 ×
7.40
1.96 ×
4.11
1.83 ×
3.48
3.11 ×
6.15
Domino
3.98 ×
9.21
4.21 ×
9.04
3.33 ×
7.25
3.31 ×
6.92
3.62 ×
6.68
3.41 ×
6.89
2.23 ×
4.94
2.04 ×
4.30
3.27 ×
6.90
Table 1: Single-H800 system comparison on Qwen3-4B and Qwen3-8B. Each temperature reports speedup S over AR with the same target and temperature , and committed tokens per round τ . LCB denotes LiveCodeBench. The reported average is the arithmetic mean over the eight listed datasets.
Speculative decoding mitigates the latency of sequential generation in autoregressive Large Language Models (LLMs) by interleaving draft generation with target verification. However, existing parallel drafting backends often suffer from rapid accuracy degradation over long horizons, leading to high rejection rates during verification and suboptimal wall-clock speedups. We observe that drafting errors are not uniformly distributed but typically stem from localized high-uncertainty tokens that destabilize downstream generation trajectories. Motivated by this token error pattern, we propose CURE, a budget-aware dynamic repair tree designed to repair errors at uncertainty focal points without incurring prohibitive tree-verification overheads. Specifically, our method uses predictive confidence margins to dynamically locate candidate error tokens within a block-parallel draft, expands bounded repair paths only at these fragile nodes, and employs a novel repair resynchronization mechanism to realign draft states post-verification. Evaluations on code-generation benchmarks (HumanEval, MBPP, and LiveCodeBench-lite) and mathematical reasoning benchmark (GSM8K) demonstrate that CURE increases the average accepted length by 4.2-7.5% over parallel baselines without repair, translating to an end-to-end speedup of 2.66−3.49× over target-only decoding. Furthermore, we provide a plug-and-play repair module compatible with standard parallel drafting frameworks. We also characterize the trade-off between draft compute and verification efficiency.
Speculative decoding accelerates LLM inference by drafting multiple tokens and verifying them in parallel. Block-parallel drafters such as DFlash further improve drafting efficiency by predicting an entire block in one pass, but their position-wise predictions lack explicit intra-block causal conditioning. Recent methods such as Domino and DSpark attempt to introduce such causality into block-parallel drafting, but they require training the draft model from scratch, which limits their flexibility and increases training cost. We propose DeLS-Spec, a decoupled long-short context speculative decoding method. DeLS-Spec treats the fixed DFlash model as a long-context expert and introduces a lightweight local head as a short-context expert. The local head can be trained independently with a standard next-token prediction objective, without joint training with the target model or the DFlash backbone, leading to extremely low training cost. At inference time, DeLS-Spec combines long-context and short-context logits, and the local head is not tied to a specific DFlash checkpoint, making the method more modular and flexible. Experiments on Qwen3 models show that DeLS-Spec consistently improves speedup and average acceptance length over DFlash across math, code, and dialogue benchmarks.
Hong-Kai Zheng, Piji Li
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics, China · MIIT Key Laboratory of Pattern Analysis and Machine Intelligence, Nanjing, China · The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, Nanjing, China
Speculative decoding accelerates language-model inference by drafting future tokens the target model verifies in parallel. A diffusion-style drafter such as DFlash drafts an entire block in one forward pass. It is trained on the per-position marginals rather than on the joint distribution over the block, so the tokens it emits are individually plausible yet jointly incoherent. We introduce LiLiCorr, a Lightweight Likelihood-based model that Correlates the per-position marginals such a drafter produces. It keeps the top-K tokens at each position and processes them jointly, emitting an in and an out vector for each. Two candidates at consecutive positions match when the earlier out vector aligns, in cosine similarity, with the later in vector. Training scores the correct pairings highest and pushes competing ones down, so coherent blocks outscore incoherent ones. The joint distribution over the block, exponential in its length, is never materialized. One lightweight network pass produces all the vectors, the pairwise scores follow as batched matrix operations, leaving only a cheap greedy walk sequential. We co-train the DFlash drafter with LiLiCorr, so it proposes candidates that correlate into longer accepted sequences. Over the vanilla DFlash drafter it builds on, LiLiCorr accepts more and serves faster at all 72 settings we test: nine benchmarks at two target sizes under greedy and temperature-one decoding, plus a throughput sweep over six concurrencies, two input lengths and three output-entropy tiers. It raises acceptance length by 7 to 19%, while its single-pass scoring head costs only about 3% of the per-block latency. Against three concurrently developed methods that also restore coherence at draft time, all equally optimized on a common stack, LiLiCorr holds the highest throughput in 63 of those settings, ties within a measured noise floor in 6, and trails in only 3.