Reciprocal Guidance: Orchestrating Draft and Verify Budgets for Advancing the Diffusion-AR Self-Speculation Frontier
Authors: Linye Wei, Shutian Zheng, Haoyu Zeng, Meng Li
Organizations: Institute for Artificial Intelligence, Peking University · School of Integrated Circuits, Peking University · College of Engineering, Peking University · School of Airspace Science and Engineering, Shandong University · Beijing Advanced Innovation Center for Integrated Circuits
Diffusion drafting with autoregressive (AR) verification has emerged as a promising paradigm for efficient speculative decoding. Recent self-speculation models, represented by Nemotron-Labs-Diffusion, further simplify the speculative pipeline by unifying drafting and verification within a shared backbone, while enabling longer acceptance lengths. However, the Pareto frontier between aggregate and per-request throughput remains underexplored. At low concurrency, sequential draft-verify execution requires two model forward passes per round, limiting the effective tokens per forward (TPF). By contrast, at high concurrency, longer drafts incur increasingly expensive computation, forcing individual requests to operate under constrained speculation budgets and preventing full exploitation of the full-backbone drafter. Our key observation indicates that drafting and verification exhibit reciprocal predictability. Draft logits can anticipate likely verification mismatches, while recent verification outcomes predict future drafting utility and suitable block sizes. Building on this observation, we introduce Reciprocal Guidance (RecGuide), a runtime draft-verify orchestration framework that adapts speculative decoding to varying serving loads. RecGuide exploits spare compute capacity through verification-overlapped drafting at low concurrency, while dynamically allocating request-specific draft block sizes as the workload becomes increasingly compute-intensive. Experiments across a wide range of concurrency levels demonstrate consistent throughput improvements over vanilla self-speculation, achieving up to 1.8× speedup.
Figures & tables
Figure 1: Comparison between RecGuide and existing diffusion-based speculative decoding.
Figure 2: Acceptance variability and block-dependent forward cost of self-speculation. (a) Acceptance length distribution with B=32 . (b) Forward time across concurrency levels for B∈{8,16,32} .
Figure 3: Cumulative distribution functions (CDFs) of draft confidence and top-2 probability margin during drafting rounds.
At\At+1
[1,8]
(8,16]
(16,32]
[1,8]
70.30%
18.11%
11.60%
(8,16]
58.39%
22.42%
19.19%
(16,32]
20.83%
12.51%
66.66%
Table 1: Transition probabilities of acceptance intervals between adjacent decoding rounds.
Figure 4: Overview of RecGuide runtime orchestration. Draft uncertainty guides predictive draft-verify overlap through correction and continuation branches, while verification feedback and hardware profiles guide request-specific block size allocation.
Model
Method
Metric
Math
Code
QA & IF
Overall
GSM8K
MATH-500
AIME25
HumanEval
MBPP
LCB
GPQA
IFEval
Avg.
NLD-8B
LinearSS
τ
9.94
11.21
10.49
9.71
8.15
9.78
10.62
9.57
9.93
Speedup
3.03 ×
2.95 ×
2.42 ×
2.91 ×
2.67 ×
2.36 ×
2.58 ×
2.53 ×
2.68 ×
DFlash
τ
4.26
4.31
4.90
4.28
3.78
4.64
5.11
3.90
4.40
Speedup
2.80 ×
2.87 ×
3.13 ×
2.83 ×
2.47 ×
3.09 ×
3.39 ×
2.57 ×
2.89 ×
TiDAR
τ
13.44
15.80
14.54
12.92
9.18
13.48
15.18
12.52
13.38
Table 2: Average acceptance length ( τ ) per two-forward round and decoding speedup over autoregressive decoding on NLD-8B/14B at C=1 , B=16 . RecGuide can verify two blocks in this window.
Model
Branch Outcome
Math
Code
QA & IF
Overall
NLD-8B
Correction
32.46 / 21.46
41.63 / 25.81
33.31 / 17.25
36.11 / 22.04
Continuation
43.85 / 42.92
27.89 / 27.14
48.55 / 47.88
39.04 / 38.24
No Hit
23.68
30.48
18.14
24.85
NLD-14B
Correction
32.37 / 21.98
30.99 / 20.38
35.44 / 18.61
32.62 / 20.54
Continuation
37.12 / 35.95
43.81 / 36.53
39.22 / 38.88
40.15 / 36.90
No Hit
30.51
25.20
25.33
27.22
Table 3: Branch outcomes in draft-verify overlap. Correction and Continuation report coverage / reuse rate (%); No Hit denotes rounds unmatched by any prospective branch.
Figure 5: Aggregate and per-request throughput trade-offs under different concurrency levels on (a) NLD-8B and (b) NLD-14B. We compare fixed block sizes B∈{8,16,32} with RecGuide.
Model
Concurrency
Math
Code
QA & IF
Overall
NLD-8B
2
3.61 / 1.28 / 95.11
4.63 / 1.51 / 93.86
24.97 / 5.97 / 69.07
9.33 / 2.54 / 88.13
8
0.63 / 26.91 / 72.46
0.66 / 39.18 / 60.16
12.85 / 48.73 / 38.43
3.70 / 36.96 / 59.34
32
28.83 / 52.95 / 18.22
40.45 / 47.70 / 11.86
60.76 / 22.35 / 16.89
41.17 / 43.33 / 15.50
128
72.46 / 17.25 / 10.30
82.03 / 11.11 / 6.87
81.12 / 5.65 / 13.23
78.21 / 12.04 / 9.75
NLD-14B
2
1.40 / 3.11 / 95.49
0.32 / 1.92 / 97.76
12.14 / 12.83 / 75.03
3.68 / 5.09 / 91.23
8
0.56 / 42.26 / 57.18
0.04 / 37.14 / 62.82
6.27 / 58.77 / 34.97
1.79 / 44.47 / 53.74
Table 4: Dynamic block-size distributions across serving loads. Each entry reports the percentage of requests selecting B=8/16/32 , respectively; Overall is the macro average over eight benchmarks.
Method
NLD-8B
NLD-14B
LinearSS
595
470
+ 1st Correction
799 (+34.29%)
605 (+28.72%)
+ 2nd Correction
811 (+36.30%)
612 (+30.21%)
+ Continuation
903 (+51.76%)
702 (+49.36%)
Table 5: Ablation of prospective branches in Predictive Draft-Verify Overlap at C=1 . TPS is averaged over all benchmarks.
Method
NLD-8B
NLD-14B
Fixed Block
4725
3756
+ st
5221 (+10.49%)
4041 (+7.60%)
+ zt
5254 (+11.19%)
4071 (+8.40%)
+ Bucketing
6608 (+39.84%)
4987 (+32.80%)
Table 6: Ablation of verification guided adaptive drafting. TPS is averaged over all benchmarks and concurrency levels.
Speculative decoding accelerates large language model inference by drafting multiple tokens for parallel verification, with efficiency critically determined by the speculative length selected at each decoding round. Existing dynamic speculation methods select the speculation length by estimating how many tokens will be accepted, which is reasonable for autoregressive drafters that generates tokens sequentially. The recent wave of diffusion-based drafters, however, generates candidate blocks in parallel at substantially lower drafting cost, shifting the key question from how many tokens to generate to how many generated tokens are worth verifying. We therefore reformulate dynamic speculative-length selection as expected-speedup optimization and derive a marginal criterion that extends the speculative sequence only when its acceptance gain outweighs the additional verification cost. Building on this criterion, we develop \textit{LibraSpec}, a training-free and plug-and-play algorithm that iteratively determines the speculative length using drafter confidence scores. Theoretically, we prove that LibraSpec monotonically converges toward the optimal speculative length. Experiments across six target models, three diffusion-based speculative decoding methods, and math, coding, and chat benchmarks show consistent improvements under both greedy and sampling settings, achieving a further 0.5∼1.5× improvement over baselines and up to 8.49× speedup over autoregressive decoding.
Zexun Lin, Yuan Feng, Junlin Lv +2
Suzhou Institute for Advanced Research, University of Science and Technology of China Suzhou, Jiangsu, China
Speculative decoding accelerates autoregressive large language model inference by drafting multiple tokens and verifying them in a single target-model forward pass. Recent diffusion-based drafters generate an entire block of tokens in parallel but usually commit to a single draft sequence per verification: once the first mismatch occurs, all subsequent draft tokens are discarded, resulting in a limited acceptance rate. Naively batching more draft candidate sequences only introduces a marginal improvement, as redundant or poorly placed branches increase the cost of drafting and verification without proportionally increasing the number of accepted tokens. We propose D^2SD, a dual diffusion draft speculative decoding framework that organizes candidates into a confidence-guided prefix tree, where the first diffusion drafter generates a block along with per-position confidence scores that are used to identify the most likely rejection boundary and select the top-K prefix ranges for recovery; the second variable-prefix diffusion drafter re-anchors at each selected prefix and proposes alternative continuations in one batched pass; the resulting shared-prefix candidates are jointly verified via cascade attention. Empirically, D^2SD shows clear improvements over both the underlying diffusion approach and strong autoregressive speculative decoding baselines.
Liyuan Zhang, Jiarui Zhang, Jinwei Yao +6
Peking University · Tsinghua University · HKUST +2
Block-diffusion drafters have recently emerged as a powerful alternative for speculative decoding by predicting multiple future-token distributions in a single parallel step. However, since these parallel predictions are sampled from position-wise marginals rather than fully conditioned sequences, committing to a single greedy path often fails to capture the target model's preferred trajectory. To address this, we propose BASTION, a budget-aware speculative decoding framework with tree-based diffusion drafting. Unlike existing methods that rely on static tree topologies, BASTION dynamically constructs query-dependent trees by balancing draft quality against hardware constraints. Our framework integrates three synergistic components: (1) an acceptance surrogate that estimates expected accepted length via path confidence, (2) an online latency estimator that calibrates a hardware-aware roofline model, and (3) an adaptive best-first expansion that grows the tree until marginal gains no longer justify incremental verification costs. BASTION is training-free, preserves the target model's distribution, and requires no per-setting tuning. Across diverse benchmarks and GPU architectures, BASTION achieves up to a 6.61x speedup over standard autoregressive decoding, outperforming state-of-the-art block-diffusion baselines by 39%.
Soowon Oh, Nam Cao, Yujin Kim +4
KAIST AI · Samsung Advanced Institute of Technology