Reciprocal Guidance: Orchestrating Draft and Verify Budgets for Advancing the Diffusion-AR Self-Speculation Frontier
Authors: Linye Wei, Shutian Zheng, Haoyu Zeng, Meng Li
Organizations: Institute for Artificial Intelligence, Peking University · School of Integrated Circuits, Peking University · College of Engineering, Peking University · School of Airspace Science and Engineering, Shandong University · Beijing Advanced Innovation Center for Integrated Circuits
Diffusion drafting with autoregressive (AR) verification has emerged as a promising paradigm for efficient speculative decoding. Recent self-speculation models, represented by Nemotron-Labs-Diffusion, further simplify the speculative pipeline by unifying drafting and verification within a shared backbone, while enabling longer acceptance lengths. However, the Pareto frontier between aggregate and per-request throughput remains underexplored. At low concurrency, sequential draft-verify execution requires two model forward passes per round, limiting the effective tokens per forward (TPF). By contrast, at high concurrency, longer drafts incur increasingly expensive computation, forcing individual requests to operate under constrained speculation budgets and preventing full exploitation of the full-backbone drafter. Our key observation indicates that drafting and verification exhibit reciprocal predictability. Draft logits can anticipate likely verification mismatches, while recent verification outcomes predict future drafting utility and suitable block sizes. Building on this observation, we introduce Reciprocal Guidance (RecGuide), a runtime draft-verify orchestration framework that adapts speculative decoding to varying serving loads. RecGuide exploits spare compute capacity through verification-overlapped drafting at low concurrency, while dynamically allocating request-specific draft block sizes as the workload becomes increasingly compute-intensive. Experiments across a wide range of concurrency levels demonstrate consistent throughput improvements over vanilla self-speculation, achieving up to 1.8× speedup.
Figures & tables
Figure 1: Comparison between RecGuide and existing diffusion-based speculative decoding.
Figure 2: Acceptance variability and block-dependent forward cost of self-speculation. (a) Acceptance length distribution with B=32 . (b) Forward time across concurrency levels for B∈{8,16,32} .
Figure 3: Cumulative distribution functions (CDFs) of draft confidence and top-2 probability margin during drafting rounds.
At\At+1
[1,8]
(8,16]
(16,32]
[1,8]
70.30%
18.11%
11.60%
(8,16]
58.39%
22.42%
19.19%
(16,32]
20.83%
12.51%
66.66%
Table 1: Transition probabilities of acceptance intervals between adjacent decoding rounds.
Figure 4: Overview of RecGuide runtime orchestration. Draft uncertainty guides predictive draft-verify overlap through correction and continuation branches, while verification feedback and hardware profiles guide request-specific block size allocation.
Model
Method
Metric
Math
Code
QA & IF
Overall
GSM8K
MATH-500
AIME25
HumanEval
MBPP
LCB
GPQA
IFEval
Avg.
NLD-8B
LinearSS
τ
9.94
11.21
10.49
9.71
8.15
9.78
10.62
9.57
9.93
Speedup
3.03 ×
2.95 ×
2.42 ×
2.91 ×
2.67 ×
2.36 ×
2.58 ×
2.53 ×
2.68 ×
DFlash
τ
4.26
4.31
4.90
4.28
3.78
4.64
5.11
3.90
4.40
Speedup
2.80 ×
2.87 ×
3.13 ×
2.83 ×
2.47 ×
3.09 ×
3.39 ×
2.57 ×
2.89 ×
TiDAR
τ
13.44
15.80
14.54
12.92
9.18
13.48
15.18
12.52
13.38
Table 2: Average acceptance length ( τ ) per two-forward round and decoding speedup over autoregressive decoding on NLD-8B/14B at C=1 , B=16 . RecGuide can verify two blocks in this window.
Model
Branch Outcome
Math
Code
QA & IF
Overall
NLD-8B
Correction
32.46 / 21.46
41.63 / 25.81
33.31 / 17.25
36.11 / 22.04
Continuation
43.85 / 42.92
27.89 / 27.14
48.55 / 47.88
39.04 / 38.24
No Hit
23.68
30.48
18.14
24.85
NLD-14B
Correction
32.37 / 21.98
30.99 / 20.38
35.44 / 18.61
32.62 / 20.54
Continuation
37.12 / 35.95
43.81 / 36.53
39.22 / 38.88
40.15 / 36.90
No Hit
30.51
25.20
25.33
27.22
Table 3: Branch outcomes in draft-verify overlap. Correction and Continuation report coverage / reuse rate (%); No Hit denotes rounds unmatched by any prospective branch.
Figure 5: Aggregate and per-request throughput trade-offs under different concurrency levels on (a) NLD-8B and (b) NLD-14B. We compare fixed block sizes B∈{8,16,32} with RecGuide.
Model
Concurrency
Math
Code
QA & IF
Overall
NLD-8B
2
3.61 / 1.28 / 95.11
4.63 / 1.51 / 93.86
24.97 / 5.97 / 69.07
9.33 / 2.54 / 88.13
8
0.63 / 26.91 / 72.46
0.66 / 39.18 / 60.16
12.85 / 48.73 / 38.43
3.70 / 36.96 / 59.34
32
28.83 / 52.95 / 18.22
40.45 / 47.70 / 11.86
60.76 / 22.35 / 16.89
41.17 / 43.33 / 15.50
128
72.46 / 17.25 / 10.30
82.03 / 11.11 / 6.87
81.12 / 5.65 / 13.23
78.21 / 12.04 / 9.75
NLD-14B
2
1.40 / 3.11 / 95.49
0.32 / 1.92 / 97.76
12.14 / 12.83 / 75.03
3.68 / 5.09 / 91.23
8
0.56 / 42.26 / 57.18
0.04 / 37.14 / 62.82
6.27 / 58.77 / 34.97
1.79 / 44.47 / 53.74
Table 4: Dynamic block-size distributions across serving loads. Each entry reports the percentage of requests selecting B=8/16/32 , respectively; Overall is the macro average over eight benchmarks.
Method
NLD-8B
NLD-14B
LinearSS
595
470
+ 1st Correction
799 (+34.29%)
605 (+28.72%)
+ 2nd Correction
811 (+36.30%)
612 (+30.21%)
+ Continuation
903 (+51.76%)
702 (+49.36%)
Table 5: Ablation of prospective branches in Predictive Draft-Verify Overlap at C=1 . TPS is averaged over all benchmarks.
Method
NLD-8B
NLD-14B
Fixed Block
4725
3756
+ st
5221 (+10.49%)
4041 (+7.60%)
+ zt
5254 (+11.19%)
4071 (+8.40%)
+ Bucketing
6608 (+39.84%)
4987 (+32.80%)
Table 6: Ablation of verification guided adaptive drafting. TPS is averaged over all benchmarks and concurrency levels.