Nucleus Speculative Decoding: Plausibility-Aware Verification Beyond Exact Distribution
Organizations: EIT-NLP Lab, Eastern Institute of Technology, Ningbo · The Hong Kong Polytechnic University · Independent Researcher
Abstract
Speculative decoding accelerates autoregressive generation by using a lightweight draft model to propose multiple tokens that are verified by a target model in parallel. However, the standard acceptance rule focuses on exact distribution correction and rejects tokens that remain highly plausible under the target model when the draft model assigns excess probability. This conservative verification limits the number of draft tokens retained after each verification forward pass. We introduce Nucleus Speculative Decoding (NSD), a relaxed verification method that incorporates target-model plausibility into speculative decoding. NSD accepts a draft token if it satisfies the standard acceptance rule or belongs to the target model's nucleus. We theoretically characterize the distributional deviation introduced by our method and show that the single-step error is exactly determined by the draft model's excess probability within the target nucleus. We further derive sequence-level fidelity bounds that quantify how local deviations accumulate over autoregressive decoding. Experiments across multiple target models and proposal mechanisms demonstrate that NSD consistently improves speculative decoding efficiency while maintaining competitive task performance. Our method achieves throughput speedups of up to over autoregressive decoding and up to over standard speculative decoding. These improvements coincide with longer accepted lengths, allowing more output tokens to share the cost of each target verification pass. Analysis shows that plausibility-aware verification provides an effective approach for relaxed verification and speculative decoding efficiency. Our code is available at https://github.com/EIT-NLP/Nucleus-Speculative-Decoding.
Figures & tables
| Nucleus threshold | ||||||||||||||||
| 0.95 | 0.925 | 0.90 | 0.875 | 0.85 | ||||||||||||
| Dataset | Spd. | Acc. | Spd. | Acc. | Spd. | Acc. | Spd. | Acc. | Spd. | Acc. | ||||||
| Nemo-8B | ||||||||||||||||
| GSM8K | 8 | 6.61 | 91.66 | 6.61 | 91.66 | 6.67 | 91.66 | 6.51 | 92.42 | 6.54 | 92.87 | |||||
| 16 | 10.70 | 91.58 | 9.42 | 93.48 | 9.66 | 93.94 | 9.56 | 93.10 | 9.46 | 92.72 | ||||||
| 24 | 11.90 | 92.87 | 11.74 | 90.90 | 12.09 | 90.14 | 12.03 | 90.98 | 12.47 | 91.58 | ||||||
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Selection or intervention | Recovery and scope |
|---|---|---|
| SD | Stochastic ratio acceptance . | Sample from normalized at rejection; preserves the chosen target law under the usual assumptions. |
| NSD | Retain ratio acceptance and accept in-nucleus failures. | Keep the same residual; local added acceptance and TV error equal . No request-level budget. |
| ASD [ 10 ] | Greedy exceptions constrained by logit regret, a block cap, and a persistent request budget. | Target argmax recovery or bonus under the committed prefix. Zero budget restores greedy verification; regret accounting is not an NSD TV bound. |
| MARS [ 23 ] | Accept the target top-1 token, or its top-2 token when the raw-logit ratio . | On failure, emit the target top-1 token and discard the remaining draft. The stated rule relaxes token matching; it does not use the stochastic residual. |
| Cactus [ 13 ] | Select an adjusted distribution under a divergence constraint. | Uses generalized acceptance/recovery equations. NSD instead fixes the original residual and analyzes the induced deviation. |
| Acc. (%) | |||||
|---|---|---|---|---|---|
| Proposal | Benchmark | SD | NSD | SD | NSD |
| DFlash-b16 | MATH500 | 83.88 | 84.28 | 5.75 | 6.32 |
| AIME25 | 17.33 | 18.67 | 4.61 | 5.13 | |
| HumanEval | 84.51 | 83.17 | 5.48 | 6.20 | |
| MBPP | 62.41 | 63.50 | 4.77 | 5.29 | |
| EAGLE-3-TTT7 | MATH500 | 83.64 | 83.32 | 4.54 | 5.32 |
| Proposal | Benchmark | SD | NSD |
|---|---|---|---|
| DFlash-b16 | MATH500 | 115.92 | 140.77 |
| AIME25 | 87.48 | 115.15 | |
| HumanEval | 118.54 | 141.18 | |
| MBPP | 102.36 | 121.65 | |
| EAGLE-3-TTT7 | MATH500 | 157.25 | 188.61 |
| AIME25 | 115.63 | 155.62 |
| Acc. (%) | ||||
|---|---|---|---|---|
| Benchmark | SD | NSD | SD | NSD |
| MATH500 | 92.73 | 92.33 | 2.63 | 2.95 |
| AIME25 | 35.56 | 34.44 | 2.57 | 2.94 |
| HumanEval | 78.25 | 80.89 | 2.55 | 2.95 |
| MBPP | 55.12 | 57.85 | 2.42 | 2.92 |