cs.CROct 7, 2026

Speedbumps: Rejection Attacks on Speculative Decoding

Authors: Adam Y. J. Jones, Yu Yuan, Sergio Maffeis

Organizations: Imperial College London London, UK

Abstract

Speculative decoding is a popular technique for increasing the speed and reducing the costs of large language model (LLM) inference by verifying multiple draft tokens in a single target-model forward pass. The resulting benefit depends on the ability of the drafter to approximate the target model's distribution. In this work, we study Speculative Rejection Attacks (SRAs), a novel class of attacks that cause draft and target models to disagree more often, resulting in fewer draft tokens being accepted per draft cycle. This leads to more target model forward passes needed per generated token, slowing down inference and increasing costs for the victim. We introduce two attacks which append an adversarial suffix to attacker-controlled content to degrade speculative decoding on a victim's prompts. Both attacks optimise the expected length of the accepted speculative prefix, estimating per-depth acceptance from the target's probability of the drafted proposals (Speedbump-P) or from the overlap between the draft and target distributions (Speedbump-D). In some cases, attacks degrade speculative decoding to the point of being slower than autoregressive decoding. The degradation reduces the output quality - regularisation restores output quality but gives up most of the degradation, trading effectiveness for stealthiness. Additionally, the suffixes remain effective under sampling, and transfer across drafters (Speedbump-P) or across target models sharing a drafter (Speedbump-D). These findings identify the draft-target interaction of speculative decoding as a realistic attack surface through which adversarial inputs can inflate inference costs.

Explore similar work

CardsList
  1. Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding

    May 13, 2026Shuoyang Sun, Chang Dai, Hao Fang +6Adversarial Attacks on LLMsSpeculative Decoding

  2. Adversarial Prompts for Acceptance Collapse in Speculative Decoding

    Jul 23, 2026Run Wang, Chaoyi Zhou, Xi Liu +7Adversarial Prompt GenerationSpeculative Decoding

  3. Secure Speculative Decoding for Large Language Models

    Oct 6, 2026Yichi Zhang, Zhiqi Wang, Neil Gong +1Adversarial Attacks on LLMsSpeculative Decoding