cs.CRFeb 8, 2026

Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model

Authors: Tianyi Wang, Huawei Fan, Yuanchao Shu, Peng Cheng, Cong Wang

Organizations: Zhejiang University

Abstract

LLM inference is inherently expensive, even a modest slowdown can translate into substantial operating costs and severe availability risks. Recently, a growing body of research known as latency attacks focuses on crafting inputs to trigger worst-case output lengths. However, we report a contrary finding that these algorithmic-level latency attacks are largely ineffective against modern LLM serving systems. We reveal that system-level optimization such as continuous batching provides a logical isolation to mitigate contagious latency impact on co-located users. Thus, in this paper, we shift our focus from the algorithm to the system layer, and introduce a new Fill and Squeeze attack strategy targeting the state transition of the scheduler. Fill'' first exhausts the global KV cache to induce Head-of-Line blocking, while Squeeze'' forces the system into repetitive preemption. By manipulating output lengths using different attack prompts, and leveraging side-channel probing of memory status, we demonstrate that the attack can succeed in a practical black-box setting with much less cost. Extensive evaluations on vLLM indicate up to 75−742×75-742\times TTFT degradation relative to benign baselines and 1.5−4×1.5-4\times average slowdown on Time Per Output Token compared to existing attacks with 30-40% lower attack cost. Code: https://github.com/Phil-Fan/FS-attack

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Continuous Discovery of Vulnerabilities in LLM Serving Systems with Fuzzing

    May 11, 2026Yunze Zhao, Yibo Zhao, Yuchen Zhang +2LLM ServingLLM Security

  2. LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving

    Aug 9, 2026Shuowei Jin, Xueshen Liu, Jiaxin Shan +4LLM Serving

  3. Token Latency Fairness: Performance Isolation for Multi-Tenant LLM Serving

    Sep 16, 2026Dev Bali, Soujanya Ponnapalli, Yichuan Wang +3LLM ServingLLM Inference Scheduling