Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model
Organizations: Zhejiang University
Abstract
LLM inference is inherently expensive, even a modest slowdown can translate into substantial operating costs and severe availability risks. Recently, a growing body of research known as latency attacks focuses on crafting inputs to trigger worst-case output lengths. However, we report a contrary finding that these algorithmic-level latency attacks are largely ineffective against modern LLM serving systems. We reveal that system-level optimization such as continuous batching provides a logical isolation to mitigate contagious latency impact on co-located users. Thus, in this paper, we shift our focus from the algorithm to the system layer, and introduce a new Fill and Squeeze attack strategy targeting the state transition of the scheduler. Fill'' first exhausts the global KV cache to induce Head-of-Line blocking, while Squeeze'' forces the system into repetitive preemption. By manipulating output lengths using different attack prompts, and leveraging side-channel probing of memory status, we demonstrate that the attack can succeed in a practical black-box setting with much less cost. Extensive evaluations on vLLM indicate up to TTFT degradation relative to benign baselines and average slowdown on Time Per Output Token compared to existing attacks with 30-40% lower attack cost. Code: https://github.com/Phil-Fan/FS-attack
Figures & tables
| Target Model | Attack Method | TTFT (↑) | TTFT P99 (↑) | TPOT (↑) | TPOT P99 (↑) | Preempt# | Attack Request# (↓) | Cost ($) (↓) |
| Qwen3-8B | Benign (No Attack) | 0.15 | 0.27 | 34.78 | 49.66 | 0 | 0 | — |
| Engorgio | 0.12 | 0.24 | 48.17 | 59.51 | 0 | 4613 | 1.18 | |
| LoopLLM | 1.80 | 58.05 | 114.12 | 1249.20 | 1104 | 295 | 0.80 | |
| ExtendAttack | 0.85 | 11.90 | 80.03 | 474.01 | 538 | 529 | 0.92 | |
| F&S+plain-text | 11.05 | 400.13 | 194.45 | 2159.74 | 618 | 180 | 0.75 | |
| F&S+ExtendAttack | 35.95 | 408.86 | 181.93 | 991.19 | 524 | 73 | 0.42 |
| Model | Workload | TTFT | T | TPOT | P |
| Qwen3-8B | Alpaca | 11.05 | 75.6 | 194 | 5.6 |
| ShareGPT | 2.57 | 41.0 | 254 | 7.3 | |
| Gemma3-12B-it | Alpaca | 97.83 | 742.8 | 244 | 3.6 |
| ShareGPT | 65.89 | 720.6 | 359 | 5.3 | |
| DeepSeek-R1- D-Llama-8B | Alpaca | 5.31 | 68.2 | 90 | 2.4 |
| ShareGPT | 5.28 | 67.8 | 80 | 2.1 |
| Model | Method | TTFT | TPOT | Cost($) |
| Qwen3-8B | Engorgio | 0.10 | 45.9 | 1.23 |
| ExtendAttack | 20.11 | 191.0 | 0.75 | |
| LoopLLM | 2.11 | 64.7 | 0.95 | |
| F&S+plain-text | 21.32 | 135.2 | 0.78 | |
| F&S+ExtendAttack | 27.65 | 196.99 | 0.68 | |
| DeepSeek* | Engorgio | 0.10 | 50.19 | 1.13 |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Package | Version | Package | Version |
| Python | 3.11.0 | Transformers | 4.57.6 |
| PyTorch | 2.9.0 | vLLM | 0.11.2 |
| TorchVision | 0.24.0 | xFormers | 0.0.33.post1 |
| Torchaudio | 2.9.0 | LightGBM | 4.6.0 |
| CUDA | 13.0 | Scikit-learn | 1.7.2 |
| Accelerate | 1.12.0 | NumPy | 1.26.4 |
| Setting | TTFT (s) | TTFT P99 (s) | TPOT (ms) | TPOT P99 (ms) | Preempt# | Req# | Cost ($) |
| Full F&S | 12.124 | 403.282 | 205.90 | 2892.10 | 268 | 205 | 0.74 |
| Fill-only | 3.572 | 92.553 | 111.39 | 1122.38 | 854 | 508 | 0.81 |
| Squeeze-only | 0.095 | 0.139 | 48.94 | 58.93 | 0 | 3462 | 1.19 |
| No adaptive back-off | 0.803 | 24.608 | 78.75 | 322.67 | 583 | 448 | 0.94 |
| Non-tiered adaptive | 0.272 | 0.921 | 73.39 | 246.61 | 384 | 510 | 0.99 |
| F&S mechanism | vLLM (v0.11.2) | SGLang (commit 99c0b62 ) |
| Admission gate | can_allocate() | add_one_req() budget_state() |
| HOL-blocking trigger | Free blocks | rem_total_tokens |
| Budget type | Block count (reactive) | Token count (forward-looking) |
| Budget reservation | None | CLIP_MAX_NEW_TOKENS per req |
| Default queue order | FCFS | LPM (reverts to FCFS at queue ) |
| Long-output attack amplification | KV growth over time | KV growth plus immediate budget inflation |
| Target Model | Attack Method | TTFT (↑) | TPOT (↑) | Attack Request# (↓) | Cost ($) (↓) |
| Qwen3-8B | Engorgio | 0.09 | 48.40 | 4487 | 3.88 |
| LoopLLM | 12.65 | 55.40 | 607 | 3.01 | |
| ExtendAttack | 6.74 | 57.55 | 422 | 3.31 | |
| F&S+plain-text | 42.40 | 56.44 | 202 | 2.71 | |
| Gemma3-12B-it | Engorgio | 201.03 | 65.88 | 317 | 0.85 |
| LoopLLM | 267.30 | 70.26 | 262 | 0.99 |
| Concurrency | ITL (ms) | KV (%) | |
| C = 1 | 0.649 | 7 – 18 | 0 – 1.14 |
| C = 2 | 0.848 | 7 – 22 | 0 – 2.29 |
| C = 4 | 0.958 | 5 – 45 | 0 – 4.57 |
| C = 8 | 0.978 | 8 – 80 | 0 – 9.13 |
| C = 16 | 0.980 | 8 – 65 | 0 – 9.13 |
| Model | Acc. | F1 | Train (s) | Infer (s) |
| LightGBM | 0.87 | 0.88 | 6.94 | 0.01 |
| tsai-LSTM | 0.80 | 0.79 | 83.25 | 0.55 |
| tsai-GRU | 0.80 | 0.79 | 83.25 | 0.55 |
| tsai-Transformer | 0.75 | 0.75 | 158.91 | 0.56 |
| Setting | Accuracy | Macro-F1 |
| Controlled L40S, Qwen3-8B | 0.870 | 0.880 |
| Production H20, single replica | 0.771 | 0.765 |
| Stage | Requests (completed / dispatched) | Observed KV usage | Windows |
| Low | 320 / 320 | 0.19–1.35% | 4,031 |
| Mid | 637 / 640 | 2.43–34.73% | 153,328 |
| High | 950 / 960 | 2.96–62.58% | 179,108 |
| Total | 1,907 / 1,920 | 0.19–62.58% | 336,467 |
| Test set | Windows | Accuracy | Macro-F1 |
| Single-replica in-domain reference | 15,355 | 0.771 | 0.765 |
| Two-replica round-robin | 336,467 | 0.681 | 0.646 |
| Stage | Chunks scanned | Max ITL | 200 ms | 500 ms | 1000 ms |
| Low | 120,119 | 65.5 ms | 0 | 0 | 0 |
| Mid | 3,505,123 | 410.9 ms | 2,954 | 0 | 0 |
| High | 3,955,238 | 1,332.9 ms | 828 | 222 | 56 |
| Method | TTFT | TPOT | Preempt | Req.# | Cost |
| (ms) | (ms) | (#) | (#) | ($) | |
| Benign (No Attack) | 144.16 | 32.28 | 0 | 0 | — |
| Engorgio | 122.28 | 39.22 | 0 | 5576 | 1.25 |
| LoopLLM | 137.16 | 57.53 | 86 | 2489 | 4.16 |
| ExtendAttack | 131.05 | 45.10 | 0 | 1850 | 2.10 |
| F&S+plain-text | 126.68 | 54.35 | 27 | 1006 | 3.73 |