cs.CROct 7, 2026

Understanding and Mitigating Token-Pruning-Induced Vulnerabilities in VLMs

Authors: Shuailong Wang, Xinyu Lyu, Shengming Yuan, Jingkuan Song, Heng Tao Shen, Lianli Gao

Organizations: University of Electronic Science and Technology of China, Chengdu, China · Southwestern University of Finance and Economics, Chengdu, China · Tongji University, Shanghai, China · Shanghai Innovation Institute, Shanghai, China

Abstract

Token-Pruning accelerates Vision-Language Models by removing redundant visual tokens, yet its safety implications remain underexplored. In this work, we present the first comprehensive safety evaluation of Token-Pruning mechanisms and find that: most pruning strategies significantly degrade safety as pruning ratios increase, whereas Query-based Compression shows the opposite, with extreme pruning (up to 99.8%), unexpectedly improves model safety. This sharp contrast prompts a key question: How do different Token-Pruning strategies reshape model safety behavior, and is it possible to enhance safety without sacrificing acceleration? To answer this, we identify an unrecognized mechanism, termed Pruning-Induced Malicious Amplification, where removal of background tokens triggers a side effect: forcing the model's attention to collapse onto a few retained malicious anchors within the foreground, inadvertently amplifying their toxic semantics under jailbreak. To address that, we propose an inference-time and plug-and-play Safety-Aware Pruning (SAP) mechanism that counteracts such dominance via three steps: (1) identifying malicious anchors, (2) restoring pruned benign tokens, and (3) reallocating excessive attention from malicious anchors to benign tokens. Extensive experiments across three safety and four utility benchmarks demonstrate that SAP mitigates pruning-induced vulnerabilities, i.e., reducing ASR by up to 62%, without compromising efficiency or utility.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models

    Apr 27, 2026Rinyoichi Takezoe, Yaqian Li, Zihao Bo +3Visual Token Pruning

  2. Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles

    Aug 5, 2026Hyeonyu Kim, Sehwan Lim, Youngwon Choi +2Visual Token PruningLens

  3. Beyond Surrogate Gradients: Fully Differentiable Token Pruning for Vision-Language Models

    May 27, 2026Landi He, Mingde Yao, Shawn Young +1Visual Token PruningRecent Vision-Language Models