cs.CVSep 30, 2026

Feature-Aware Token Attack for Compression-Triggered Stealthy Failures in Large Vision-Language Models

Authors: Shilinlu Yan, Bowen Chen, Yuechen Zhang, Zhenhong Zhou, Li Sun, Sen Su

Organizations: Beijing University of Posts and Telecommunications · Jiangnan University · Nanyang Technological University · Chongqing University of Posts and Telecommunications

Abstract

Visual-token compression improves the efficiency of large vision-language models, but can expose failures that full-token evaluation misses. We study adversarial images that preserve full-token correctness yet induce errors after compression, even when both inference paths succeed on the clean image. Creating such failures is challenging because perturbing token importance can also damage the visual content needed for full-token inference. We propose Feature-Aware Token Attack (FATA), which couples attention suppression with cosine-based feature preservation on a fixed set of salient clean-image tokens. In the primary LLaVA-1.5-7B setting, FATA uses only vision-encoder gradients, without access to the deployed compressor, token budget, or downstream task. Across four visually dependent task subsets and four compressors under a controlled reconstruction protocol, FATA achieves SR = 96.3% full-token accuracy retention and CBR = 22.1% conditional blinding, compared with 89.8% and 15.7% for CAA. Ablations support the role of both objectives in balancing compressed-path failure against full-token preservation. FATA also has the lowest measured detection rate among four attacks across three evaluated detectors at a 5% false-positive rate. These findings motivate assessing adversarial robustness jointly across full-token and compressed inference.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Still There, No Longer Seen: Exposing Compression-Induced Risk in Large Vision-Language Models

    Sep 28, 2026Qiankun Li, Yuechen Zhang, Bowen Chen +4Token CompressionInference Cost

  2. CoViST: Visual Token Compression via Composable States

    Sep 27, 2026Qi Zhang, Xiandong Meng, Ronggang Wang +1Token CompressionData Compression Methods

  3. Steal the Patch Size: Adversarially Manipulate Vision-Language Models

    Jun 30, 2026Kai Hu, Akash Bharadwaj, Weichen Yu +1Visual TokenizersPreprocessing