Feature-Aware Token Attack for Compression-Triggered Stealthy Failures in Large Vision-Language Models
Organizations: Beijing University of Posts and Telecommunications · Jiangnan University · Nanyang Technological University · Chongqing University of Posts and Telecommunications
Abstract
Visual-token compression improves the efficiency of large vision-language models, but can expose failures that full-token evaluation misses. We study adversarial images that preserve full-token correctness yet induce errors after compression, even when both inference paths succeed on the clean image. Creating such failures is challenging because perturbing token importance can also damage the visual content needed for full-token inference. We propose Feature-Aware Token Attack (FATA), which couples attention suppression with cosine-based feature preservation on a fixed set of salient clean-image tokens. In the primary LLaVA-1.5-7B setting, FATA uses only vision-encoder gradients, without access to the deployed compressor, token budget, or downstream task. Across four visually dependent task subsets and four compressors under a controlled reconstruction protocol, FATA achieves SR = 96.3% full-token accuracy retention and CBR = 22.1% conditional blinding, compared with 89.8% and 15.7% for CAA. Ablations support the role of both objectives in balancing compressed-path failure against full-token preservation. FATA also has the lowest measured detection rate among four attacks across three evaluated detectors at a 5% false-positive rate. These findings motivate assessing adversarial robustness jointly across full-token and compressed inference.
Figures & tables
| TextVQA-Open | VQAv2-Open | ScienceQA-MC | VQAv2-MC | |||||||||||||||||
| Method | Clean | Base | CAA | CAGE | FATA | Clean | Base | CAA | CAGE | FATA | Clean | Base | CAA | CAGE | FATA | Clean | Base | CAA | CAGE | FATA |
| Upper Bound (576 Tokens) | ||||||||||||||||||||
| None | 92.0 | 39.9 | 77.4 | 19.8 | 86.6 | 88.2 | 66.1 | 78.9 | 28.3 | 86.0 | 94.5 | 57.8 | 89.2 | 40.6 | 90.1 | 98.0 | 82.5 | 89.6 | 48.7 | 96.2 |
| Retain 192 Tokens ( ) | ||||||||||||||||||||
| VisionZIP | 89.1 | 35.7 | 82.5 | 16.2 | 83.3 | 86.6 | 61.8 | 56.2 | 25.8 | 84.2 | 91.3 | 54.5 | 86.3 | 38.2 | 86.3 | 96.9 | 79.1 | 94.8 | 45.8 | 95.1 |
| VisPruner | 90.6 | 36.1 | 83.9 | 17.9 | 83.3 | 87.4 | 65.1 | 84.9 | 26.6 | 84.3 | 91.1 | 56.5 | 87.1 | 40.2 | 88.7 | 96.8 | 79.8 | 95.1 | 45.0 | 95.0 |
| Attack | SR | CER | CBR |
| Base (Attn.-only) | 66.7 | 54.78 | 32.9 |
| CAA | 89.8 | 24.65 | 15.7 |
| CAGE | 36.6 | 75.58 | 43.1 |
| FATA | 96.3 | 30.38 | 22.1 |
| Variant | Objective | SR | CBR |
| Attention-only | 66.7 | 32.9 | |
| Semantic-only | 98.5 | 14.5 | |
| FATA Full | 96.3 | 22.1 |
| Feature Squeezing | Mahalanobis-Max | ML-ATD | |||||||
| Attack | AUROC | AUPR | TPR@5FPR | AUROC | AUPR | TPR@5FPR | AUROC | AUPR | TPR@5FPR |
| FATA | 0.560 | 0.382 | 0.024 | 0.695 | 0.509 | 0.134 | 0.828 | 0.724 | 0.497 |
| CAGE | 0.797 | 0.655 | 0.147 | 0.936 | 0.882 | 0.647 | 0.913 | 0.854 | 0.694 |
| CAA | 0.564 | 0.385 | 0.027 | 0.833 | 0.765 | 0.468 | 0.802 | 0.721 | 0.533 |
| Base | 0.665 | 0.494 | 0.064 | 0.813 | 0.660 | 0.278 | 0.952 | 0.918 | 0.801 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| .25 | .5 | .75 | 1 | 1.25 | 1.5 | 1.75 | |
| 576 | 96.0 | 95.0 | 94.3 | 94.5 | 94.2 | 95.1 | 94.3 |
| 192 | 94.1 | 92.7 | 93.4 | 93.8 | 93.8 | 92.9 | 94.0 |
| 128 | 93.0 | 91.5 | 92.1 | 92.1 | 92.8 | 91.7 | 92.7 |
| 64 | 88.9 | 87.6 | 88.4 | 88.5 | 89.4 | 88.8 | 88.7 |
| 32 | 78.8 | 80.1 | 80.7 | 80.6 | 82.0 | 81.4 | 80.9 |
| 16 | 69.4 | 68.8 | 67.3 | 69.4 | 68.1 | 69.2 | 71.0 |
| 576 | 94.1 | 95.0 | 95.2 | 95.6 | 96.0 |
| 192 | 92.0 | 92.7 | 93.3 | 94.2 | 94.3 |
| 128 | 91.8 | 91.5 | 93.1 | 93.3 | 92.2 |
| 64 | 87.8 | 87.6 | 87.5 | 89.0 | 90.1 |
| 32 | 77.4 | 80.1 | 81.1 | 81.4 | 81.0 |
| 16 | 65.4 | 68.8 | 68.8 | 70.3 | 70.7 |
| Compressor | TextVQA-Open | VQAv2-Open | ScienceQA-MC | VQAv2-MC |
| VisionZIP | 64 | 64 | 32 | 32 |
| VisPruner | 64 | 64 | 32 | 32 |
| PruMerge | 64 | 64 | 32 | 32 |
| FlowCut | 64 | 64 | 32 | 64 |
| Model | Full | Practical Open | Practical MC | Valid images per dataset |
| LLaVA-1.5-7B | or | |||
| Qwen3.5-9B | ||||
| InternVL3.5-8B-HF | or |
| Task | Task credit | Binary correctness |
| TextVQA-Open and VQAv2-Open | Normalized VQA credit against the human reference answers, with partial credit allowed. | for a normalized exact match to an accepted reference answer; otherwise. |
| ScienceQA-MC | if the parsed letter (A–F) matches the annotated answer index converted to a letter; otherwise. | The same option-letter decision as . |
| VQAv2-MC | if the parsed letter (A–D) matches the fixed target derived from the modal human answer; otherwise. | The same option-letter decision as . |
| Dataset | |||||
| TextVQA-Open | 0.70 | 5.22 | 8.60 | 4.52 | 7.90 |
| VQAv2-Open | 1.00 | 10.00 | 14.57 | 9.00 | 13.57 |
| ScienceQA-MC | -0.20 | 9.47 | 23.70 | 9.67 | 23.90 |
| VQAv2-MC | 1.00 | 8.05 | 19.60 | 7.05 | 18.60 |
| Budget | Clean | FATA | Damage (pp) | Amplification (pp) |
| 80.77 | 68.49 | 12.28 | 7.05 | |
| 70.78 | 56.20 | 14.58 | 9.34 | |
| 60.07 | 47.74 | 12.33 | 7.09 | |
| 49.97 | 41.90 | 8.06 | 2.83 |
| Compressor | Flip rate | Jaccard |
| VisionZIP | 0.854 | 0.079 |
| VisPruner | 0.925 | 0.040 |
| FlowCut | 0.890 | 0.059 |
| PruMerge | 0.617 | 0.240 |