Feature-Aware Token Attack for Compression-Triggered Stealthy Failures in Large Vision-Language Models
Authors: Shilinlu Yan, Bowen Chen, Yuechen Zhang, Zhenhong Zhou, Li Sun, Sen Su
Organizations: Beijing University of Posts and Telecommunications · Jiangnan University · Nanyang Technological University · Chongqing University of Posts and Telecommunications
Visual-token compression improves the efficiency of large vision-language models, but can expose failures that full-token evaluation misses. We study adversarial images that preserve full-token correctness yet induce errors after compression, even when both inference paths succeed on the clean image. Creating such failures is challenging because perturbing token importance can also damage the visual content needed for full-token inference. We propose Feature-Aware Token Attack (FATA), which couples attention suppression with cosine-based feature preservation on a fixed set of salient clean-image tokens. In the primary LLaVA-1.5-7B setting, FATA uses only vision-encoder gradients, without access to the deployed compressor, token budget, or downstream task. Across four visually dependent task subsets and four compressors under a controlled reconstruction protocol, FATA achieves SR = 96.3% full-token accuracy retention and CBR = 22.1% conditional blinding, compared with 89.8% and 15.7% for CAA. Ablations support the role of both objectives in balancing compressed-path failure against full-token preservation. FATA also has the lowest measured detection rate among four attacks across three evaluated detectors at a 5% false-positive rate. These findings motivate assessing adversarial robustness jointly across full-token and compressed inference.
Figures & tables
Figure 1: FATA: attention suppression with feature preservation. A illustrates the compression-triggered failure setting; B presents the two attack objectives; C compares full-token retention and conditional blinding; D shows how answers to the same perturbed TextVQA image vary across compression paths.
Figure 2: FATA overview. Fixed clean targets guide input-only PGD with attention and semantic losses (A–C), followed by full-token and compressed-path evaluation of the same image (D).
TextVQA-Open
VQAv2-Open
ScienceQA-MC
VQAv2-MC
Method
Clean
Base
CAA
CAGE
FATA
Clean
Base
CAA
CAGE
FATA
Clean
Base
CAA
CAGE
FATA
Clean
Base
CAA
CAGE
FATA
Upper Bound (576 Tokens)
None
92.0
39.9
77.4
19.8
86.6
88.2
66.1
78.9
28.3
86.0
94.5
57.8
89.2
40.6
90.1
98.0
82.5
89.6
48.7
96.2
Retain 192 Tokens ( Kmodel=192 )
VisionZIP
89.1
35.7
82.5
16.2
83.3
86.6
61.8
56.2
25.8
84.2
91.3
54.5
86.3
38.2
86.3
96.9
79.1
94.8
45.8
95.1
VisPruner
90.6
36.1
83.9
17.9
83.3
87.4
65.1
84.9
26.6
84.3
91.1
56.5
87.1
40.2
88.7
96.8
79.8
95.1
45.0
95.0
Table 1: Accuracy (%) across four benchmarks and visual-token budgets.
Figure 3: Failure-stage decomposition over 14,316 shared Clean-full-correct cases.
Attack
SR ↑
CER ↑
CBR ↑
Base (Attn.-only)
66.7
54.78
32.9
CAA
89.8
24.65
15.7
CAGE
36.6
75.58
43.1
FATA
96.3
30.38
22.1
Table 2: Overall stealth–lethality trade-off
Variant
Objective
SR ↑
CBR ↑
Attention-only
Lattn
66.7
32.9
Semantic-only
Lsem
98.5
14.5
FATA Full
Lattn+λLsem
96.3
22.1
Table 3: Dual-objective ablation of FATA.
Figure 4: Qwen3.5 practical-budget CBR by compressor and dataset.
Feature Squeezing
Mahalanobis-Max
ML-ATD
Attack
AUROC
AUPR
TPR@5FPR
AUROC
AUPR
TPR@5FPR
AUROC
AUPR
TPR@5FPR
FATA
0.560
0.382
0.024
0.695
0.509
0.134
0.828
0.724
0.497
CAGE
0.797
0.655
0.147
0.936
0.882
0.647
0.913
0.854
0.694
CAA
0.564
0.385
0.027
0.833
0.765
0.468
0.802
0.721
0.533
Base
0.665
0.494
0.064
0.813
0.660
0.278
0.952
0.918
0.801
Table 8: Cross-attack detection performance with clean and random negatives. Higher values indicate easier detection; protocol differences and limitations are detailed in Appendix H .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
K
.25
.5
.75
1
1.25
1.5
1.75
576
96.0
95.0
94.3
94.5
94.2
95.1
94.3
192
94.1
92.7
93.4
93.8
93.8
92.9
94.0
128
93.0
91.5
92.1
92.1
92.8
91.7
92.7
64
88.9
87.6
88.4
88.5
89.4
88.8
88.7
32
78.8
80.1
80.7
80.6
82.0
81.4
80.9
16
69.4
68.8
67.3
69.4
68.1
69.2
71.0
Appendix
Table 9: FATA accuracy on unfiltered TextVQA-Open across token budgets and step sizes α .
K
λ=.5
1
2
4
8
576
94.1
95.0
95.2
95.6
96.0
192
92.0
92.7
93.3
94.2
94.3
128
91.8
91.5
93.1
93.3
92.2
64
87.8
87.6
87.5
89.0
90.1
32
77.4
80.1
81.1
81.4
81.0
16
65.4
68.8
68.8
70.3
70.7
Appendix
Table 10: FATA accuracy on unfiltered TextVQA-Open across token budgets and semantic weights λ .
Compressor
TextVQA-Open
VQAv2-Open
ScienceQA-MC
VQAv2-MC
VisionZIP
64
64
32
32
VisPruner
64
64
32
32
PruMerge
64
64
32
32
FlowCut
64
64
32
64
Appendix
Table 11: Clean-only practical token budgets Kprac(d,c) for LLaVA-1.5-7B. Each entry is the smallest K∈{16,32,64,128,192} that retains at least 80% of the corresponding full-token clean accuracy.
Model
Full
Practical Open
Practical MC
Valid images per dataset
LLaVA-1.5-7B
576
64
32 or 64
1,000
Qwen3.5-9B
Ni
r=1/9
r=1/18
1,000
InternVL3.5-8B-HF
Ni
r=1/9
r=1/18
998 or 1,000
Appendix
Table 12: Model-specific budgets and sample counts used in evaluation. Practical budgets are given separately for open-ended and multiple-choice tasks.
Task
Task credit vi
Binary correctness ci
TextVQA-Open and VQAv2-Open
Normalized VQA credit against the human reference answers, with partial credit allowed.
1 for a normalized exact match to an accepted reference answer; 0 otherwise.
ScienceQA-MC
1 if the parsed letter (A–F) matches the annotated answer index converted to a letter; 0 otherwise.
The same option-letter decision as vi .
VQAv2-MC
1 if the parsed letter (A–D) matches the fixed target derived from the modal human answer; 0 otherwise.
The same option-letter decision as vi .
Appendix
Table 13: Task-level scoring rules for reported accuracy and binary correctness. CBR uses the binary decisions in the final column.
Figure 5: Average task performance versus visual-token budget. The exact per-method values are reported in Table 1 in the main paper.
Figure 9: Loss ablation across 16 dataset–compressor settings, with paired lines and mean diamonds; shading marks full FATA. Error bars show ±1 sample standard deviation across settings
Dataset
DFull
D1/3
DPrac
A1/3
APrac
TextVQA-Open
0.70
5.22
8.60
4.52
7.90
VQAv2-Open
1.00
10.00
14.57
9.00
13.57
ScienceQA-MC
-0.20
9.47
23.70
9.67
23.90
VQAv2-MC
1.00
8.05
19.60
7.05
18.60
Appendix
Table 14: Qwen3.5 compression-specific damage and amplification by dataset (percentage points). Positive amplification means that compression increases the attack-induced accuracy gap relative to full-token inference.
Budget
Clean
FATA
Damage (pp)
Amplification (pp)
2/9
80.77
68.49
12.28
7.05
1/9
70.78
56.20
14.58
9.34
1/18
60.07
47.74
12.33
7.09
1/36
49.97
41.90
8.06
2.83
Appendix
Table 15: InternVL3.5 accuracy and amplification at additional retention fractions. Each row applies the same fraction to all four datasets and averages over the four compressors.
Compressor
Flip rate
Jaccard
VisionZIP
0.854
0.079
VisPruner
0.925
0.040
FlowCut
0.890
0.059
PruMerge
0.617
0.240
Appendix
Table 16: InternVL3.5 retained-set diagnostics at the practical budget. For each image, Sclean and SFATA contain the selected source-token indices, using representative indices for merging methods. Flip rate is ∣Sclean∖SFATA∣/∣Sclean∣ , the fraction of clean-selected tokens removed after attack; Jaccard is ∣Sclean∩SFATA∣/∣Sclean∪SFATA∣ . Values are fractions, macro-averaged by Eq. 37 .
Figure 10: ML-ATD reference fitting, feature readout, and stage-wise detection. Visual, Projector, and LLM trajectories are scored and evaluated separately, with stage metrics averaged only for reporting.