Adversarial perturbations can alter the predictions of frozen vision-language models (VLMs) while leaving their confidence and image--text similarity patterns seemingly plausible. We investigate whether we can identify adversarial inputs based on the broader way an image interacts with a collection of general semantic prompts. Our detector summarizes these responses using category-level statistics, relationships among prompts, deviations from clean reference distributions, and stability under weak image transformations, producing a compact response profile that is classified by a lightweight model while the VLM remains fixed. We evaluate the approach on multiple public image datasets, several CLIP-style visual backbones, and a range of gradient-based, optimization-based, automated, and spatial attacks. The detector achieves strong discrimination in attack-specific settings and retains substantial performance when evaluated on attacks not seen during training. Under a controlled detector-specific protocol, the response-profile representation outperforms the evaluated embedding-geometry baselines. Additional analyses show that the feature groups provide complementary information and that the method remains effective under variations in the prompt configuration. We also examine inference cost and performance against detector-aware adaptive attacks. Overall, the results indicate that response patterns across semantic prompts provide a useful complementary signal for adversarial image detection in frozen VLMs.
Figures & tables
Fig. 1: Conceptual example of image–text similarity computation for a CIFAR-10 test sample labeled as a ship using a CLIP-style VLM with ViT-B/16 as the backbone. The image encoder maps the input to the normalized image embedding zI(x) , while the text encoder maps the semantic prompts to text embeddings stacked in ZT . The semantic response vector is computed as s(x)=ZTzI(x) and contains one similarity value per prompt. For visualization, the figure shows only selected high-scoring prompts and their associated semantic categories rather than the complete prompt set. The full detector uses all seven semantic categories defined in the method.
Fig. 2: Example of feature-vector construction for a CIFAR-10 test image using a CLIP-style VLM with ViT-B/16 as backbone. The figure shows the actual feature groups computed by the pipeline for one sample: category statistics over the seven semantic prompt groups, prompt-graph features including graph energy and local deviations, clean-distribution residuals for the two dominant semantic categories, and stability features under small image transformations. These components are concatenated into the final feature vector ϕ(x) and passed to the MLP detector. For this example, the computed values include H(x)=1.9230 , γ(x)=0.0116 , EG(x)=1.2067 , Rtop1=4.8525 , Rtop2=6.8666 , S(x)=0.6781 , and ΔA(x)=0.1241 .
Fig. 3: Visualization of the prompt graph used in the proposed detector. Nodes correspond to semantic prompts and are colored by semantic category. Black edges connect top- r nearest neighbors in the frozen text-embedding space and are symmetrized before graph construction. Node size is proportional to graph degree. Blue guide lines connect selected prompt labels to their corresponding nodes and are included only for readability; they are not graph edges. The graph provides the fixed structure used to compute graph-based response features.
Fig. 4: Architecture of the lightweight MLP detector. The input is the whitened response-profile feature vector ϕw(x)∈Rdϕ . The detector uses three hidden layers with dimensions 256 , 128 , and 64 , each followed by ReLU activation and dropout, and a final linear layer that outputs a scalar detector score. Higher scores indicate a higher likelihood of adversarial input.
Standard
Cross-attack
Attack
ViT-B/16
ViT-B/32
ConvNeXt
ViT-B/16
ViT-B/32
ConvNeXt
CIFAR-10
PGD
98.87
97.78
98.96
–
–
–
CW
99.59
97.68
99.95
90.35
92.40
96.90
MI-FGSM
98.58
97.93
98.53
95.07
95.00
92.39
Flow
99.91
99.83
99.93
96.60
94.80
92.06
TABLE I: ROC-AUC (%) on clean + adversarial samples. Left: standard attack-specific evaluation. Right: cross-attack evaluation with detectors trained on PGD.
Fig. 5: Cross-attack ROC-AUC (%) for detectors trained on PGD adversarial examples and evaluated on unseen attack families. Results are shown across three datasets and three VLM backbones. Higher values indicate stronger transfer of the learned response-profile detector across attacks.
Dataset
Method
Latency (ms/image)
FPS
CIFAR-10
MSP
5.88
170.10
GeoDetect-KNN
6.01
166.50
GeoDetect-Mahalanobis
6.12
163.30
Proposed, one view
12.83
77.97
Proposed, four views
29.27
34.16
STL-10
MSP
5.84
171.14
TABLE III: Batch-size-one inference latency and throughput on an NVIDIA A100.
Dataset / Attack
ViT-B/16
ViT-B/32
ConvNeXt
CIFAR-10
PGD
89.95
91.68
94.04
MI-FGSM
92.59
93.86
93.79
CW
70.39
65.89
73.02
Flow
94.04
93.69
95.83
AutoAttack
94.37
94.07
94.88
TABLE IV: ROC-AUC (%) of MSP-based detection across datasets and models.
Variant
Dim.
Acc.
F1
AUC
TPR@5
Category only
30
87.08
87.30
94.69
74.10
Category + graph
38
90.75
90.82
96.71
83.95
Category + residual
32
88.12
88.05
95.29
76.45
Category + stability
32
92.10
92.05
97.78
88.25
Category + graph + residual
40
90.90
90.76
97.18
86.55
Category + graph + stability
40
93.60
93.53
98.55
92.20
TABLE V: Ablation study on CIFAR-10 with ViT-B/16 under PGD. Dim. denotes feature dimension, Acc. denotes accuracy, and TPR@5 denotes the true-positive rate at 5% false-positive rate. Acc., F1, AUC, and TPR@5 are reported in percent.
Vision language models (VLMs) employ both visual and textual modalities to enable advanced vision-language inference. However, incorporating visual modalities expands the attack surface of VLMs, making them more susceptible to security threats such as adversarial perturbations and indirect prompt injection, wherein crafted malicious image prompts can elicit unintended model outputs. Existing defense methods against malicious image prompts remain insufficient as they typically demand extensive datasets for retraining or the deployment of additional, complex classifiers. Most critically, there is a profound lack of specialized defense mechanisms specifically targeting indirect prompt injections, a gap that serves as a primary motivation for this work. To address these limitations, we introduce DE-FIVE, a novel training-free framework for detecting malicious image prompts by leveraging Fourier features and the hidden state representations of the visual encoder (image vector embeddings) across perturbations. Specifically, we develop a hybrid detection strategy consisting of a black-box detector that operates on Fourier-domain features and a white-box detector that exploits image vector embeddings derived from only a few-shot malicious set. Extensive experiments demonstrate that the proposed framework consistently outperforms state-of-the-art baselines against malicious image prompts.
Vision-language models (VLMs) have advanced rapidly and are increasingly deployed in real-world applications, especially with the rise of agent-based systems. However, their safety has received relatively limited attention. Even the latest proprietary and open-weight VLMs remain highly vulnerable to adversarial attacks, leaving downstream applications exposed to significant risks. In this work, we propose a novel and lightweight adversarial attack detection framework based on sparse autoencoders (SAEs), termed SAEgis. By inserting an SAE module into a pretrained VLM and training it with standard reconstruction objectives, we find that the learned sparse latent features naturally capture attack-relevant signals. These features enable reliable classification of whether an input image has been adversarially perturbed, even for previously unseen samples. Extensive experiments show that SAEgis achieves strong performance across in-domain, cross-domain, and cross-attack settings, with particularly large improvements in cross-domain generalization compared to existing baselines. In addition, combining signals from multiple layers further improves robustness and stability. To the best of our knowledge, this is the first work to explore SAE as a plug-and-play mechanism for adversarial attack detection in VLMs. Our method requires no additional adversarial training, introduces minimal overhead, and provides a practical approach for improving the safety of real-world VLM systems.
Hao Wang, Yiqun Sun, Pengfei Wei +2
2Waseda University · 1Magellan Technology Research Institute (MTRI)
Existing adversarial attacks on vision-language models (VLMs) can steer model outputs toward attacker-specified target responses, but their effectiveness often degrades when the same perturbed input is paired with different textual queries. This paper studies cross-query response manipulation, where a single adversarial example is expected to remain effective across diverse user queries. We first analyze the limitations of existing attacks and find that successful transfer is closely associated with preserving an image-dominant attention pattern during response generation. Motivated by the observation, we propose \textbf{Attention Hijacking}, a novel adversarial attack that explicitly steers internal attention distributions toward a persistent image-dominant pattern. By amplifying the influence of visual tokens on target response tokens while suppressing the competing influence of textual tokens, our method reduces the dependence of the manipulated output on the specific wording of the query. Extensive experiments on widely used VLMs show that Attention Hijacking substantially improves cross-query transferability across diverse target responses and unseen queries. The method also extends effectively to multiple attack scenarios, offering new insights into the role of attention stability in transferable response manipulation for VLMs.
Zhiqiang Wang, Dongrui Liu, Yan Li +4
Hong Kong University of Science and Technology · Shanghai Jiao Tong University · Beihang University