Adversarial perturbations can alter the predictions of frozen vision-language models (VLMs) while leaving their confidence and image--text similarity patterns seemingly plausible. We investigate whether we can identify adversarial inputs based on the broader way an image interacts with a collection of general semantic prompts. Our detector summarizes these responses using category-level statistics, relationships among prompts, deviations from clean reference distributions, and stability under weak image transformations, producing a compact response profile that is classified by a lightweight model while the VLM remains fixed. We evaluate the approach on multiple public image datasets, several CLIP-style visual backbones, and a range of gradient-based, optimization-based, automated, and spatial attacks. The detector achieves strong discrimination in attack-specific settings and retains substantial performance when evaluated on attacks not seen during training. Under a controlled detector-specific protocol, the response-profile representation outperforms the evaluated embedding-geometry baselines. Additional analyses show that the feature groups provide complementary information and that the method remains effective under variations in the prompt configuration. We also examine inference cost and performance against detector-aware adaptive attacks. Overall, the results indicate that response patterns across semantic prompts provide a useful complementary signal for adversarial image detection in frozen VLMs.
Figures & tables
Fig. 1: Conceptual example of image–text similarity computation for a CIFAR-10 test sample labeled as a ship using a CLIP-style VLM with ViT-B/16 as the backbone. The image encoder maps the input to the normalized image embedding zI(x) , while the text encoder maps the semantic prompts to text embeddings stacked in ZT . The semantic response vector is computed as s(x)=ZTzI(x) and contains one similarity value per prompt. For visualization, the figure shows only selected high-scoring prompts and their associated semantic categories rather than the complete prompt set. The full detector uses all seven semantic categories defined in the method.
Fig. 2: Example of feature-vector construction for a CIFAR-10 test image using a CLIP-style VLM with ViT-B/16 as backbone. The figure shows the actual feature groups computed by the pipeline for one sample: category statistics over the seven semantic prompt groups, prompt-graph features including graph energy and local deviations, clean-distribution residuals for the two dominant semantic categories, and stability features under small image transformations. These components are concatenated into the final feature vector ϕ(x) and passed to the MLP detector. For this example, the computed values include H(x)=1.9230 , γ(x)=0.0116 , EG(x)=1.2067 , Rtop1=4.8525 , Rtop2=6.8666 , S(x)=0.6781 , and ΔA(x)=0.1241 .
Fig. 3: Visualization of the prompt graph used in the proposed detector. Nodes correspond to semantic prompts and are colored by semantic category. Black edges connect top- r nearest neighbors in the frozen text-embedding space and are symmetrized before graph construction. Node size is proportional to graph degree. Blue guide lines connect selected prompt labels to their corresponding nodes and are included only for readability; they are not graph edges. The graph provides the fixed structure used to compute graph-based response features.
Fig. 4: Architecture of the lightweight MLP detector. The input is the whitened response-profile feature vector ϕw(x)∈Rdϕ . The detector uses three hidden layers with dimensions 256 , 128 , and 64 , each followed by ReLU activation and dropout, and a final linear layer that outputs a scalar detector score. Higher scores indicate a higher likelihood of adversarial input.
Standard
Cross-attack
Attack
ViT-B/16
ViT-B/32
ConvNeXt
ViT-B/16
ViT-B/32
ConvNeXt
CIFAR-10
PGD
98.87
97.78
98.96
–
–
–
CW
99.59
97.68
99.95
90.35
92.40
96.90
MI-FGSM
98.58
97.93
98.53
95.07
95.00
92.39
Flow
99.91
99.83
99.93
96.60
94.80
92.06
TABLE I: ROC-AUC (%) on clean + adversarial samples. Left: standard attack-specific evaluation. Right: cross-attack evaluation with detectors trained on PGD.
Fig. 5: Cross-attack ROC-AUC (%) for detectors trained on PGD adversarial examples and evaluated on unseen attack families. Results are shown across three datasets and three VLM backbones. Higher values indicate stronger transfer of the learned response-profile detector across attacks.
Dataset
Method
Latency (ms/image)
FPS
CIFAR-10
MSP
5.88
170.10
GeoDetect-KNN
6.01
166.50
GeoDetect-Mahalanobis
6.12
163.30
Proposed, one view
12.83
77.97
Proposed, four views
29.27
34.16
STL-10
MSP
5.84
171.14
TABLE III: Batch-size-one inference latency and throughput on an NVIDIA A100.
Dataset / Attack
ViT-B/16
ViT-B/32
ConvNeXt
CIFAR-10
PGD
89.95
91.68
94.04
MI-FGSM
92.59
93.86
93.79
CW
70.39
65.89
73.02
Flow
94.04
93.69
95.83
AutoAttack
94.37
94.07
94.88
TABLE IV: ROC-AUC (%) of MSP-based detection across datasets and models.
Variant
Dim.
Acc.
F1
AUC
TPR@5
Category only
30
87.08
87.30
94.69
74.10
Category + graph
38
90.75
90.82
96.71
83.95
Category + residual
32
88.12
88.05
95.29
76.45
Category + stability
32
92.10
92.05
97.78
88.25
Category + graph + residual
40
90.90
90.76
97.18
86.55
Category + graph + stability
40
93.60
93.53
98.55
92.20
TABLE V: Ablation study on CIFAR-10 with ViT-B/16 under PGD. Dim. denotes feature dimension, Acc. denotes accuracy, and TPR@5 denotes the true-positive rate at 5% false-positive rate. Acc., F1, AUC, and TPR@5 are reported in percent.