cs.CVOct 7, 2026

Detecting Adversarial Images through Response Profiles of Vision-Language Models

Authors: Arash Vashagh, Roozbeh Razavi-Far

Organizations: Trustworthy and Secure AI (TSAI) Lab, Faculty of Computer Science, University of New Brunswick, Canada

Abstract

Adversarial perturbations can alter the predictions of frozen vision-language models (VLMs) while leaving their confidence and image--text similarity patterns seemingly plausible. We investigate whether we can identify adversarial inputs based on the broader way an image interacts with a collection of general semantic prompts. Our detector summarizes these responses using category-level statistics, relationships among prompts, deviations from clean reference distributions, and stability under weak image transformations, producing a compact response profile that is classified by a lightweight model while the VLM remains fixed. We evaluate the approach on multiple public image datasets, several CLIP-style visual backbones, and a range of gradient-based, optimization-based, automated, and spatial attacks. The detector achieves strong discrimination in attack-specific settings and retains substantial performance when evaluated on attacks not seen during training. Under a controlled detector-specific protocol, the response-profile representation outperforms the evaluated embedding-geometry baselines. Additional analyses show that the feature groups provide complementary information and that the method remains effective under variations in the prompt configuration. We also examine inference cost and performance against detector-aware adaptive attacks. Overall, the results indicate that response patterns across semantic prompts provide a useful complementary signal for adversarial image detection in frozen VLMs.

Figures & tables

Explore similar work

CardsList
  1. Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs

    May 8, 2026Hao Wang, Yiqun Sun, Pengfei Wei +2VLM RobustnessAdversarial Attacks on VLMs

  2. Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models

    May 17, 2026Zhiqiang Wang, Dongrui Liu, Yan Li +4Adversarial Attacks on VLMsAdversarial Attacks