The rapid integration of Vision Language Models (VLMs) into sensitive systems introduces critical safety vulnerabilities that remain unexplored in exist studies. While adversarial attack robustness has been extensively studied for image-based models, the susceptibility of VLMs to temporally-aware adversarial attacks against video in driving context poses a distinct and under examined threat. In this paper, we introduce novel adversarial attack against video targeting VLM models used for autonomous driving scenes named Spatial Temporal Coherence Adversarial Attack (STCA). Our attack comprise from three stages: modalities expansion, Spatial attack, and STCA attack. In modalities expansion, we propose caption-guided frame selection method in order to ensure that adversarial perturbation target the most semantically significant frames. Secondly.In spatial attack, we craft effective perturbation and preserve high similarity. Then the perturbed video generated fed into STCA stage that disrupt cross-frame temporal coherence using motion guided mask. Our method operate under black box threat model against victim target VLMs, relying solely on transferability from white-box surrogate model.We conduct our experiments on the BDD100K and nuScenes autonomous driving datasets across three VLM models: Video LLaVA-7B, Qwen2.5-VL-7B, and Dolphin. Experimental results demonstrate spatial attack achieves an ASR with high SSIM. Our finding reveal that existing video language model, remain highly susceptible to adversarial attack in autonomous driving scenarios, underscoring the urgent need for robust defense for VLM models.
Figures & tables
Fig. 1: Enter Caption
Fig. 2: Diversity of textual prompts
Fig. 3: Diversity of visual scenario
Fig. 4: Our proposed framework consisting from three main stages: Modalities expansion, Spatial attack and temporal attack.
Fig. 5: Caption selected strategy
Fig. 6: Sample of Mask Generation
Fig. 7: Captions resulted from Video-LLaVA model before and after attack
Fig. 8: Captions resulted from Qwen2.5-VL model before and after attack
Fig. 9: Captions resulted from Dolphins model before and after attack
Model
Spatial Attack
Spatial + Temporal Attack
Video LLaVA
32.6%
71%
Qwen2.5-VL
45%
84.2%
Dolphin
46.9%
46.2%
SSIM
0.93
0.82
TABLE I: Results of our attack in BDD100K dataset
Model
Spatial Attack
spatial + Temporal Attack
Video LLaVA
57.6%
83%
Qwen2.5-VL
64.7%
96.5%
Dolphin
37.6 %
47.1%
SSIM
0.98
0.79
TABLE II: Results of our attack in nuScene dataset
Fig. 10: SSIM demonstration between original frames and adversarial frames
Method
Video-LLaVa
Qwen2.5-VL
Dolphins
PGD
47.2%
44.8%
36.2%
FGSM
25.2%
35.4%
18.1%
Our (Spatial attack only)
32.6 %
45%
46.9%
Our (Full pipeline)
71%
84.2%
46.2%
TABLE III: Compare our proposed method with baseline methods
Dataset
Victim Model
Reference
Attack Type
ASR %
nuScene dataset
Video-LLaVA
IN et al. (2024) [ 38 ]
Physical backdoor with common object trigger
different based on trigger ranged from 65.3 % to 89.3%
Wang et al. (2026) [ 32 ]
Fine grained object-semantic attack
75%
Spatial Attack only
Transerable Spatial Attack
57.6%
Spaital+temporal
Transerable Spatial-Temporal Attack
83%
Qwen2-VL
Liu et al. (2025) [ 30 ]
Natural-reflection backdoor
70.92%
Spatial Attack only
Transerable Spatial Attack
64.7%
TABLE IV: Comparison of our framework with state of the art methods
Fig. 11: (a) The relationship between ASR and Perturbation budget. (b) The relationship between the number of ASR and the number of steps.