Exploiting Vulnerabilities: Universal Adversarial Attacks on Vision-Language-Action Models in Robotics
Authors: Songhua Yang, Ziyu Liu, Yuanwei Liu, Xuetao Li, Xuanye Fei, He Huang, Zheng Wang, Miao Li
Organizations: School of Computer Science, Wuhan University. · School of Robotics, Wuhan University. · School of Power and Mechanical Engineering, Wuhan University.
Recently, Vision-Language-Action (VLA) models have revolutionized robotic manipulation by seamlessly integrating visual perception, language understanding, and action generation in an end-to-end learning framework. However, since these models are designed to interact directly with the physical world and humans, their security is critical, and even small vulnerabilities can lead to catastrophic failures. In this work, we propose the Universal Adversarial Object, a sphere with optimized surface texture that significantly degrades task success rates when placed within the robot's field of view. Specifically, our approach introduces a multi-level attack framework that jointly disrupts trajectory planning, task execution, and action control. We validate our method in both simulated and real-world robotic settings. Experimental results demonstrate that the adversarial object reduces the average task success rates by 31.2%-39.9% for two representative VLA models (Pi0 and RDT), with success rates dropping to near zero in complex scenarios. Index Terms--Vision-Language-Action models, adversarial attack, robotic security, universal adversarial object
Figures & tables
Fig. 1: An example of implementing an attack on VLA with universal adversarial objects. The normal inputs to the VLA include the robot status St , image observation Ii,t , and instruction l , which enable the robot to successfully perform the grasping task. However, by applying an adversarial attack, an adversarial object Op is added to these normal inputs, causing the robot’s grasping attempt to fail.
Fig. 2: The framework for generating universal adversarial objects via multi-level attack optimization. Starting from expert demonstrations, we filter successful trajectories and optimize adversarial textures using four complementary losses that target different aspects of robot behavior—from high-level task failure to low-level motion perturbations. The resulting texture pattern is mapped onto a sphere to create a physically universal adversarial object.
Task
RDT
Pi0
Easy
Hard
Easy
Hard
Clean SR
Adv. SR
Δ SR↓
TD
Clean SR
Adv. SR
Δ SR↓
TD
Clean SR
Adv. SR
Δ SR↓
TD
Clean SR
Adv. SR
Δ SR↓
TD
Adjust Bottle
83%
38%
45%
0.24
73%
5%
68%
0.29
88%
44%
44%
0.26
58%
2%
56%
0.31
Beat Block Hammer
75%
35%
40%
0.22
39%
3%
36%
0.27
45%
22%
23%
0.18
19%
0%
19%
0.25
Click Alarmclock
62%
30%
32%
0.19
10%
0%
10%
0.21
61%
28%
33%
0.20
13%
0%
13%
0.23
Grab Roller
72%
33%
39%
0.21
45%
4%
41%
0.26
94%
46%
48%
0.27
78%
8%
70%
0.32
TABLE I: Attack performance on different VLA models and tasks. SR: Success Rate, Δ SR: Success Rate Drop, TD: Trajectory Deviation. All metrics are averaged over 10 independent runs.
Fig. 3: The image of universal adversarial objects in simulator and real-world.
Method
RDT (Easy)
RDT (Hard)
Pi0 (Easy)
Pi0 (Hard)
Δ SR↓
TD
Δ SR↓
TD
Δ SR↓
TD
Δ SR↓
TD
Full Attack
32.3 %
0.19
31.2 %
0.25
39.9 %
0.24
33.8 %
0.27
w/o Ltraj
22.9%
0.12
24.8%
0.16
26.8%
0.15
26.2%
0.18
w/o Ltask
18.6%
0.18
18.2%
0.23
22.4%
0.22
18.6%
0.25
w/o Lshake
25.7%
0.14
27.9%
0.19
32.9%
0.18
30.1%
0.21
w/o Lreg
32.3%
0.20
31.2%
0.26
39.2%
0.25
33.7%
0.28
TABLE II: Ablation study on different attack components. w/o indicates that the component is removed.
Visibility
RDT (Easy)
RDT (Hard)
Pi0 (Easy)
Pi0 (Hard)
Δ SR↓
TD
Δ SR↓
TD
Δ SR↓
TD
Δ SR↓
TD
3 Cameras (All)
32.3 %
0.19
31.2 %
0.25
39.9 %
0.24
33.8 %
0.27
2 Cameras
26.1%
0.15
24.7%
0.20
31.2%
0.19
26.9%
0.22
1 Camera
15.8%
0.09
14.3%
0.12
18.6%
0.11
15.2%
0.13
Top Camera Only
12.4%
0.08
11.8%
0.11
14.9%
0.10
12.7%
0.12
Wrist Cameras Only
17.9%
0.10
16.2%
0.13
21.3%
0.12
17.4%
0.14
TABLE III: Attack performance under different camera visibility conditions. The adversarial object is placed to be visible in different numbers of cameras.
Fig. 4: Effect of Patch Size and Regularization Strength on SR Drop and Trajectory Deviation. Moderate patch sizes (3cm) maximize attack impact while increased regularization ( γ>0.1 ) significantly diminishes success rate reduction but paradoxically increases trajectory perturbation.
Task
RDT
Pi0
Clean SR
Adv. SR
Δ SR↓
Transfer Rate
Clean SR
Adv. SR
Δ SR↓
Transfer Rate
Grab Roller
55%
23%
32%
82.1%
76%
36%
40%
83.3%
Adjust Bottle
65%
28%
37%
82.2%
72%
35%
37%
84.1%
Pick Dual Bottles
28%
10%
18%
81.8%
45%
20%
25%
80.6%
Place Burger Fries
38%
15%
23%
79.3%
62%
28%
34%
81.0%
Press Stapler
30%
12%
18%
81.8%
46%
23%
23%
82.1%
TABLE IV: Real-world validation of the adversarial attack on physical Aloha robot platform. Transfer Rate indicates the percentage of attack effectiveness retained from simulation.
Vision-Language-Action (VLA) models have shown strong capabilities in controlling robots across diverse manipulation tasks. However, their adversarial robustness remains largely underexplored, and exploiting this weakness can lead to physical-world harm. Existing attacks on VLA models often rely on pixel-space perturbations or white-box access, resulting in noticeable artifacts and limited deployability in real-world robotic systems. In this work, we propose DURA, a diffusion-based unrestricted robotic attack that generates visually natural adversarial patches for VLA models. DURA supports both white-box and black-box attack settings, where the black-box setting requires only the predicted actions of the victim model. By optimizing along the latent trajectory of a pretrained diffusion model, DURA generates visually natural patches while steering the robot toward attacker-specified target actions. Extensive experiments in both simulation and the real physical world show that DURA consistently outperforms existing methods. Our findings expose a safety risk for physically deployed VLA models and call for stronger defenses.
Jiahui Han, Yuhui Yao, Xin Wang +6
1Xi’an Jiaotong University · 2Shanghai AI Laboratory · University of Science and Technology of China
Vision-language-action (VLA) models are gaining attention in robotics, yet their robustness to adversarial attacks remains largely unexplored. Existing work shows that adversarial patches can mislead VLA-based robots but assumes full access to the entire execution trajectory, an unrealistic requirement in practice. We address this limitation by formulating a partially observable threat model, where the adversary can exploit only a short prefix of the trajectory to generate a fixed patch applied to all subsequent frames. Under this setting, we propose a two-phase framework. First, we localize the patch using the model's attention maps to identify visually critical regions that correspond to the full instruction. Then, we optimize the patch to disrupt the semantic grounding of target objects and increase the curvature of action trajectories, thereby compounding failures in both perception and control. Extensive experiments in simulation and real-world robotic environments show that our method sustains adversarial effects under partial observability, inducing long-horizon disruptions and significantly reducing task success rates.
Xiaofei Wang, Mingliang Han, Tianyu Hao +3
Department of Automation, University of Science and Technology of China, Hefei, China · SmartMore Corporation · Cyberspace Institute of Advanced Technology, Guangzhou University, Guangzhou, China +2
Vision-Language-Action (VLA) models have emerged as generalist robotic policies capable of following diverse language instructions and performing a wide range of manipulation tasks. However, their direct control over embodied agents also exposes them to adversarial interference that may cause unsafe physical behaviors. Existing attacks on robotic policies are typically optimized for a single task or instruction, leaving the cross-task vulnerabilities of multitask VLAs largely unexplored. We introduce UniTexture, a cross-task universal adversarial texture attack that uses a single textured 3D object to induce targeted deviations in VLA action predictions across multiple tasks. UniTexture backpropagates gradients from the policy's action outputs to surface texture parameters through a differentiable renderer. It jointly optimizes the shared texture over a distribution of tasks, instructions, states, and viewpoints using a targeted action-space objective, steering predicted actions toward attacker-defined targets without optimizing a separate texture for each task. We evaluate UniTexture on OpenVLA and π0.5 across diverse manipulation tasks and multiple evaluation settings. UniTexture reduces the mean task success rate from 90.0% under benign conditions to 48.4% under attack, induces target-aligned action shifts, and further exhibits cross-suite and cross-model transfer without re-optimization. Together, these findings reveal shared cross-task vulnerabilities in multitask VLAs that can be systematically exploited through a single adversarial surface texture.
Yukun Dai, Mingzhe Dai, Tianshi Wang +3
Tongji University · Mohamed bin Zayed University of Artificial Intelligence · University of Electronic Science and Technology of China