Vision-language-action (VLA) models adapt pretrained vision-language models (VLMs) for closed-loop robot control, transferring their perceptual and semantic capabilities to action prediction. Despite strong in-distribution performance, however, VLAs often degrade under deployment shifts. Adversarial training (AT) offers a model-adaptive approach to robustness without explicitly anticipating individual shifts, but its effect on natural distribution-shift generalization in multi-view VLAs remains unclear. We study this question using a multi-view VLA directly adapted from a pretrained VLM and evaluate generalization across seven LIBERO-Plus shift axes. Direct AT substantially improves Camera Viewpoint and Sensor Noise, the two shifts affecting only the third-person view, yet produces mixed or negative effects on other shifts. Controlled view interventions reveal a surprising failure mode that we term view collapse: Direct AT can shift cross-view reliance so strongly that the policy becomes dominated by the wrist view. This exposes a \textit{robustness shortcut}: apparent robustness to a shifted view can arise from reduced use of that view rather than more robust perception of it. This motivates a distinction between robust perception, extracting reliable information under within-view shifts, and robust fusion, adapting reliance across views according to their reliability. To reduce fixed view reliance, we use a simple View Swap intervention and then re-evaluate AT. With View Swap, AT further improves Camera Viewpoint, Sensor Noise, and Robot Initial State, while its effects remain mixed on other shifts. Our results show that multi-view robustness requires separating improved perception from changes in cross-view reliance, and that AT provides selective rather than generic distribution-shift benefits.
Figures & tables
Original
LIBERO-Plus
Total
Camera
Robot
Language
Light
Background
Noise
Layout
Total
SFT
92.7
39.8
55.7
64.4
84.8
82.2
53.3
71.7
62.7
Direct AT
92.9
88.4
36.9
69.2
68.2
66.7
86.8
68.6
69.6
↑0.2
↑48.6
↓18.8
↑4.8
↓16.6
↓15.5
↑33.5
↓3.1
↑6.9
Table 1: Adversarial Training (AT) preserves original-task performance and improves LIBERO-Plus performance, but its effects are highly shift-dependent. It strongly improves Camera (camera viewpoint) and Noise (sensor noise), the two shifts affecting only the third-person view, while degrading several others. Values are four-suite success rates (%).
Figure 1: Controlled interventions reveal wrist-dominant view collapse. (a) Direct AT remains functional with only the wrist view but fails with only the third-person view. (b) Graded blur confirms the same asymmetric reliance under progressive degradation. (c) Intermediate checkpoints show that wrist-dominant reliance can emerge together with task competence. All panels use LIBERO-Object.
Views
Success (%)
Multi-view
92.7
wrist only
80.6
third-person only
72.2
Table 2: Matched single-view SFT.
Figure 2: Observed robustness to perturbing one view is mechanistically ambiguous. Such robustness may arise from robust perception , robust fusion , or the failure mode of view collapse . Here, fixed reliance means reliance that does not adapt to relative view reliability.
Figure 3: Controlled view interventions reveal patterns of view reliance. Left: black-out removes either the third-person or wrist view. Right: graded blur progressively degrades either view (strengths 0, 6, and 8). The horizontal and vertical axes report success under third-person- and wrist-view corruption, respectively, with the upper-right indicating flexible fallback on either view.
Original
LIBERO-Plus
Training
Total
Camera
Robot
Language
Light
Background
Noise
Layout
Total
SFT
95.6
38.8
55.7
71.0
88.9
88.5
59.1
69.8
65.3
SFT → AT
97.8
35.8
55.9
71.8
86.2
86.8
56.8
75.7
65.0
↑2.2
↓3.0
↑0.2
↑0.8
↓2.7
↓1.7
↓2.3
↑5.8
↓0.3
SFT → Swap
96.3
77.9
57.9
70.5
87.5
88.8
85.7
77.0
77.0
↑0.7
↑39.0
↑2.3
↓0.5
↓1.4
↑0.3
↑26.6
↑7.2
↑11.7
Table 3: Comparison of robust adaptation strategies for Qwen3.5-0.8B-OFT under matched training steps. Four-suite success rates (%); green/red mark gains/losses over the SFT baseline.
Original
LIBERO-Plus
Training
Total
Camera
Robot
Language
Light
Background
Noise
Layout
Total
Direct VLM-to-VLA adaptation
Qwen3.5-0.8B-PI
SFT
95.5
27.7
51.5
70.5
80.5
61.2
54.3
71.3
58.4
SFT → Swap
97.2
43.2
56.9
55.4
85.6
47.7
77.5
73.4
62.5
↑1.7
↑15.4
↑5.4
↓15.1
↑5.1
↓13.6
↑23.1
↑2.2
↑4.1
Table 4: Comparison of adaptation strategies across model configurations. Values are four-suite success rates (%); green/red mark gains/losses over matched SFT baselines.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Configuration
VLM backbone
Qwen3.5-0.8B
Action head
OFT-style lightweight MLP head
Action chunk
8 steps
Input resolution
224×224 raw frames; internally resized to 256×256
VLM adaptation
LoRA
LoRA rank / α / dropout
32 / 64 / 0.05
Appendix
Table 5: Model and optimization configuration for the primary setting.
Experiment
Initialization
Objective
Steps
AT norm
ϵ
AT / Swap
Sec. 4 : SFT
Pretrained VLM
SFT
30k
–
–
–
Sec. 4 : Direct AT
Pretrained VLM
Fast-AT
30k
ℓ∞
1/255
100% / 0%
Sec. 5 : SFT
SFT (3,750 steps)
SFT
+10k
–
–
–
Sec. 5 : SFT → AT
SFT (3,750 steps)
Fast-AT
+10k
ℓ∞
1/255
100% / 0%
Sec. 5 : SFT → Swap
SFT (3,750 steps)
View Swap
+10k
–
–
0% / 25%
Sec. 5 : SFT → AT+Swap
SFT (3,750 steps)
Fast-AT + View Swap
+10k
ℓ∞
1/255
75% / 25%
Appendix
Table 6: Training configurations for Qwen3.5-0.8B-OFT. Sec. 4 compares SFT and Direct AT from the pretrained VLM initialization, whereas Sec. 5 compares matched continuations from a shared SFT checkpoint.
Model
Initial SFT updates
Further updates
Qwen3.5-0.8B-OFT
3,750
10,000
Qwen3.5-0.8B-PI
7,500
10,000
Qwen3.5-2B-OFT
15,000
10,000
PaliGemma-OFT
20,000
10,000
π0
30,000
10,000
Appendix
Table 7: Training schedules for the matched adaptation comparisons. All counts denote optimizer updates. Initial SFT refers to LIBERO training before the compared variants begin.
Run / Evaluation
A100-h
Training
SFT (30k steps)
88
Direct AT (30k steps)
181
SFT initialization (3,750 steps)
11
SFT → AT (+10k)
49
SFT → Swap (+10k)
28
Appendix
Table 8: Approximate computational cost for the primary Qwen3.5-0.8B-OFT experiments. Training costs are reported per run and evaluation costs per checkpoint, in A100-hours.
Category
A100-h
Training
2,167
Original LIBERO evaluation
115
LIBERO-Plus evaluation
1,536
Additional evaluations
1,669
Total
5,487
Appendix
Table 9: Approximate aggregate computational cost of the experiments.
Training
Step
Clean
Wrist black-out
3rd black-out
3rd-only-AT
1000
83.0
0.0
78.4
2000
87.4
0.0
87.6
3000
93.4
0.0
88.0
wrist-only-AT
1000
21.4
23.2
0.0
2000
52.2
51.6
0.0
3000
49.2
50.0
0.0
Appendix
Table 10: Attack placement reverses view reliance. Success rates (%) on LIBERO-Object. Each policy remains effective when the attacked view is removed, but fails when the unattacked view is removed.
Checkpoint
Δℓg
Δℓw
dˉ
t
Δℓg>Δℓw
Initialization
0.01610
0.00750
+0.00860
+21.15
86%
Direct AT, step 1000
−0.00002
0.00003
−0.00005
−2.11
45%
Direct AT, step 2000
0.00022
0.00267
−0.00245
−9.87
15%
Direct AT, step 3000
0.00037
0.00334
−0.00298
−10.76
11%
Appendix
Table 11: Per-view adversarial sensitivity before and during joint Direct AT. Positive dˉ indicates a larger loss increase from perturbing the third-person view. The initial ordering disappears by step 1000 and clearly reverses by steps 2000–3000.
Original
LIBERO-Plus
Training
Total
Camera
Robot
Language
Light
Background
Noise
Layout
Total
SFT
92.7
39.8
55.7
64.4
84.8
82.2
53.3
71.7
62.7
Direct AT ( ℓ∞ , ϵ=2/255 )
92.1
75.7
35.5
59.3
13.0
29.8
83.0
60.7
53.8
↓0.6
↑35.9
↓20.2
↓5.1
↓71.8
↓52.4
↑29.7
↓11.0
↓8.9
Direct AT ( ℓ∞ , ϵ=1/255 )
92.9
88.4
36.9
69.2
68.2
66.7
86.8
68.6
69.6
↑0.2
↑48.6
↓18.8
↑4.8
↓16.6
↓15.5
↑33.5
↓3.1
↑6.9
Appendix
Table 12: The shift-specific robustness pattern persists across adversarial threat models and perturbation budgets. Across multiple configurations, Direct AT produces pronounced gains on Camera Viewpoint and Sensor Noise, together with mixed gains and degradations across the remaining shifts. Values are four-suite success rates (%).
Figure 4: View-intervention results for additional Direct AT configurations. The two settings that substantially improve Camera Viewpoint exhibit near-complete wrist-dominant collapse, whereas ℓ2 , ϵ=1 shows a qualitatively different pattern with greater sensitivity to third-person degradation.
Figure 5: View reliance after 3k steps of Direct AT across model configurations. Success rates under third-person-view corruption (horizontal axis) and wrist-view corruption (vertical axis). Under graded blur, all configurations show greater sensitivity to wrist-view degradation, while black-out responses vary. Perturbation budgets are shown in the legend.
Original
LIBERO-Plus
Method
Total
Camera
Robot
Language
Light
Background
Noise
Layout
Total
SFT → AT+Dropout
97.3
63.6
60.4
72.3
86.1
91.0
85.1
76.9
75.4
SFT → AT+Swap
96.5
84.1
65.2
68.4
88.8
87.7
88.3
79.5
79.7
Δ (Swap − Dropout)
-0.8
+20.5
+4.8
-3.8
+2.7
-3.3
+3.1
+2.6
+4.3
Appendix
Table 13: View Swap versus matched View Dropout on LIBERO-Plus. Both variants use the same SFT initialization, adversarial-training configuration, and view-intervention frequency. Values are four-suite success rates (%).
Method
Size
Camera
Robot
Language
Light
Background
Noise
Layout
Total
Direct VLM-to-VLA adaptation on standard LIBERO
Qwen3-VL-OFT
4B
47.0
60.1
87.0
96.3
95.3
73.1
79.2
75.0
Qwen3-VL-PI
4B
64.3
57.2
82.8
94.2
94.0
79.6
78.2
77.0
Qwen3.5-0.8B-OFT (SFT)
0.8B
38.8
55.7
71.0
88.9
88.5
59.1
69.8
65.3
Qwen3.5-0.8B-OFT (SFT → AT+Swap)
0.8B
84.1
65.2
68.4
88.8
87.7
88.3
79.5
79.7
Robot-pretrained VLAs finetuned on standard LIBERO
Appendix
Table 14: Comparison with prior methods on LIBERO-Plus. Values are four-suite pooled success rates (%). Bold indicates the best result among models directly adapted from a VLM using only standard LIBERO training data.
Figure 6: Visualization of the training perturbations. Adversarial training perturbs pixels within each view, while View Swap perturbs the correspondence across views.
VLAs combine pretrained vision-language representations with action generation to enable language-guided control across diverse tasks, becoming a mainstream paradigm in embodied intelligence. However, multiple studies have reported VLA's substantial declines in task success under camera shifts, revealing a key vulnerability that limits reliable deployment. To address this vulnerability, existing methods collect paired observations of the same scene from different viewpoints to fine-tune the VLA or train visual adaptation modules. Unfortunately, they require additional data collection and VLA training costs. In this paper, we first identify \emph{cross-view coordination breakdown} under external camera shifts: the robot may rely too heavily on wrist-view cues and consequently execute subtasks in the wrong order when losing global view. Motivated by this, we propose CoRe-VLA, a plug-and-play framework requiring neither additional multi-view data collection nor VLA fine-tuning, which can incorporate with exsiting VLAs. It reconstructs a scene point cloud and renders the observation from the VLA's training viewpoint to restore cross-view coordination. In CoRe-VLA, Render-to-Camera (R2C) Restoration reduces rendering-induced visual degradation, while Execution-Trajectory-Conditioned Alignment (ETCA) reduces robot idle time and mitigates motion conflicts during asynchronous execution. Experiments on 5 real-robot tasks, LIBERO-100 and LIBERO-Plus demonstrate CoRe-VLA substantially improves task success across mainstream VLAs under camera shifts. For example, CoRe-VLA raises PI0.5's success rate from 13.3% to 83.3% at a 1.6m camera shift in real-robot environment.
Tianhang Pan, Xuanhao Wang, Yiwen Pang +4
Southeast University, Nanjing, China · National Center of Technology Innovation for EDA, Nanjing, China
Vision-Language-Action (VLA) models have shown strong capabilities in controlling robots across diverse manipulation tasks. However, their adversarial robustness remains largely underexplored, and exploiting this weakness can lead to physical-world harm. Existing attacks on VLA models often rely on pixel-space perturbations or white-box access, resulting in noticeable artifacts and limited deployability in real-world robotic systems. In this work, we propose DURA, a diffusion-based unrestricted robotic attack that generates visually natural adversarial patches for VLA models. DURA supports both white-box and black-box attack settings, where the black-box setting requires only the predicted actions of the victim model. By optimizing along the latent trajectory of a pretrained diffusion model, DURA generates visually natural patches while steering the robot toward attacker-specified target actions. Extensive experiments in both simulation and the real physical world show that DURA consistently outperforms existing methods. Our findings expose a safety risk for physically deployed VLA models and call for stronger defenses.
Jiahui Han, Yuhui Yao, Xin Wang +6
1Xi’an Jiaotong University · 2Shanghai AI Laboratory · University of Science and Technology of China
It is infeasible to encompass all possible disturbances within the training dataset. This raises a critical question regarding the robustness of Vision-Language-Action (VLA) models when encountering unseen real-world visual disturbances, particularly under imperfect visual conditions. In this work, we conduct a systematic study based on recent state-of-the-art VLA models and reveal a significant performance drop when visual disturbances absent from the training data are introduced. To mitigate this issue, we propose a lightweight adapter module grounded in information theory, termed the Information Bottleneck Adapter (IB-Adapter), which selectively filters potential noise from visual inputs. Without requiring any extra data or augmentation strategies, IB-Adapter consistently improves over the baseline by an average of 30%, while adding fewer than 10M parameters, demonstrating notable efficiency and effectiveness. Furthermore, even with a 14x smaller backbone (0.5B parameters) and no pre-training on the Open X-Embodiment dataset, our model StableVLA achieves robustness competitive with 7B-scale state-of-the-art VLAs. With negligible parameter overhead (<10M), our approach maintains accuracy on long-horizon tasks and surpasses OpenPi under both synthetic and physical visual corruptions.
Yiyang Fu, Chubin Zhang, Shukai Gong +7
Peking University · Tsinghua University · Astribot +2