cs.AISep 27, 2026

Does Adversarial Training Improve Generalization in Multi-View VLAs? Revealing and Mitigating View Collapse

Authors: Futa Waseda, Shuhei Kurita, Isao Echizen

Organizations: National Institute of Informatics Tokyo, Japan

Abstract

Vision-language-action (VLA) models adapt pretrained vision-language models (VLMs) for closed-loop robot control, transferring their perceptual and semantic capabilities to action prediction. Despite strong in-distribution performance, however, VLAs often degrade under deployment shifts. Adversarial training (AT) offers a model-adaptive approach to robustness without explicitly anticipating individual shifts, but its effect on natural distribution-shift generalization in multi-view VLAs remains unclear. We study this question using a multi-view VLA directly adapted from a pretrained VLM and evaluate generalization across seven LIBERO-Plus shift axes. Direct AT substantially improves Camera Viewpoint and Sensor Noise, the two shifts affecting only the third-person view, yet produces mixed or negative effects on other shifts. Controlled view interventions reveal a surprising failure mode that we term view collapse: Direct AT can shift cross-view reliance so strongly that the policy becomes dominated by the wrist view. This exposes a \textit{robustness shortcut}: apparent robustness to a shifted view can arise from reduced use of that view rather than more robust perception of it. This motivates a distinction between robust perception, extracting reliable information under within-view shifts, and robust fusion, adapting reliance across views according to their reliability. To reduce fixed view reliance, we use a simple View Swap intervention and then re-evaluate AT. With View Swap, AT further improves Camera Viewpoint, Sensor Noise, and Robot Initial State, while its effects remain mixed on other shifts. Our results show that multi-view robustness requires separating improved perception from changes in cross-view reliance, and that AT provides selective rather than generic distribution-shift benefits.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CoRe-VLA: Preserving Cross-View Coordination in VLAs under Camera Shifts

    Sep 29, 2026Tianhang Pan, Xuanhao Wang, Yiwen Pang +4Cross-ViewPhotometric Supervision

  2. Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action Models

    Aug 11, 2026Jiahui Han, Yuhui Yao, Xin Wang +6Diffusion-Based Vision-Language-Actions

  3. StableVLA: Towards Robust Vision-Language-Action Models without Extra Data

    May 18, 2026Yiyang Fu, Chubin Zhang, Shukai Gong +7Diffusion-Based Vision-Language-ActionsVision-Language Model Adaptation