Does Adversarial Training Improve Generalization in Multi-View VLAs? Revealing and Mitigating View Collapse
Organizations: National Institute of Informatics Tokyo, Japan
Abstract
Vision-language-action (VLA) models adapt pretrained vision-language models (VLMs) for closed-loop robot control, transferring their perceptual and semantic capabilities to action prediction. Despite strong in-distribution performance, however, VLAs often degrade under deployment shifts. Adversarial training (AT) offers a model-adaptive approach to robustness without explicitly anticipating individual shifts, but its effect on natural distribution-shift generalization in multi-view VLAs remains unclear. We study this question using a multi-view VLA directly adapted from a pretrained VLM and evaluate generalization across seven LIBERO-Plus shift axes. Direct AT substantially improves Camera Viewpoint and Sensor Noise, the two shifts affecting only the third-person view, yet produces mixed or negative effects on other shifts. Controlled view interventions reveal a surprising failure mode that we term view collapse: Direct AT can shift cross-view reliance so strongly that the policy becomes dominated by the wrist view. This exposes a \textit{robustness shortcut}: apparent robustness to a shifted view can arise from reduced use of that view rather than more robust perception of it. This motivates a distinction between robust perception, extracting reliable information under within-view shifts, and robust fusion, adapting reliance across views according to their reliability. To reduce fixed view reliance, we use a simple View Swap intervention and then re-evaluate AT. With View Swap, AT further improves Camera Viewpoint, Sensor Noise, and Robot Initial State, while its effects remain mixed on other shifts. Our results show that multi-view robustness requires separating improved perception from changes in cross-view reliance, and that AT provides selective rather than generic distribution-shift benefits.
Figures & tables
| Original | LIBERO-Plus | ||||||||
| Total | Camera | Robot | Language | Light | Background | Noise | Layout | Total | |
| SFT | 92.7 | 39.8 | 55.7 | 64.4 | 84.8 | 82.2 | 53.3 | 71.7 | 62.7 |
| Direct AT | 92.9 | 88.4 | 36.9 | 69.2 | 68.2 | 66.7 | 86.8 | 68.6 | 69.6 |
| Views | Success (%) |
|---|---|
| Multi-view | 92.7 |
| wrist only | 80.6 |
| third-person only | 72.2 |
| Original | LIBERO-Plus | ||||||||
| Training | Total | Camera | Robot | Language | Light | Background | Noise | Layout | Total |
| SFT | 95.6 | 38.8 | 55.7 | 71.0 | 88.9 | 88.5 | 59.1 | 69.8 | 65.3 |
| SFT AT | 97.8 | 35.8 | 55.9 | 71.8 | 86.2 | 86.8 | 56.8 | 75.7 | 65.0 |
| SFT Swap | 96.3 | 77.9 | 57.9 | 70.5 | 87.5 | 88.8 | 85.7 | 77.0 | 77.0 |
| Original | LIBERO-Plus | ||||||||
| Training | Total | Camera | Robot | Language | Light | Background | Noise | Layout | Total |
| Direct VLM-to-VLA adaptation | |||||||||
| Qwen3.5-0.8B-PI | |||||||||
| SFT | 95.5 | 27.7 | 51.5 | 70.5 | 80.5 | 61.2 | 54.3 | 71.3 | 58.4 |
| SFT Swap | 97.2 | 43.2 | 56.9 | 55.4 | 85.6 | 47.7 | 77.5 | 73.4 | 62.5 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Configuration |
|---|---|
| VLM backbone | Qwen3.5-0.8B |
| Action head | OFT-style lightweight MLP head |
| Action chunk | 8 steps |
| Input resolution | raw frames; internally resized to |
| VLM adaptation | LoRA |
| LoRA rank / / dropout | 32 / 64 / 0.05 |
| Experiment | Initialization | Objective | Steps | AT norm | AT / Swap | |
|---|---|---|---|---|---|---|
| Sec. 4 : SFT | Pretrained VLM | SFT | 30k | – | – | – |
| Sec. 4 : Direct AT | Pretrained VLM | Fast-AT | 30k | 100% / 0% | ||
| Sec. 5 : SFT | SFT (3,750 steps) | SFT | +10k | – | – | – |
| Sec. 5 : SFT AT | SFT (3,750 steps) | Fast-AT | +10k | 100% / 0% | ||
| Sec. 5 : SFT Swap | SFT (3,750 steps) | View Swap | +10k | – | – | 0% / 25% |
| Sec. 5 : SFT AT+Swap | SFT (3,750 steps) | Fast-AT + View Swap | +10k | 75% / 25% |
| Model | Initial SFT updates | Further updates |
|---|---|---|
| Qwen3.5-0.8B-OFT | 3,750 | 10,000 |
| Qwen3.5-0.8B-PI | 7,500 | 10,000 |
| Qwen3.5-2B-OFT | 15,000 | 10,000 |
| PaliGemma-OFT | 20,000 | 10,000 |
| 30,000 | 10,000 |
| Run / Evaluation | A100-h |
| Training | |
| SFT (30k steps) | 88 |
| Direct AT (30k steps) | 181 |
| SFT initialization (3,750 steps) | 11 |
| SFT AT (+10k) | 49 |
| SFT Swap (+10k) | 28 |
| Category | A100-h |
|---|---|
| Training | 2,167 |
| Original LIBERO evaluation | 115 |
| LIBERO-Plus evaluation | 1,536 |
| Additional evaluations | 1,669 |
| Total | 5,487 |
| Training | Step | Clean | Wrist black-out | 3rd black-out |
|---|---|---|---|---|
| 3rd-only-AT | 1000 | 83.0 | 0.0 | 78.4 |
| 2000 | 87.4 | 0.0 | 87.6 | |
| 3000 | 93.4 | 0.0 | 88.0 | |
| wrist-only-AT | 1000 | 21.4 | 23.2 | 0.0 |
| 2000 | 52.2 | 51.6 | 0.0 | |
| 3000 | 49.2 | 50.0 | 0.0 |
| Checkpoint | |||||
|---|---|---|---|---|---|
| Initialization | |||||
| Direct AT, step 1000 | |||||
| Direct AT, step 2000 | |||||
| Direct AT, step 3000 |
| Original | LIBERO-Plus | ||||||||
| Training | Total | Camera | Robot | Language | Light | Background | Noise | Layout | Total |
| SFT | 92.7 | 39.8 | 55.7 | 64.4 | 84.8 | 82.2 | 53.3 | 71.7 | 62.7 |
| Direct AT ( , ) | 92.1 | 75.7 | 35.5 | 59.3 | 13.0 | 29.8 | 83.0 | 60.7 | 53.8 |
| Direct AT ( , ) | 92.9 | 88.4 | 36.9 | 69.2 | 68.2 | 66.7 | 86.8 | 68.6 | 69.6 |
| Original | LIBERO-Plus | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | Total | Camera | Robot | Language | Light | Background | Noise | Layout | Total |
| SFT AT+Dropout | 97.3 | 63.6 | 60.4 | 72.3 | 86.1 | 91.0 | 85.1 | 76.9 | 75.4 |
| SFT AT+Swap | 96.5 | 84.1 | 65.2 | 68.4 | 88.8 | 87.7 | 88.3 | 79.5 | 79.7 |
| (Swap Dropout) | -0.8 | +20.5 | +4.8 | -3.8 | +2.7 | -3.3 | +3.1 | +2.6 | +4.3 |
| Method | Size | Camera | Robot | Language | Light | Background | Noise | Layout | Total |
|---|---|---|---|---|---|---|---|---|---|
| Direct VLM-to-VLA adaptation on standard LIBERO | |||||||||
| Qwen3-VL-OFT | 4B | 47.0 | 60.1 | 87.0 | 96.3 | 95.3 | 73.1 | 79.2 | 75.0 |
| Qwen3-VL-PI | 4B | 64.3 | 57.2 | 82.8 | 94.2 | 94.0 | 79.6 | 78.2 | 77.0 |
| Qwen3.5-0.8B-OFT (SFT) | 0.8B | 38.8 | 55.7 | 71.0 | 88.9 | 88.5 | 59.1 | 69.8 | 65.3 |
| Qwen3.5-0.8B-OFT (SFT AT+Swap) | 0.8B | 84.1 | 65.2 | 68.4 | 88.8 | 87.7 | 88.3 | 79.5 | 79.7 |
| Robot-pretrained VLAs finetuned on standard LIBERO | |||||||||