CoRe-VLA: Preserving Cross-View Coordination in VLAs under Camera Shifts
Organizations: Southeast University, Nanjing, China · National Center of Technology Innovation for EDA, Nanjing, China
Abstract
VLAs combine pretrained vision-language representations with action generation to enable language-guided control across diverse tasks, becoming a mainstream paradigm in embodied intelligence. However, multiple studies have reported VLA's substantial declines in task success under camera shifts, revealing a key vulnerability that limits reliable deployment. To address this vulnerability, existing methods collect paired observations of the same scene from different viewpoints to fine-tune the VLA or train visual adaptation modules. Unfortunately, they require additional data collection and VLA training costs. In this paper, we first identify \emph{cross-view coordination breakdown} under external camera shifts: the robot may rely too heavily on wrist-view cues and consequently execute subtasks in the wrong order when losing global view. Motivated by this, we propose CoRe-VLA, a plug-and-play framework requiring neither additional multi-view data collection nor VLA fine-tuning, which can incorporate with exsiting VLAs. It reconstructs a scene point cloud and renders the observation from the VLA's training viewpoint to restore cross-view coordination. In CoRe-VLA, Render-to-Camera (R2C) Restoration reduces rendering-induced visual degradation, while Execution-Trajectory-Conditioned Alignment (ETCA) reduces robot idle time and mitigates motion conflicts during asynchronous execution. Experiments on 5 real-robot tasks, LIBERO-100 and LIBERO-Plus demonstrate CoRe-VLA substantially improves task success across mainstream VLAs under camera shifts. For example, CoRe-VLA raises PI0.5's success rate from 13.3% to 83.3% at a 1.6m camera shift in real-robot environment.
Figures & tables
| ID | Category | Task instruction |
|---|---|---|
| T1 | Basic | Pick up the green glue stick and place it in the pink cup. |
| T2 | Long-horizon | Place the yellow tape measure and the pink cup into the basket in sequence. |
| T3 | Long-horizon | Open the lid, then pick up the yellow tape measure and place it in the basket. |
| T4 | Spatial | Pick up the green glue stick and place it on the side of the pink cup farther from the basket. |
| T5 | Spatial | Pick up the object closest to the basket and place it inside. |
| Shift (m) | VLA | Metric | T1 | T2 | T3 | T4 | T5 | Overall | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Raw | CoRe | Raw | CoRe | Raw | CoRe | Raw | CoRe | Raw | CoRe | Raw | CoRe | |||
| 0.0 | SR | 100.0 | 83.3 | 83.3 | 83.3 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 96.7 | 93.3 | |
| F2 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ||
| OpenVLA-OFT | SR | 83.3 | 83.3 | 83.3 | 100.0 | 100.0 | 83.3 | 100.0 | 100.0 | 83.3 | 83.3 | 90.0 | 90.0 | |
| F2 | 0.0 | 0.0 | 16.7 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 16.7 | 16.7 | 6.7 | 3.3 | ||
| GR00T | SR | 83.3 | 66.7 | 66.7 | 66.7 | 66.7 | 66.7 | 100.0 | 83.3 | 100.0 | 100.0 | 83.3 | 76.7 | |
| Method | Backbone | SR (%) |
| VLA based methods | ||
| Cross-View AC ( Huang et al., 2026b ) | 87.2 | |
| GAM ( Han et al., 2026 ) | DA3-Giant | 83.1 |
| Anchor-Align ( Dalal et al., 2026 ) | Prismatic- Qwen2.5-0.5B | 96.3 |
| AVA-VLA ( Xiao et al., 2026 ) | OpenVLA-OFT | 69.4 |
| SRPO ( Fei et al., 2026a ) | OpenVLA | 83.4 |
| Method | Latency (ms) | Success (%) | Completion (s) | Idle (s) |
|---|---|---|---|---|
| 244 | 96.7 | 47.2 | 22.5 | |
| + ETCA | 244 | 96.7 | 34.4 | 9.0 |
| + CoRe-VLA w/o ETCA | 478 | 93.3 | 64.5 | 39.1 |
| + CoRe-VLA | 478 | 93.3 | 43.7 | 13.8 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Success criterion |
|---|---|
| T1 | The robot grasps the green glue stick and releases it inside the pink cup. |
| T2 | The robot places and releases the yellow tape measure inside the basket, then places and releases the pink cup inside the basket. |
| T3 | The robot first opens the lid so that it no longer covers the basket, then places and releases the yellow tape measure inside the basket. |
| T4 | The robot grasps the green glue stick and releases it on the side of the pink cup farther from the basket, within a sector about the basket-to-cup direction. |
| T5 | The robot grasps the object initially closest to the basket and releases it inside the basket. |