VLAs combine pretrained vision-language representations with action generation to enable language-guided control across diverse tasks, becoming a mainstream paradigm in embodied intelligence. However, multiple studies have reported VLA's substantial declines in task success under camera shifts, revealing a key vulnerability that limits reliable deployment. To address this vulnerability, existing methods collect paired observations of the same scene from different viewpoints to fine-tune the VLA or train visual adaptation modules. Unfortunately, they require additional data collection and VLA training costs. In this paper, we first identify \emph{cross-view coordination breakdown} under external camera shifts: the robot may rely too heavily on wrist-view cues and consequently execute subtasks in the wrong order when losing global view. Motivated by this, we propose CoRe-VLA, a plug-and-play framework requiring neither additional multi-view data collection nor VLA fine-tuning, which can incorporate with exsiting VLAs. It reconstructs a scene point cloud and renders the observation from the VLA's training viewpoint to restore cross-view coordination. In CoRe-VLA, Render-to-Camera (R2C) Restoration reduces rendering-induced visual degradation, while Execution-Trajectory-Conditioned Alignment (ETCA) reduces robot idle time and mitigates motion conflicts during asynchronous execution. Experiments on 5 real-robot tasks, LIBERO-100 and LIBERO-Plus demonstrate CoRe-VLA substantially improves task success across mainstream VLAs under camera shifts. For example, CoRe-VLA raises PI0.5's success rate from 13.3% to 83.3% at a 1.6m camera shift in real-robot environment.
Figures & tables
Figure 1: (a) The external camera shifts while the wrist-camera mounting remains unchanged. The accompanying curves compare task success rates on 5 real-robot tasks with and without CoRe-VLA. (b) Under the shifted view, the robot grasps the tape measure visible in the wrist view while the basket lid remains closed, violating the required subtask order. (c) The core idea of CoRe-VLA: using point-cloud reconstruction and re-rendering to restore the external observation and recover cross-view coordination.
ID
Category
Task instruction
T1
Basic
Pick up the green glue stick and place it in the pink cup.
T2
Long-horizon
Place the yellow tape measure and the pink cup into the basket in sequence.
T3
Long-horizon
Open the lid, then pick up the yellow tape measure and place it in the basket.
T4
Spatial
Pick up the green glue stick and place it on the side of the pink cup farther from the basket.
T5
Spatial
Pick up the object closest to the basket and place it inside.
Table 1: Real-world task suite used to diagnose VLA failures under external-camera shifts. T1 evaluates basic manipulation, T2–T3 long-horizon execution, and T4–T5 spatial reasoning.
Figure 2: Outcome distributions on five real-robot tasks under external-camera shifts for (a) π0.5 , (b) OpenVLA-OFT, and (c) GR00T. Failure Type 1 denotes failure to complete any meaningful subtask, while Failure Type 2 denotes local manipulation inconsistent with the global task state.
Figure 3: Overview of CoRe-VLA, comprising three components: geometric viewpoint restoration, R2C for image restoration, and ETCA for asynchronous control.
Figure 4: ETCA matches recorded robot motion to a new action chunk and selects where execution resumes.
Shift (m)
VLA
Metric
T1
T2
T3
T4
T5
Overall
Raw
CoRe
Raw
CoRe
Raw
CoRe
Raw
CoRe
Raw
CoRe
Raw
CoRe
0.0
π0.5
SR
100.0
83.3
83.3
83.3
100.0
100.0
100.0
100.0
100.0
100.0
96.7
93.3
F2
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
OpenVLA-OFT
SR
83.3
83.3
83.3
100.0
100.0
83.3
100.0
100.0
83.3
83.3
90.0
90.0
F2
0.0
0.0
16.7
0.0
0.0
0.0
0.0
0.0
16.7
16.7
6.7
3.3
GR00T
SR
83.3
66.7
66.7
66.7
66.7
66.7
100.0
83.3
100.0
100.0
83.3
76.7
Table 2: Real-robot outcome rates (%) across external-camera shifts. Raw and CoRe denote the shifted view and CoRe-VLA-restored view, respectively. SR denotes success rate, and F2 denotes the proportion of all trials classified as Failure Mode 2 (task-state inconsistency).
Figure 5: Real-robot performance across 5 manipulation tasks under external-camera shifts. (a) Task success rates. (b–d) Outcome proportions for π0.5 , OpenVLA-OFT, and GR00T, respectively. Failure Types 1 and 2 denote local execution failure and task-state inconsistency, respectively. G0 and Gs denote raw external-camera observations captured at the nominal and shifted positions, respectively, while Gr denotes observations transformed by CoRe-VLA.
Figure 6: Task success rates on LIBERO-100 under external-camera shifts, without CoRe-VLA (solid) and with CoRe-VLA (dashed).
Method
Backbone
SR (%) ↑
VLA based methods
Cross-View AC ( Huang et al., 2026b )
π0.5
87.2
GAM ( Han et al., 2026 )
DA3-Giant
83.1
Anchor-Align ( Dalal et al., 2026 )
Prismatic- Qwen2.5-0.5B
96.3
AVA-VLA ( Xiao et al., 2026 )
OpenVLA-OFT
69.4
SRPO ( Fei et al., 2026a )
OpenVLA
83.4
Table 3: Success rates (%) on LIBERO-Plus Camera.
Figure 7: Effect of R2C Restoration on π0.5 task success under external-camera shifts.
Method
Latency (ms)
Success (%)
Completion (s)
Idle (s)
π0.5
244
96.7
47.2
22.5
π0.5 + ETCA
244
96.7
34.4
9.0
π0.5 + CoRe-VLA w/o ETCA
478
93.3
64.5
39.1
π0.5 + CoRe-VLA
478
93.3
43.7
13.8
Table 4: Effect of ETCA on execution efficiency.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Success criterion
T1
The robot grasps the green glue stick and releases it inside the pink cup.
T2
The robot places and releases the yellow tape measure inside the basket, then places and releases the pink cup inside the basket.
T3
The robot first opens the lid so that it no longer covers the basket, then places and releases the yellow tape measure inside the basket.
T4
The robot grasps the green glue stick and releases it on the side of the pink cup farther from the basket, within a ±45∘ sector about the basket-to-cup direction.
T5
The robot grasps the object initially closest to the basket and releases it inside the basket.
Appendix
Table 5: Success criteria for the five real-robot tasks.
Figure 8: Representative real-robot execution sequences under nominal (Raw), shifted, and CoRe-VLA-restored external views. Arrows indicate action order; check marks and crosses denote success and failure, respectively.
Figure 9: Qualitative ablation of the R2C training losses. From left to right: original image, degraded input, and restorations using color loss, color and pixel losses, and the full objective including perceptual loss.