Organizations: The Hong Kong University of Science and Technology (Guangzhou) · The Hong Kong University of Science and Technology · ETH Zurich · Zhejiang University · EPFL · University of Zurich · The Chinese University of Hong Kong
World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit. This makes it harder to connect global scene context with the local geometry required for interaction. We introduce the Multi-View Geometry-Aware World-Action Model (MVG-WAM), which organizes these observations as related projections of one physical world rather than separate images on a canvas. Our model combines an epipolar-constrained global state with view-indexed geometric states jointly inferred from synchronized observations. Camera-aware routing supplies each video region with its corresponding geometric context and the shared global state, explicitly structuring the representation used for action prediction. We further ground the geometry-aware representation in metric scale through multi-horizon future-depth supervision, without requiring depth decoding during action rollout. MVG-WAM achieves average success rates of 99.1% on LIBERO and 92.07% on RoboTwin 2.0, demonstrating competitive performance across both benchmarks. Real-world experiments on Cobot Magic further demonstrate a 91.3% success rate across 150 trials spanning three manipulation tasks.
Figures & tables
Fig. 2: Architecture of our MVG-WAM. (a) Dense DINOv3 features are aggregated under calibrated epipolar constraints to provide cross-view geometric context. (b) Compact VGGT- Ω registers form view-indexed geometric states, which are routed to the corresponding camera regions. (c) The global and view-indexed states condition the video stream of the video–action diffusion transformer, while action tokens access the geometry-enhanced video representation through mixed attention. (d) During training, future metric-depth supervision provides metric-scale grounding for intermediate denoising features across views and horizons; the depth decoder is not used during action rollout.
LIBERO
RoboTwin 2.0
Method
Emb. PT.
Spatial
Object
Goal
Long
Avg.
Clean
Rand.
Avg.
π0.5 [ 26 , 6 ]
Yes
98.8
98.2
98.0
92.4
96.9
82.74
76.76
79.75
X-VLA [ 27 , 3 ]
Yes
98.2
98.6
97.8
97.6
98.1
72.80
72.84
72.82
ABot-M0 [ 28 ]
Yes
98.8
99.8
99.0
96.6
98.6
80.42
81.16
80.79
StarVLA- α [ 29 ]
No
99.0
99.8
98.5
94.1
97.9
88.20
88.30
88.25
Motus [ 3 , 6 ]
Yes
96.8
99.8
96.6
97.6
97.7
88.66
87.02
87.84
TABLE I: Success rate (%) on LIBERO and RoboTwin 2.0. Emb. PT. indicates additional embodied/robot-data pretraining before target-benchmark training. Bold and underlined values denote the highest and second-highest values in each benchmark column, respectively.
Shift
OpenVLA
OpenVLA-OFT
FastWAM-Joint
MVG-WAM
Camera
0.8
56.4
39.9
51.9
Robot
3.5
31.9
65.1
65.3
Language
23.0
79.5
94.7
90.1
Light
8.1
88.7
92.1
88.4
Background
34.8
93.3
58.1
62.3
Noise
15.2
75.8
56.2
57.6
TABLE II: LIBERO-Plus success rate (%) under seven distribution shifts. OpenVLA and OpenVLA-OFT results follow LIBERO-Plus [ 37 ] ; FastWAM-Joint [ 6 ] and MVG-WAM use our evaluation. Bold and underlined values denote the highest and second-highest values in each row, respectively.
X-WAM
FastWAM
AHA-WAM
FastWAM-Joint
MVG-WAM
C2C
70.0
77.8
64.3
69.8
72.6
C2R
25.8
1.9
3.2
1.3
29.2
Avg.
47.9
39.9
33.8
35.6
50.9
TABLE III: RoboTwin 2.0 Clean2Clean (C2C), Clean2Random (C2R), and average success rates (%). X-WAM, FastWAM, and AHA-WAM results follow the RoboTwin 2.0 leaderboard snapshot [ 40 ] ; FastWAM-Joint and MVG-WAM are evaluated in our clean-only setting. Bold and underlined values denote the highest and second-highest values in each row, respectively.
Fig. 3: Real-world evaluation on Cobot Magic. (a) Physical manipulation platform with the wrist-mounted scene-view camera. (b)–(d) Representative execution sequences for radish-to-plate, cup stacking, and pot opening followed by pepper placement, respectively. The third task consists of two stages: removing the pot lid and subsequently placing the pepper into the pot. Frames progress from left to right within each sequence. The shown frames are captured from an external recording viewpoint rather than from the camera observations supplied to the policy. Quantitative success rates over 50 trials per task are reported in Table V .
Variant
All 50
Hard 13
MVG-WAM (Ours)
92.05
77.65
w/o Epipolar Residual
91.05
75.00
w/o Epipolar Constraint
90.86
74.65
w/o View Routing
90.93
74.92
w/o Depth Supervision
90.96
74.81
FastWAM-Joint (Baseline)
89.44
71.96
TABLE IV: RoboTwin 2.0 ablation success rates (%) under the reduced 50-clean/200-random training regime. Hard-13 is fixed by independent FastWAM-Joint validation. Rows beginning with “w/o” denote MVG-WAM variants with the corresponding component removed.
Model
R
C
O
Mean
π0.5 [ 26 ]
92
88
88
89.3
FastWAM-Joint [ 6 ]
86
78
72
78.7
MVG-WAM w/o Epipolar Residual
94
86
82
87.3
MVG-WAM (Ours)
96
90
88
91.3
TABLE V: Cobot Magic success rate (%), 50 trials per task. R, C, and O denote radish-to-plate, cup stacking, and pot opening followed by pepper placement. “w/o” denotes without.
Fig. 4: Qualitative short-horizon RGB prediction on RoboTwin 2.0. FastWAM-Joint and MVG-WAM use the same current observation and are compared against the same future target. Dashed boxes indicate gripper–object interaction regions, with enlarged crops shown in red. MVG-WAM produces more faithful local predictions, preserving sharper object boundaries and more coherent gripper–object geometry.
Model / ref.
MAE ↓ (cm)
AbsRel ↓ (%)
CV ↓ (cm)
Cov. (%)
PSNR ↑ (dB)
SSIM ↑
FastWAM-Joint
—
—
—
—
28.05
0.842
MVG-WAM (Ours)
2.12
5.80
1.68
16.70
31.79
0.914
GT reference
0.00∗
0.00∗
0.25
19.00
∞∗
1.000∗
TABLE VI: Short-horizon RGB and teacher-forced depth diagnostics on RoboTwin 2.0, averaged equally over Clean and Random domains. A dash denotes not applicable: FastWAM-Joint has no trained native depth head. Starred GT entries are analytical self-comparisons; CV and coverage are reprojection references.
World Action Models (WAMs) have shown strong potential for robotic manipulation by jointly modeling visual future dynamics and executable action sequences. However, existing video-action co-training methods primarily optimize appearance-oriented video latents, which may insufficiently capture the temporally evolving geometry required for precise manipulation. We propose MECo-WAM, a Multi-Expert Co-Training World Action Model that injects action-relevant 4D geometric priors into video-action representations while preserving the original lightweight inference graph. During training, MECo-WAM combines video and action experts with a lightweight 4D expert supervised by relational targets from a frozen VGGT encoder. Asymmetric expert visibility prevents non-causal shortcuts from auxiliary geometry to action generation. To transfer geometric knowledge into the deployed video-action pathway, we introduce decayed 4D read-mask attention, which provides restricted current-frame geometric guidance early in training and progressively removes this dependency. We further propose action-aware temporal geometric distillation, which aligns within-frame geometric relations and their temporal evolution while emphasizing visual regions most relevant to robot actions. At deployment, all auxiliary 4D components are removed. Experiments on LIBERO (98.2%), RoboTwin 2.0 (92.6%), and challenging real-world manipulation tasks show that MECo-WAM improves manipulation performance without increasing inference cost.
Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but often degrade under visual and environmental shifts. Latent world modeling offers a promising approach to improving robustness, yet existing methods commonly encode camera views independently and predict holistic scene dynamics without explicitly modeling their geometric relationships. We propose GWM-VLA, a geometry-aware latent world modeling framework for VLA learning. GWM-VLA combines geometry-aware multi-view state encoding, global context-conditioned target-view prediction, and shared latent-action representations grounded by robot-action supervision. Specifically, VGGT-Ω jointly aggregates multi-view observations at each timestep to construct geometry-aware multi-view states. The latent world model predicts the next-step patch tokens of a selected target view using patch and register tokens obtained after multi-view aggregation, thereby retaining multi-view geometric information without predicting the complete multi-view state. We use the wrist view as the target in our experiments, placing greater emphasis on end-effector motion and local gripper-object interactions. Finally, the shared latent-action representations condition both the latent world model and the flow-matching action head, allowing latent-prediction supervision and ground-truth robot-action supervision to jointly shape the same latent-action representations. Experiments across both simulation and real-world environments demonstrate the effectiveness and robustness of GWM-VLA.
Achieving robust and generalizable manipulation across diverse environments remains a fundamental challenge in embodied robotics. Recent world action models achieve strong in-domain performance, yet their gains do not extend proportionally to out-of-distribution scenarios. We attribute this to a structural mismatch between visual and action modalities, whose intrinsically heterogeneous manifolds cause joint optimization to disproportionately degrade action robustness under distribution shift. To address this, we propose MV-WAM, a novel end-to-end framework that jointly models visual prediction, action generation, and value estimation designed to effectively leverage video priors during both training and inference for enhanced action generalization. Key to this unification is a cross-modality causal mask that hierarchically grounds actions in predicted video frames and value function tokens in both modalities. To further narrow the generalization gap, MV-WAM adopts a manifold-aware optimization scheme that explicitly accounts for the structural heterogeneity across modalities. Finally, MV-WAM introduces a progress-value regulation mechanism that estimates task completion and detects misalignment between predicted frames and generated actions, enabling the policy to autonomously identify execution deviations and recover through value-guided rollback. On the RoboTwin simulation, MV-WAM achieves a 55.7% mean success rate on random scenarios without any randomized action supervision, outperforming the strongest baseline by 29.3%. MV-WAM achieves a 77.5% mean success rate across four real-world tasks of varying difficulty on a dual-arm robot. Our results demonstrate that manifold-aware cross-modal alignment is essential for robust policy generalization, offering a path toward deployable robotic manipulation.
Jintao Chen, Peidong Jia, Qingpo Wuwu +13
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · Beijing Innovation Center of Humanoid Robotics