Organizations: The Hong Kong University of Science and Technology (Guangzhou) · The Hong Kong University of Science and Technology · ETH Zurich · Zhejiang University · EPFL · University of Zurich · The Chinese University of Hong Kong
World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit. This makes it harder to connect global scene context with the local geometry required for interaction. We introduce the Multi-View Geometry-Aware World-Action Model (MVG-WAM), which organizes these observations as related projections of one physical world rather than separate images on a canvas. Our model combines an epipolar-constrained global state with view-indexed geometric states jointly inferred from synchronized observations. Camera-aware routing supplies each video region with its corresponding geometric context and the shared global state, explicitly structuring the representation used for action prediction. We further ground the geometry-aware representation in metric scale through multi-horizon future-depth supervision, without requiring depth decoding during action rollout. MVG-WAM achieves average success rates of 99.1% on LIBERO and 92.07% on RoboTwin 2.0, demonstrating competitive performance across both benchmarks. Real-world experiments on Cobot Magic further demonstrate a 91.3% success rate across 150 trials spanning three manipulation tasks.
Figures & tables
Fig. 2: Architecture of our MVG-WAM. (a) Dense DINOv3 features are aggregated under calibrated epipolar constraints to provide cross-view geometric context. (b) Compact VGGT- Ω registers form view-indexed geometric states, which are routed to the corresponding camera regions. (c) The global and view-indexed states condition the video stream of the video–action diffusion transformer, while action tokens access the geometry-enhanced video representation through mixed attention. (d) During training, future metric-depth supervision provides metric-scale grounding for intermediate denoising features across views and horizons; the depth decoder is not used during action rollout.
LIBERO
RoboTwin 2.0
Method
Emb. PT.
Spatial
Object
Goal
Long
Avg.
Clean
Rand.
Avg.
π0.5 [ 26 , 6 ]
Yes
98.8
98.2
98.0
92.4
96.9
82.74
76.76
79.75
X-VLA [ 27 , 3 ]
Yes
98.2
98.6
97.8
97.6
98.1
72.80
72.84
72.82
ABot-M0 [ 28 ]
Yes
98.8
99.8
99.0
96.6
98.6
80.42
81.16
80.79
StarVLA- α [ 29 ]
No
99.0
99.8
98.5
94.1
97.9
88.20
88.30
88.25
Motus [ 3 , 6 ]
Yes
96.8
99.8
96.6
97.6
97.7
88.66
87.02
87.84
TABLE I: Success rate (%) on LIBERO and RoboTwin 2.0. Emb. PT. indicates additional embodied/robot-data pretraining before target-benchmark training. Bold and underlined values denote the highest and second-highest values in each benchmark column, respectively.
Shift
OpenVLA
OpenVLA-OFT
FastWAM-Joint
MVG-WAM
Camera
0.8
56.4
39.9
51.9
Robot
3.5
31.9
65.1
65.3
Language
23.0
79.5
94.7
90.1
Light
8.1
88.7
92.1
88.4
Background
34.8
93.3
58.1
62.3
Noise
15.2
75.8
56.2
57.6
TABLE II: LIBERO-Plus success rate (%) under seven distribution shifts. OpenVLA and OpenVLA-OFT results follow LIBERO-Plus [ 37 ] ; FastWAM-Joint [ 6 ] and MVG-WAM use our evaluation. Bold and underlined values denote the highest and second-highest values in each row, respectively.
X-WAM
FastWAM
AHA-WAM
FastWAM-Joint
MVG-WAM
C2C
70.0
77.8
64.3
69.8
72.6
C2R
25.8
1.9
3.2
1.3
29.2
Avg.
47.9
39.9
33.8
35.6
50.9
TABLE III: RoboTwin 2.0 Clean2Clean (C2C), Clean2Random (C2R), and average success rates (%). X-WAM, FastWAM, and AHA-WAM results follow the RoboTwin 2.0 leaderboard snapshot [ 40 ] ; FastWAM-Joint and MVG-WAM are evaluated in our clean-only setting. Bold and underlined values denote the highest and second-highest values in each row, respectively.
Fig. 3: Real-world evaluation on Cobot Magic. (a) Physical manipulation platform with the wrist-mounted scene-view camera. (b)–(d) Representative execution sequences for radish-to-plate, cup stacking, and pot opening followed by pepper placement, respectively. The third task consists of two stages: removing the pot lid and subsequently placing the pepper into the pot. Frames progress from left to right within each sequence. The shown frames are captured from an external recording viewpoint rather than from the camera observations supplied to the policy. Quantitative success rates over 50 trials per task are reported in Table V .
Variant
All 50
Hard 13
MVG-WAM (Ours)
92.05
77.65
w/o Epipolar Residual
91.05
75.00
w/o Epipolar Constraint
90.86
74.65
w/o View Routing
90.93
74.92
w/o Depth Supervision
90.96
74.81
FastWAM-Joint (Baseline)
89.44
71.96
TABLE IV: RoboTwin 2.0 ablation success rates (%) under the reduced 50-clean/200-random training regime. Hard-13 is fixed by independent FastWAM-Joint validation. Rows beginning with “w/o” denote MVG-WAM variants with the corresponding component removed.
Model
R
C
O
Mean
π0.5 [ 26 ]
92
88
88
89.3
FastWAM-Joint [ 6 ]
86
78
72
78.7
MVG-WAM w/o Epipolar Residual
94
86
82
87.3
MVG-WAM (Ours)
96
90
88
91.3
TABLE V: Cobot Magic success rate (%), 50 trials per task. R, C, and O denote radish-to-plate, cup stacking, and pot opening followed by pepper placement. “w/o” denotes without.
Fig. 4: Qualitative short-horizon RGB prediction on RoboTwin 2.0. FastWAM-Joint and MVG-WAM use the same current observation and are compared against the same future target. Dashed boxes indicate gripper–object interaction regions, with enlarged crops shown in red. MVG-WAM produces more faithful local predictions, preserving sharper object boundaries and more coherent gripper–object geometry.
Model / ref.
MAE ↓ (cm)
AbsRel ↓ (%)
CV ↓ (cm)
Cov. (%)
PSNR ↑ (dB)
SSIM ↑
FastWAM-Joint
—
—
—
—
28.05
0.842
MVG-WAM (Ours)
2.12
5.80
1.68
16.70
31.79
0.914
GT reference
0.00∗
0.00∗
0.25
19.00
∞∗
1.000∗
TABLE VI: Short-horizon RGB and teacher-forced depth diagnostics on RoboTwin 2.0, averaged equally over Clean and Random domains. A dash denotes not applicable: FastWAM-Joint has no trained native depth head. Starred GT entries are analytical self-comparisons; CV and coverage are reprojection references.
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · Beijing Innovation Center of Humanoid Robotics