The diversity of robot embodiments and action spaces makes it challenging to build robot world models that generalize across different embodiments. We introduce the Latent Action-Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse embodiments. This unified action space improves the world model's performance when adapted to previously unseen robot embodiments. We compare LAC-WM with an Explicit Action-Conditioned World Model (EAC-WM), which conditions on explicit motion labels. Our results show that explicit action conditioning leads to disjoint action representations across embodiments, limiting downstream performance when adapting to new robots. We evaluate both models on dexterous manipulation tasks and a modified LIBERO benchmark. LAC-WM improves downstream performance over EAC-WM by up to 46.7% on dexterous manipulation and 11.7% on LIBERO. Crucially, the unified latent action space allows LAC-WM's downstream performance to scale positively with the number of embodiments used during pretraining. In contrast, the disjoint action space in EAC-WM leads to decreased performance as the number of pretraining embodiments increases. These results highlight the importance of a unified action space for efficient cross-embodiment learning, addressing a key challenge in robotics. Project website: https://lacwm.github.io/
Figures & tables
Figure 1 : LAC-WM Architecture . During pretraining (yellow area), the inverse dynamics model (IDM) encodes cross-embodiment image observations into a unified latent action space, as the condition for the forward dynamics model (FDM). A motion decoder translates the latent actions into motion labels, including delta human hand poses, robot end-effector actions, and camera motion. During finetuning (green area), an action projector encodes raw robot actions of unseen embodiments into the pre-trained unified latent action space.
Figure 2 : UMAP visualization of 7,000 action embeddings from three datasets: Droid (red), Agibot (green), and Egodex (blue). 1 : Embeddings from separate Explicit Action Encoders (EAE), which are distinctly separated by embodiment. 2 : Embeddings from a shared EAE (EAE*), still showing dataset-level separation. 3 : Embeddings from IDM without the motion decoder, where Agibot and Egodex cluster together but separate from Droid. 4 : Embeddings from the IDM, showing strong cross-dataset alignment.
Figure 3 : Cross-embodiment action transfer results : The left panels illustrate robot-to-human transfer, while the right panels show human-to-robot transfer. The first row presents reference videos. Rows 2–4 display FDM generated videos using embeddings from EAE, IDM without MD, and IDM, respectively. Blue dotted lines trace end-effector trajectories from the references and are aligned across generated videos except for initial positions. IDM embeddings yield trajectories that closely match the references, whereas the other methods exhibit inaccurate motion patterns and visual artifacts (highlighted by red circles).
Eval Split
Model
PSNR ↑
LPIPS ↓
FID ↓
FVD ↓
Unseen Instances
EAC-WM-S
22.212
0.077
10.648
42.257
EAC-WM
24.252
0.057
9.312
30.349
LAC-WM
27.484
0.037
8.439
23.213
Unseen Categories
EAC-WM-S
22.062
0.081
10.642
40.477
EAC-WM
24.091
0.060
9.144
29.702
LAC-WM
27.333
0.040
8.509
24.456
Table 1 : Action conditioned imagined roll-out performance . Image and video quality metrics are reported across evaluation splits on unseen instances and unseen categories. LAC-WM performs the best compared to models trained from scratch (EAC-WM-S) or pretrained with explicit actions (EAC-WM).
Val Set
Model
δf↓
δfg↓
S.R.C ↑
S.R.L ↑
S.R. ↑
Unseen Instances
VLA-mean
26.49 ± 0.06
27.07 ± 0.02
0.84 ± 0.04
0.56 ± 0.00
0.20 ± 0.03
VLA-random
26.52 ± 0.06
27.23 ± 0.12
0.82 ± 0.03
0.55 ± 0.03
0.18 ± 0.03
EAC-WM-S
26.74 ± 0.34
27.38 ± 0.35
0.84 ± 0.02
0.58 ± 0.04
0.19 ± 0.04
EAC-WM
26.79 ± 0.26
27.37 ± 0.27
0.84 ± 0.01
0.56 ± 0.04
0.16 ± 0.04
LAC-WM
26.29 ± 0.15
26.87 ± 0.13
0.86 ± 0.03
0.60 ± 0.03
0.25 ± 0.03
Unseen Categories
VLA-mean
26.94 ± 0.09
27.86 ± 0.17
0.84 ± 0.04
0.49 ± 0.04
0.14 ± 0.03
Table 2: Robot planning performance of a VLA with action selection using a world model (EAC-WM-S, EAC-WM, LAC-WM) versus without action selection (VLA-mean, VLA-random). Mean and standard deviation across three seeds are shown.
Figure 4 : We show the image latent embedding distance δf (left) and δfg (middle) and task success rate S.R. (right), for LAC-WM (blue) and EAC-WM (orange) pre-trained with different number of embodiments. The x axis shows the number of embodiments and dataset size used in pre-training, from left to right of the x axis, the model is trained with Egodex only, Egodex and Agibot, and Egodex, Agibot and Droid dataset.
Figure 5 : Experiments of scaling number of embodiments or data samples. The first two points shows the scaling from two embodiments (Egodex and Droid) to three embodiments (Egodex, Droid and Agibot) with same amount of trajectories, while the last two points shows the scaling from 24k trajectories to 150k trajectories across three embodiments.
Model
δf↓
δfg↓
S.R.C↑
S.R.L↑
S.R.↑
LAC-WM-MD-CA
26.98
27.80
0.83
0.46
0.13
LAC-WM
26.73
27.58
0.85
0.56
0.18
Table 3: Ablation study of motion decoding loss and cross-augmentation inputs on the validation set of unseen categories.
Task
VLA-random
VLA-mean
EAC-WM
LAC-WM
book
0.76
0.76
0.80
0.92
white bowl
0.76
0.92
0.92
0.96
wine bottle
0.72
0.84
0.84
0.96
alphabet soup
0.80
0.72
0.80
0.88
salad dressing
0.68
0.84
0.76
0.88
Average
0.744
0.816
0.824
0.920
Table 4: Success rate on five unseen LIBERO tasks over 125 total episodes . All methods share the same π0.5 VLA backbone, sampling 20 candidate 50-step action chunks.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Example of action conditioned imagined roll-out from the world models . The top row displays the ground truth video, while the subsequent rows show action-conditioned imagined rollouts from EAC-WM-S, EAC-WM, and LAC-WM. While LAC-WM accurately predicts future observations, both EAC-WM-S and EAC-WM exhibit overshooting robot actions, highlighted by red circles, where the robot arm moves excessively. This suggests that LAC-WM better captures the robot dynamics.
Hyperparameter
Value
Image Tokenizer
V-JEPA 2 (1B)
Action Projector
1,550,528
Image Decoder
153,845,952
IDM:
# Attention Blocks
4
# Attention Heads
16
Appendix
Table 5 : Hyperparameters for LAC-WM model architecture.
Figure 7 : Examples of action selection rollout . The leftmost image shows the initial observation, and the rightmost image depicts the goal state. Numbers in the top-left corner indicate time steps. In this example, three action sequences are sampled and rolled out by the world model, producing eight future predictions. The action sequence whose predicted final image has the smallest embedding distance ( δfg ) to the goal image is selected for execution. Here, the action sequence in the top row is chosen.
Val Set
Model
δf↓
δfg↓
S.R.C ↑
S.R.L ↑
S.R ↑
Unseen Instance
EAC-WM-S
25.00 ± 0.04
26.53 ± 0.06
0.59 ± 0.03
0.22 ± 0.02
0.11 ± 0.02
EAC-WM-TFT
25.05 ± 0.01
26.49 ± 0.10
0.54 ± 0.03
0.15 ± 0.06
0.09 ± 0.02
EAC-WM
25.03 ± 0.04
26.53 ± 0.12
0.61 ± 0.03
0.21 ± 0.02
0.11 ± 0.01
LAC-WM-DFT
24.82 ± 0.08
26.32 ± 0.20
0.61 ± 0.03
0.19 ± 0.02
0.07 ± 0.01
LAC-WM-(2,3)
24.88 ± 0.04
26.51 ± 0.08
0.54 ± 0.09
0.21 ± 0.01
0.09 ± 0.01
LAC-WM-(1,2)
24.90 ± 0.09
26.37 ± 0.12
0.59 ± 0.07
0.21 ± 0.05
0.09 ± 0.06
Appendix
Table 6: Robot planning performance for different methods using action selection from the training dataset.
Figure 8 : Example roll-out of executing actions from a pre-trained VLA. From top to bottom, the figure shows rollouts for: executing the mean of all sampled actions (VLA-mean), a random sampled action (VLA-random), and actions selected using EAC-WM-S, EAC-WM, and LAC-WM from the VLA samples. The task is to pick up the object and place it back on the counter. While LAC-WM successfully selects actions to complete the task, the other methods fail to grasp and lift the object.
Figure 9 : Latent Action PCA Analysis . The representative examples from each of the cluster show that latent action in each cluster corresponds to semantically meaningful actions (e.g. opening and closing gripper, moving left and right) across different embodiments.
Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to reveal whether they capture physical dynamics or merely memorize visual patterns. To answer whether a model can faithfully render a robot it has never seen, we introduce XEWorld, a controlled cross-embodiment testbed for world models that isolates embodiments by evaluating held-out robots within physically identical scenes. Our systematic analysis uncovers a shared architectural bottleneck: current models act primarily as 2D visual pattern matchers whose generalization is governed by visual similarity rather than physical kinematic similarity. Driven by this limitation, they struggle to translate abstract numeric joint actions into coherent visual trajectories, and fail to predict dynamic visual changes from static initial observations. Consequently, successfully rendering an unseen embodiment zero-shot strictly requires heavily grounded cues, specifically pixel-space actions and explicit spatial-temporal alignment. Even when bypassing this zero-shot barrier via few-shot adaptation, the forced appearance recovery triggers catastrophic forgetting of seen embodiments. Together, these failures expose a critical inability to apply learned physical dynamics to novel visual appearances, highlighting that achieving true cross-embodiment generalization requires architectural innovations that decouple visual appearance from underlying physical dynamics.
Yixiang Chen, Jiabing Yang, Yuan Xu +10
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Shandong University +1
As generalist robot policies gain vision and language from web-scale pretraining, demonstrations remain costly to collect and tied to the robot that recorded them. Latent action models (LAMs) address both by learning latent actions from action-free videos that can be shared across embodiments, however, in practice, LAMs are sensitive to background visual noise, and the same motion from two different robots may be encoded with different latents. One solution to the background visual noise is to add an auxiliary loss predicting the robot action from the latent action, further associating the latent action space to the embodiment specific robot action space. We study a different use of the same labels, through action-similarity supervision. The similarity between any two latent actions is trained to match the similarity of the two ground-truth robot action sequences. The ground-truth actions are never predicted by the LAM, so the latent action does not need to encode embodiment specifics. We evaluate cross-embodiment transfer on RoboTwin 2.0 in a controlled setup, two bimanual robots demonstrate disjoint task sets, a policy is trained on all the demonstrations, and each robot is evaluated closed-loop on the tasks only the other demonstrated. With the policy architecture and its hyperparameters, the dataset, and the evaluation protocol fixed, predicting latent actions instead of ground-truth actions more than doubles cross-embodiment success. Given the same ground-truth actions, similarity supervision transfers better than an auxiliary loss that predicts the ground-truth action during the LAM training. Computing the similarities on end-effector motion rather than joint-space motion, and letting the loss compare latent actions across the two robots, gives the best approach of the study.
Maxime Alvarez, Renzo Caballero, Tatsuya Matsushima +2
Graduate School of Engineering, The University of Tokyo
Action-conditioned world models (ACWMs) aim to simulate future observations conditioned on embodied actions, offering a promising foundation for robot planning, policy evaluation, and data augmentation. However, learning controllable ACWMs requires large-scale action-labeled data, which remains costly to collect in the real world. Latent action models (LAMs) mitigate this bottleneck by inferring latent actions from unlabeled videos, but existing LAMs are typically trained with reconstruction-only objectives and therefore entangle action-relevant dynamics with action-irrelevant visual factors such as backgrounds and untouched objects. In this work, we identify this action-irrelevant bias as a key obstacle to controllable ACWMs and introduce evaluation metrics to measure latent-action bias, action following, and robustness. We propose CD-LAM, a causally debiased framework for LAM-based ACWMs. CD-LAM introduces three efficient fine-tuning objectives: embodiment-centric reconstruction, action-centric contrastive learning, and latent space calibration, which together encourage embodiment-focused, action-aware, and calibrated non-collapsed latent action representations. Experiments on 2B and 14B ACWM backbones show that CD-LAM substantially improves latent-action controllability, downstream robot-action following, visual fidelity, and adaptation efficiency, requiring only 6k fine-tuning steps and more than 12× fewer robot-action adaptation updates than the baseline.