The diversity of robot embodiments and action spaces makes it challenging to build robot world models that generalize across different embodiments. We introduce the Latent Action-Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse embodiments. This unified action space improves the world model's performance when adapted to previously unseen robot embodiments. We compare LAC-WM with an Explicit Action-Conditioned World Model (EAC-WM), which conditions on explicit motion labels. Our results show that explicit action conditioning leads to disjoint action representations across embodiments, limiting downstream performance when adapting to new robots. We evaluate both models on dexterous manipulation tasks and a modified LIBERO benchmark. LAC-WM improves downstream performance over EAC-WM by up to 46.7% on dexterous manipulation and 11.7% on LIBERO. Crucially, the unified latent action space allows LAC-WM's downstream performance to scale positively with the number of embodiments used during pretraining. In contrast, the disjoint action space in EAC-WM leads to decreased performance as the number of pretraining embodiments increases. These results highlight the importance of a unified action space for efficient cross-embodiment learning, addressing a key challenge in robotics. Project website: https://lacwm.github.io/
Figures & tables
Figure 1 : LAC-WM Architecture . During pretraining (yellow area), the inverse dynamics model (IDM) encodes cross-embodiment image observations into a unified latent action space, as the condition for the forward dynamics model (FDM). A motion decoder translates the latent actions into motion labels, including delta human hand poses, robot end-effector actions, and camera motion. During finetuning (green area), an action projector encodes raw robot actions of unseen embodiments into the pre-trained unified latent action space.
Figure 2 : UMAP visualization of 7,000 action embeddings from three datasets: Droid (red), Agibot (green), and Egodex (blue). 1 : Embeddings from separate Explicit Action Encoders (EAE), which are distinctly separated by embodiment. 2 : Embeddings from a shared EAE (EAE*), still showing dataset-level separation. 3 : Embeddings from IDM without the motion decoder, where Agibot and Egodex cluster together but separate from Droid. 4 : Embeddings from the IDM, showing strong cross-dataset alignment.
Figure 3 : Cross-embodiment action transfer results : The left panels illustrate robot-to-human transfer, while the right panels show human-to-robot transfer. The first row presents reference videos. Rows 2–4 display FDM generated videos using embeddings from EAE, IDM without MD, and IDM, respectively. Blue dotted lines trace end-effector trajectories from the references and are aligned across generated videos except for initial positions. IDM embeddings yield trajectories that closely match the references, whereas the other methods exhibit inaccurate motion patterns and visual artifacts (highlighted by red circles).
Eval Split
Model
PSNR ↑
LPIPS ↓
FID ↓
FVD ↓
Unseen Instances
EAC-WM-S
22.212
0.077
10.648
42.257
EAC-WM
24.252
0.057
9.312
30.349
LAC-WM
27.484
0.037
8.439
23.213
Unseen Categories
EAC-WM-S
22.062
0.081
10.642
40.477
EAC-WM
24.091
0.060
9.144
29.702
LAC-WM
27.333
0.040
8.509
24.456
Table 1 : Action conditioned imagined roll-out performance . Image and video quality metrics are reported across evaluation splits on unseen instances and unseen categories. LAC-WM performs the best compared to models trained from scratch (EAC-WM-S) or pretrained with explicit actions (EAC-WM).
Val Set
Model
δf↓
δfg↓
S.R.C ↑
S.R.L ↑
S.R. ↑
Unseen Instances
VLA-mean
26.49 ± 0.06
27.07 ± 0.02
0.84 ± 0.04
0.56 ± 0.00
0.20 ± 0.03
VLA-random
26.52 ± 0.06
27.23 ± 0.12
0.82 ± 0.03
0.55 ± 0.03
0.18 ± 0.03
EAC-WM-S
26.74 ± 0.34
27.38 ± 0.35
0.84 ± 0.02
0.58 ± 0.04
0.19 ± 0.04
EAC-WM
26.79 ± 0.26
27.37 ± 0.27
0.84 ± 0.01
0.56 ± 0.04
0.16 ± 0.04
LAC-WM
26.29 ± 0.15
26.87 ± 0.13
0.86 ± 0.03
0.60 ± 0.03
0.25 ± 0.03
Unseen Categories
VLA-mean
26.94 ± 0.09
27.86 ± 0.17
0.84 ± 0.04
0.49 ± 0.04
0.14 ± 0.03
Table 2: Robot planning performance of a VLA with action selection using a world model (EAC-WM-S, EAC-WM, LAC-WM) versus without action selection (VLA-mean, VLA-random). Mean and standard deviation across three seeds are shown.
Figure 4 : We show the image latent embedding distance δf (left) and δfg (middle) and task success rate S.R. (right), for LAC-WM (blue) and EAC-WM (orange) pre-trained with different number of embodiments. The x axis shows the number of embodiments and dataset size used in pre-training, from left to right of the x axis, the model is trained with Egodex only, Egodex and Agibot, and Egodex, Agibot and Droid dataset.
Figure 5 : Experiments of scaling number of embodiments or data samples. The first two points shows the scaling from two embodiments (Egodex and Droid) to three embodiments (Egodex, Droid and Agibot) with same amount of trajectories, while the last two points shows the scaling from 24k trajectories to 150k trajectories across three embodiments.
Model
δf↓
δfg↓
S.R.C↑
S.R.L↑
S.R.↑
LAC-WM-MD-CA
26.98
27.80
0.83
0.46
0.13
LAC-WM
26.73
27.58
0.85
0.56
0.18
Table 3: Ablation study of motion decoding loss and cross-augmentation inputs on the validation set of unseen categories.
Task
VLA-random
VLA-mean
EAC-WM
LAC-WM
book
0.76
0.76
0.80
0.92
white bowl
0.76
0.92
0.92
0.96
wine bottle
0.72
0.84
0.84
0.96
alphabet soup
0.80
0.72
0.80
0.88
salad dressing
0.68
0.84
0.76
0.88
Average
0.744
0.816
0.824
0.920
Table 4: Success rate on five unseen LIBERO tasks over 125 total episodes . All methods share the same π0.5 VLA backbone, sampling 20 candidate 50-step action chunks.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Example of action conditioned imagined roll-out from the world models . The top row displays the ground truth video, while the subsequent rows show action-conditioned imagined rollouts from EAC-WM-S, EAC-WM, and LAC-WM. While LAC-WM accurately predicts future observations, both EAC-WM-S and EAC-WM exhibit overshooting robot actions, highlighted by red circles, where the robot arm moves excessively. This suggests that LAC-WM better captures the robot dynamics.
Hyperparameter
Value
Image Tokenizer
V-JEPA 2 (1B)
Action Projector
1,550,528
Image Decoder
153,845,952
IDM:
# Attention Blocks
4
# Attention Heads
16
Appendix
Table 5 : Hyperparameters for LAC-WM model architecture.
Figure 7 : Examples of action selection rollout . The leftmost image shows the initial observation, and the rightmost image depicts the goal state. Numbers in the top-left corner indicate time steps. In this example, three action sequences are sampled and rolled out by the world model, producing eight future predictions. The action sequence whose predicted final image has the smallest embedding distance ( δfg ) to the goal image is selected for execution. Here, the action sequence in the top row is chosen.
Val Set
Model
δf↓
δfg↓
S.R.C ↑
S.R.L ↑
S.R ↑
Unseen Instance
EAC-WM-S
25.00 ± 0.04
26.53 ± 0.06
0.59 ± 0.03
0.22 ± 0.02
0.11 ± 0.02
EAC-WM-TFT
25.05 ± 0.01
26.49 ± 0.10
0.54 ± 0.03
0.15 ± 0.06
0.09 ± 0.02
EAC-WM
25.03 ± 0.04
26.53 ± 0.12
0.61 ± 0.03
0.21 ± 0.02
0.11 ± 0.01
LAC-WM-DFT
24.82 ± 0.08
26.32 ± 0.20
0.61 ± 0.03
0.19 ± 0.02
0.07 ± 0.01
LAC-WM-(2,3)
24.88 ± 0.04
26.51 ± 0.08
0.54 ± 0.09
0.21 ± 0.01
0.09 ± 0.01
LAC-WM-(1,2)
24.90 ± 0.09
26.37 ± 0.12
0.59 ± 0.07
0.21 ± 0.05
0.09 ± 0.06
Appendix
Table 6: Robot planning performance for different methods using action selection from the training dataset.
Figure 8 : Example roll-out of executing actions from a pre-trained VLA. From top to bottom, the figure shows rollouts for: executing the mean of all sampled actions (VLA-mean), a random sampled action (VLA-random), and actions selected using EAC-WM-S, EAC-WM, and LAC-WM from the VLA samples. The task is to pick up the object and place it back on the counter. While LAC-WM successfully selects actions to complete the task, the other methods fail to grasp and lift the object.
Figure 9 : Latent Action PCA Analysis . The representative examples from each of the cluster show that latent action in each cluster corresponds to semantically meaningful actions (e.g. opening and closing gripper, moving left and right) across different embodiments.
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Shandong University +1