World action models (WAMs) predict actions and future observations, typically from a reconstruction-based representation that carries noisy, redundant information which can complicate downstream predictions. We introduce LeWAM, a bidirectional transformer for forward, backward, inverse dynamics and policy prediction, on a decoder-free JEPA latent trained end-to-end through all four modes. We see the following benefits: 1) Alignment: linear probes read robot and object state from LeWAM's latent better than from a regular Le World Model (a forward-only JEPA world model), while the latent ignores visual distractors as well as LeWM does and far better than a reconstruction-based WAM. 2) Acting: Closed-loop evaluations of LeWAM match a regular flow-matching policy trained on the same encoder at matched size, while also providing a world model. 3) Planning: Sampling raw actions when planning with WAMs lets MPC exploit dynamics-model inaccuracies; planning in the noise space of the policy head instead improves the closed-loop performance of these WAMs.
Figures & tables
decoder-
action loss
policy
forward
inverse
backward
multimodal
one
free
to encoder
head
dyn.
dyn.
dyn.
head
backbone
Dreamer
–
–
✓
✓
–
–
–
–
TD-MPC2
✓
–
✓
✓
–
–
–
–
DINO-WM / V-JEPA 2-AC
✓
–
–
✓
–
–
–
–
LeWM
✓
–
–
✓
–
–
–
–
UWM
–
–
✓
✓
✓
–
✓
✓
Table 1: World models and world action models as published: what shapes the latent and what one network does. LeWM has no policy head. Dreamer [ 13 ] ; TD-MPC2 [ 15 ] ; DINO-WM and V-JEPA 2-AC [ 40 , 2 ] ; LeWM [ 24 ] ; UWM [ 42 ] ; Cosmos Policy [ 19 ] .
Figure 1: The four training modes on one window: the mode is set by what is masked at the transformer’s input. Nothing is decoded to pixels.
Figure 2: LeWAM: a shared ViT-tiny encoder with SIGReg on its single latent per frame, and one bidirectional transformer that serves the four training modes by masking. No pixel decodeding.
Figure 3: The six tasks as the models see them, with their distractors: painted dot and bar on PushT, two wobbling boxes on the robomimic tasks (cameras and resolutions in Appendix B ).
R2=1−MSE/Var
Pearson r
MSE
Model
Task-state ↑
Distractors ↓
Task-state ↑
Distractors ↓
Task-state ↓
Distractors ↑
PushT
Dreamer
0.56 ± 0.01
0.03 ± 0.03
0.78 ± 0.00
0.30 ± 0.01
0.39 ± 0.01
0.96 ± 0.03
LeWM + diff. head
0.54 ± 0.05
0.00 ± 0.02
0.77 ± 0.03
0.06 ± 0.04
0.40 ± 0.05
1.07 ± 0.02
RecWAM
0.64 ± 0.02
0.79 ± 0.03
0.82 ± 0.02
0.89 ± 0.02
0.32 ± 0.02
0.21 ± 0.03
LeWAM (ours)
0.81 ± 0.01
0.00 ± 0.02
0.92 ± 0.00
0.06 ± 0.02
0.16 ± 0.01
1.19 ± 0.02
Table 2: Linear probe metrics across all tasks, 5 seeds (Negative R2 clipped to 0).
Figure 4: Imagined rollouts with LeWAM using ground truth actions. The top image is for PushT and the bottom one is for the Lift Robomimic task. The first 3 frames are context frames that is passed as input to the model. The 5 frames that follow are imagined frames by rolling out the LeWAM’s forward mode using ground truth actions.
Figure 5: Closed-loop success of all models and planners, over 5 training seeds × 50 evaluation seeds (250 rollouts).
Task
Variant
R2 state ↑
Policy ↑
DS-CEM ↑
Lift
All four modes
0.423±0.019
94.0±4.2
97.6±0.8
No forward
0.323±0.045
91.6±2.3
95.6±2.9
No inverse
0.386±0.031
92.0±5.5
95.6±2.9
No policy
0.359±0.019
–
0.8±1.6
No backward
0.370±0.056
88.0±4.6
93.6±2.3
Can
All four modes
0.839±0.004
82.8±6.1
78.8±4.5
Table 3: LeWAM mode ablation. One mode removed at a time from the four-mode model (41M, object-distractor datasets, 300 epochs). R2 : ridge probe from the frozen latent to task state, computed as 1−MSE/Var on z-scored targets. Policy / DS-CEM: closed-loop success (%), 50 episodes per cell. The no-policy variant has no policy head, so its DS-CEM column steers the inverse-trained flow head.
Task
With A shell
Without shell
Lift
97.6±0.8
93.0±3.0
Can
78.8±4.5
23.0±8.0
Square
87.2±4.7
75.0±3.0
Table 4: Projection ablation. DS-CEM closed-loop success (%) with LeWAM, with and without projecting proposals onto the A shell before decoding ( Section 4.4 ). mean ± std over seeds 0–4, 50 episodes per cell.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Average closed-loop performance of all models/MPC, for each task, over 5 training seeds, 50 evaluation seeds: 250 unique rollouts.
data
200 demonstrations per task, frames resized once to 224 px, frameskip 5, windows of 3 context frames + 1 target, 5-step action chunks (35-D; 70-D on Transport), z-scored
schedule
300 epochs, batch 128, AdamW (lr 5×10−5 , weight decay 10−3 ), linear warm-up then cosine, bf16, gradient clip 1.0
encoder
ViT-tiny/14 at 224 px, [CLS] token, projector MLP (hidden 2048) to a 384-D latent
regularizer
SIGReg on the frame latent, weight 0.36 (0.09 × the four modes), 17 knots, 1024 projections
World Action Models (WAMs) jointly model action generation and environment dynamics and are mostly built on pretrained Video Diffusion Models (VDMs). In VDM-based WAMs, observations are first encoded by a VAE, and the resulting compressed latents are then processed by large video diffusion backbones to extract effective features for action generation. However, this paradigm ties WAM performance and training cost to large-scale video generation pretraining, limiting WAM efficiency and scalability. In this paper, we theoretically and empirically investigate how visual representations affect action generation in WAMs. Our results show that predictive embeddings from Joint-Embedding Predictive Architecture (JEPA) encoders better support action generation than compressed VAE latents, with I-JEPA performing best in our encoder comparison. Based on these findings, we propose LeWAM, which conditions action generation on JEPA embeddings and models environment evolution by predicting future embeddings in the same space, without relying on a video diffusion backbone. We further find that imitation learning matches demonstrated actions but does not distinguish better actions from worse ones, even though small action deviations can greatly affect task success. To address this limitation without additional environment interaction or the human oversight required for resets and safety, we introduce Demonstration-Guided DPO (DemoDPO), an offline preference refinement stage that derives preference supervision directly from demonstrations. With only 0.4B trainable parameters, LeWAM achieves an average success rate of 92.28% on RoboTwin 2.0, comparable to that of state-of-the-art VLAs and WAMs, and maintains practical effectiveness on real-world manipulation tasks.
Xueji Fang, Boqiang Duan, Hua Wu +2
Zhejiang University · Westlake University · Baidu Inc.
Robust robot control benefits from explicitly modeling state transitions, but video-generation world action models (WAMs) introduce substantial deployment cost. Existing latent WAMs avoid explicit future generation, but often compress predictive representations or separate predictive modeling from the representations used for action generation. We introduce JEPA-WAM, a latent WAM built in a pretrained V-JEPA space, which couples latent transition prediction with continuous action generation through a shared predictor. JEPA-WAM predicts a spatially structured joint current-future target that captures task-shared visual temporal structure between current and future observations, while preserving dense patch-level correspondence. Through the shared predictor, transition supervision directly shapes the backbone, from which dedicated representations are extracted for action prediction. The same design can also be instantiated in pretrained VLA policies while preserving their original perception and action pathways. On LIBERO-Plus, JEPA-WAM achieves 79.2%, the best result without large-scale robot-policy pretraining, while its pretrained π0.5 instantiation reaches 86.3%, achieving the best overall performance. Experiments on RoboTwin 2.0 and real-world bimanual manipulation further demonstrate strong generalization under visual and spatial shifts.
Yihan Lin, Jiawei He, Shifeng Bao +6
School of Information, Renmin University of China, Beijing, China · XYZ Embodied AI, Beijing, China · Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China +3
World Action Models (WAMs) have emerged as a powerful paradigm for embodied intelligence, yet the prevailing reliance on pixel-level video generation creates a fundamental bottleneck. Forcing models to reconstruct task-irrelevant visual details dissipates representational capacity and renders policies vulnerable to visual distractors. In this paper, we propose LeapBot-WA, which establishes a novel Predictive-Latent paradigm for WAMs by operationalizing the Joint-Embedding Predictive Architecture (JEPA) as a World-Anchor. Departing from the traditional reliance on visual synthesis, LeapBot-WA shifts the core of world modeling to Predictive Semantic Alignment, extracting abstract physical dynamics directly within a latent foundation space. To bridge the modality gap between non-Gaussian predictive features and diffusion priors, we introduce the Isotropic Semantic Autoencoder (ISAE), which reshapes the anchor's latent space into a diffusion-friendly manifold to prevent off-manifold drift. Furthermore, we design an Asymmetric Mixture-of-Transformers (MoT) architecture. During training, an Anchor Diffusion Transformer acts as a privileged dynamics expert to guide the Action Diffusion Transformer; at inference, this heavy dynamics branch is pruned, enabling zero-overhead execution. LeapBot-WA achieves state-of-the-art performance among predictive models on LIBERO and matches top-tier generative WAMs on RoboTwin 2.0 without requiring large-scale trajectory pre-training. It further demonstrates superior zero-shot robustness to unseen environments and successful real-world transfer, establishing a highly efficient and robust latent-centric paradigm for scalable robotic control. Code: https://github.com/LeapWM/leapbot-wa.
Pei Liu, Nan Zheng, Lang Zhang +8
1The Hong Kong University of Science and Technology (Guangzhou) · 2The Hong Kong University of Science and Technology · 3Leapmotor +1