From World Models to World Action Models: Rethinking Next-State Prediction
Organizations: CASIA · UCAS · BUPT · Yinwang Intelligent Technology Co. Ltd. · XJTU · WHU · THU
Abstract
Predicting the next state is a core paradigm of World Models for modeling physical dynamics, emphasizing prediction fidelity. As World Models evolve into World-Action Models (WAMs), existing methods still fix the next state before training as RGB, a single latent feature, or a static combination of predefined targets, thereby constraining action learning to the inductive biases preserved by a particular representation. To address this limitation, we propose CF-WAM, a dynamic next-state prediction framework that samples visual, semantic, geometric, and interaction projections of the same future, standardizes them into a common video form, and supervises a unified WAM across these projections. The action-relevant constraints exposed by these projections accumulate across training steps, forcing WAM to capture the underlying state-transition structure that supports multiple projections of the same action-conditioned future. This dynamic mechanism also provides a natural cross-embodiment dynamics reference frame for Human and Robot learning. By jointly learning across different next-state parameterizations, heterogeneous Human and Robot experience can bypass appearance differences and directly contribute to shared state-transition learning, improving cross-embodiment generalization. Experiments show that CF-WAM improves both training efficiency and final control performance, while translating Human experience effectively into policy gains. CF-WAM achieves state-of-the-art performance on RoboCasa-GR1 with an average success rate of 82.50%, while reaching 82.65% on LIBERO-Plus and up to 84.00% in real-world evaluations.
Figures & tables
| Group | Method | Camera | Robot | Language | Light | Background | Noise | Layout | Avg. |
|---|---|---|---|---|---|---|---|---|---|
| VLA | GWM-VLA | 57.90 | 54.70 | 89.80 | 95.40 | 90.80 | 72.50 | 77.10 | 76.90 |
| FoMoVLA | 64.00 | 62.20 | 94.00 | 94.10 | 96.20 | 82.20 | 79.60 | 80.50 | |
| WAM | FastWAM | 16.26 | 44.13 | 66.82 | 79.77 | 52.60 | 38.54 | 61.38 | 51.36 |
| GaussianWAM | 79.11 | 56.52 | 92.18 | 89.84 | 66.17 | 86.26 | 70.55 | 77.30 | |
| Ours | CF-WAM | 66.92 | 84.90 | 90.50 | 98.25 | 67.84 | 88.94 | 81.11 | 82.65 |
| Method | Chemistry | Folding | Organization | Fruit | Insertion | Wiping | Avg. |
|---|---|---|---|---|---|---|---|
| 48.00 | 72.00 | 56.00 | 74.00 | 42.00 | 82.00 | 62.33 | |
| GR00T-N1.6 | 32.00 | 44.00 | 38.00 | 72.00 | 34.00 | 64.00 | 47.33 |
| FastWAM | 16.00 | 62.00 | 40.00 | 46.00 | 22.00 | 64.00 | 41.67 |
| LingBot-VA | 28.00 | 12.00 | 42.00 | 58.00 | 24.00 | 58.00 | 37.00 |
| CF-WAM (Visual) | 76.00 | 92.00 | 74.00 | 88.00 | 76.00 | 96.00 | 83.67 |
| CF-WAM (Semantic) | 64.00 | 78.00 | 76.00 | 90.00 | 62.00 | 88.00 | 76.33 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Configuration | Video Expert | Action Expert |
|---|---|---|
| Backbone | Wan2.2-TI2V-5B DiT | ActionDiT |
| Transformer layers | 30 | 30 |
| Residual width | 3072 | 1024 |
| Attention width | 3072 | 3072 |
| Attention heads | ||
| Positional encoding | 3D RoPE | 1D RoPE |
| Stream | Input and processing | Tokens |
|---|---|---|
| Video input | 5 frames, | – |
| VAE latent | 48 channels, | – |
| Video tokens | Patchify | 392 |
| Action chunk | 16 steps, 47 dimensions per step | 16 |
| Joint self-attention | Video + action tokens | 408 |
| Language context | umT5-XXL, up to 128 tokens | 128 |
| Item | Configuration |
|---|---|
| Optimizer | AdamW |
| Initial learning rate | |
| Learning-rate schedule | Cosine decay |
| Weight decay | |
| Numerical precision | bf16 |
| Distributed optimization | ZeRO-1 |
| Source | Data type | Physical episodes | Modality episodes | Representations | FPS | Action |
|---|---|---|---|---|---|---|
| EgoDex | Real egocentric human | 23,849 | 95,396 | RGB/Depth/Seg/DINOv3-PCA | 20 | 18 valid, padded to 47 |
| Teleop-GR1 | Simulated robot teleoperation | 24,000 | 96,000 | RGB/Depth/Seg/DINOv3-PCA | 20 | 47 dimensions |
| Total | – | 47,849 | 191,396 | 8 modality subsets | 20 | Unified to 47 |
| Task | ABot-M0 | JoyAI-RA | FastWAM | WALA | CF-WAM |
|---|---|---|---|---|---|
| CupToDrawerClose | 48.0 | 48.0 | 32.0 | 86.0 | 60.0 |
| PotatoToMicrowaveClose | 50.0 | 70.0 | 72.0 | 78.0 | 72.0 |
| MilkToMicrowaveClose | 46.0 | 84.0 | 68.0 | 78.0 | 72.0 |
| BottleToCabinetClose | 86.0 | 84.0 | 92.0 | 82.0 | 86.0 |
| WineToCabinetClose | 66.0 | 54.0 | 56.0 | 62.0 | 68.0 |
| CanToDrawerClose | 74.0 | 90.0 | 84.0 | 96.0 | 92.0 |
| Setting | Human Action | Robot Action | Future State | Avg. |
|---|---|---|---|---|
| Joint-only | Joint | Joint | Visual | 76.8 |
| EE+Joint | EE + Joint | EE + Joint | Visual | 75 |
| Human–Robot | EE | EE + Joint | Visual | 78.25 |
| CF-WAM | EE | EE + Joint | Dynamic | 82.50 |
| Setting | Inference | Human-like | Bg. Texture | Lighting | Obj. Inst. | Obj. Pos. | Avg. |
|---|---|---|---|---|---|---|---|
| CF-WAM | Visual | 82.00 | 64.00 | 74.00 | 70.00 | 88.00 | 75.60 |
| Semantic | 74.00 | 70.00 | 74.00 | 64.00 | 84.00 | 73.20 | |
| Geometric | 76.00 | 78.00 | 76.00 | 62.00 | 88.00 | 76.00 | |
| Interaction | 80.00 | 82.00 | 80.00 | 68.00 | 86.00 | 79.20 | |
| w/o Human | Visual | 38.00 | 42.00 | 50.00 | 36.00 | 62.00 | 45.60 |