cs.ROOct 8, 2026

MiniWAM: Learning Compact Future Targets for Efficient World-Action Modeling

Authors: Jie Chen, Ruofei Bai, Yuxin Cai, Yifeng Zhang, Chengyang He, Jun Li, Wei-Yun Yau, Guillaume Sartoretti

Organizations: Department of Mechanical Engineering, National University of Singapore, Singapore · A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC), Singapore · Nanyang Technological University, Singapore

Abstract

World modeling has emerged as an effective co-training objective for robot policies, giving rise to World Action Models (WAMs) that jointly predict actions and future states. However, most WAMs predict future states in the native representation space of pretrained visual backbones, resulting in high-dimensional targets with substantial training cost. We introduce MiniWAM, which instead predicts compact future representations learned from privileged current-future transitions. To construct these targets, we propose Predictive Representations via Inverse Spatiotemporal Modeling (PRISM), which combines inverse-dynamics supervision with feature reconstruction to emphasize control-relevant transition information while preserving useful future-state information. With the learned PRISM encoder frozen, MiniWAM is trained to jointly predict the resulting targets and robot actions from current observations. With 65×\times fewer native future feature tokens, MiniWAM consistently outperforms native future-feature prediction with both DINOv3 and WAN2.1 VAE features, while achieving up to an 8×\times speedup in world-action training. At 0.25B parameters, MiniWAM is already competitive with substantially larger WAMs on LIBERO, LIBERO-Plus, and RoboTwin 2.0 simulation benchmarks. Representation analyses further show that PRISM contributes behavioral structure beyond feature reconstruction alone. These results demonstrate that effective world-action modeling does not require predicting native visual futures, and that compact predictive representations provide a strong and substantially more efficient target for policy learning. The project page is available at: https://j1dan.github.io/MiniWAM.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Light-WAM: Efficient World Action Models with State-Fusion Action Decoding

    Jun 6, 2026Ziang Li, Dongzhou Cheng, Yibin Wang +5Robot Policy LearningWorld Models for Robotics

  2. MoWAM: Explicit Future Motion Prediction for Efficient World Action Models

    Sep 17, 2026Jiayu Wang, Bin Zhu, Yue Yu +1World Models for RoboticsWorld Model-Based Planning

  3. Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination

    Jun 8, 2026Jiajun Li, Tiecheng Guo, Yifan Ye +9World Models for RoboticsEfficient Inference for World Action Models