cs.ROSep 29, 2026

One from Infinity: Actualizing Futures from Pretrained World Models into Robot Actions

Authors: Bang Du, Yichen Xie, Shuqi Zhao, Yuxin Chen, Menglin Wu, Masayoshi Tomizuka

Organizations: University of California, Berkeley · Southern University of Science and Technology · Xi’an Jiaotong University

Abstract

A pretrained video world model admits many plausible futures for a scene, but a robot must realize the exact task-conditioned one. To turn world models into executable robot policies, existing methods fine-tune the heavy world model backbone using large-scale robot data and computational resources. Challenging this status quo, we argue that the expensive part has already been paid in the world model pretraining since the representation space of a video world model lays out the diverse potential futures. In this case, what remains is to select the future that accomplishes the task and to read out the actions that realize it. We formalize this task as actualization, which learns a task-conditioned selection and realization on top of a prior supplied by a frozen world model. This can be solved by a tiny actualizer model. We implement RoboActualizer with as few as 60M parameters on top of a frozen world model encoder. The actualizer is composed of two lightweight DiT experts that jointly predict future latents and actions by flow matching. The model can be trained entirely on a single GPU with 32 GB peak memory. With up to 100x fewer trainable parameters than existing WAMs and VLAs, RoboActualizer reaches great performance on simulation benchmarks including LIBERO, LIBERO-Plus, RoboTwin 2.0 and five tasks on two real-world platforms, with a low latency of 39 ms that allows real-time control.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. FutureWorlds: Learning Robotic World Models from Alternative Futures

    Oct 1, 2026Hao Wu, Shengju Qian, Weiyan Wang +5World ModelsScene Context

  2. Latent Action as Intention Enables Efficient Future Imagination for World Action Models

    Aug 25, 2026Xiang Li, Yupeng Zheng, Songen Gu +11Efficient World-Action ModelLatent Actions

  3. τ0τ_0-WM: A Unified Video-Action World Model for Robotic Manipulation

    May 31, 2026Pengfei Zhou, Shengcong Chen, Di Chen +17Efficient World-Action ModelVideo World Models