cs.ROSep 30, 2026

Magic-W0: A Structured World-Action Foundation Model for Physical Intelligence

Authors: Xuhua Chen, Zhenhan Yin, Yuan Zhang, Lingfeng Zhang, He Zheng, Tong Mu, Shun Zuo, Dian Zhou, +10 more

Organizations: Magic-Lab Team, Magiclab Robotics Inc.

Abstract

World-action models (WAMs) augment robot policies with action-conditioned environment dynamics, yet existing approaches largely rely on future observation reconstruction or generic latent prediction and lack structured, control-oriented world representations tightly coupled with action generation. We introduce Magic-W0, a world-action foundation model that jointly models structured physical state evolution and continuous actions. Magic-W0 represents interaction as a Structured World Transition consisting of Current State, Transition, and Future State. Current State combines vision-language context with Current 3D Geometry; Transition is represented by 3D Motion capturing action-induced three-dimensional changes; and Future State is represented by Future Semantics describing task-relevant outcomes. To couple prediction and control, we propose a layer-aligned world-action interaction architecture in which evolving action hypotheses condition world-transition prediction, while predicted world representations continuously inform action generation. Magic-W0 is pre-trained on large-scale egocentric human manipulation, UMI, real-robot, and simulation data, with latent supervision for geometry, 3D motion, and future semantics from pre-trained visual models. Inference-time interventions show that structured world representations respond systematically to changes in candidate actions and that action-related information propagates through shared 3D representations into future semantic predictions. On RoboDojo-Sim, Magic-W0 achieves an average Score of 27.10, the highest among the compared WAMs. Across multiple real-robot tasks, it also demonstrates strong downstream performance after fine-tuning with limited downstream data, supporting generalization and rapid adaptation.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Making Foresight Actionable: Repurposing Representation Alignment in World Action Models

    Jun 10, 2026Lu Qiu, Yizhuo Li, Yi Chen +3Efficient World-Action ModelWorld Models

  2. RepWAM: World Action Modeling with Representation Visual-Action Tokenizers

    Jun 11, 2026Junke Wang, Qihang Zhang, Shuai Yang +5Efficient World-Action ModelDiscrete Action Tokenizers

  3. From World Models to World Action Models: A Concise Tutorial for Robotics

    Jul 1, 2026Xiaoxiong Zhang, Xiong Zeng, Wei ZhangWorld ModelsAction Prediction