cs.CVOct 5, 2026

KineWorld: Action-Induced Transport Fields for Embodied World Modeling

Authors: Ziying Song, Yuchen Liu, Zhuoran Xu, Ziyang Liu, Jian Jin, Jiangtao Su, Haibao Yu, Lei Yang, +1 more

Organizations: Nanyang Technological University · North University of China · Psibot · The Hong Kong Polytechnic University · China Academy of Information and Communications Technology · University of Hong Kong · Peking University

Abstract

Embodied world models predict the visual consequences of candidate actions before execution. However, existing action-conditioned world models often adopt uniformly weighted visual generation objectives that can be misaligned with embodied prediction needs. Even with explicit motion conditioning, these objectives can underemphasize spatially sparse changes that are critical to interaction. We propose KineWorld, a transport-aware world-modeling framework that extends robot kinematics from motion conditioning to the spatial allocation of generative supervision. Kinematic Transport Lifting (KTL) constructs renderer-derived, camera-aligned transport fields from commanded robot motion. Transport-Aware World Diffusion (TAWD) calibrates their motion support on the video-latent grid and reweights future-RGB flow matching through a normalized mixture of uniform and transport-focused distributions. We train KineWorld using ALOHA-AgileX bimanual manipulation data from RoboTwin 2.0. KineWorld achieves an EWMScore-P of 68.95 in single-view evaluation and a TWB-Score of 54.82 in multi-view evaluation. These results support a shift from appearance fitting toward action-consequence modeling for embodied decision-making.

Figures & tables

Appendix figures & tables22 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. From World Models to World Action Models: A Concise Tutorial for Robotics

    Jul 1, 2026Xiaoxiong Zhang, Xiong Zeng, Wei ZhangWorld ModelsAction Prediction

  2. EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields

    May 7, 2026Zhaoyang Yang, Yurun Jin, Lizhe Qi +2Efficient World-Action ModelRecent World-Action Models