cs.CVSep 29, 2026

World2Motion: Turning Video World Models into 3D Human Motion Generators

Authors: Fangyuan Tu, Xiangyue Zhang, Yiyi Cai, Yichen Peng, Kunhang Li, Bo Zheng, Zhixiang Wang, Kaipeng Zhang, +3 more

Organizations: Japan Advanced Institute of Science and Technology · The University of Tokyo · Institute of Science Tokyo · Alaya Lab

Abstract

We present World2Motion, a framework that generates scene-aware 3D human motion and corresponding video from a single image and a text prompt. While existing 3D motion generators learn from motion datasets, their generalization is constrained by limited coverage of environments. In contrast, video world models such as Cosmos 3 offer broader environmental priors but are not designed for full-body motion generation; recovering motion from their generated videos requires costly two-stage inference. To address these, we turn Cosmos 3 into a single-stage 3D motion generator. This adaptation has two challenges: the scarcity of paired video--motion data and temporal instability in the generated motion. First, we construct a training dataset combining synthetic video--motion pairs with real videos paired with estimated 3D motion. Second, we propose a shift-decoupled noise schedule that assigns different noise levels to video and motion through shared denoising progress. This design accommodates the different denoising requirements of the two modalities, reducing motion jitter. Experiments on a multi-source interaction benchmark show that World2Motion has better motion--text alignment and scene interaction compared with the evaluated 3D motion generators. It also matches the interaction success rate of the two-stage baseline while achieving approximately 3.3×\times faster inference. Our project page is available at https://fyantu.github.io/World2Motion/.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MoVT: Video-Augmented Motion Tokenizer for Text-to-Motion Generation

    Sep 14, 2026Beibei Jing, Tianle Guo, Youjia Zhang +5Human Motion GenerationText-To-Motion Generation

  2. ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation

    May 12, 2026Inwoo Hwang, Hojun Jang, Bing Zhou +3Human Motion GenerationMotion Generation

  3. ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space

    Jul 13, 2026Mingchao Sun, Luyang Tang, Yu Liu +343D WorldWorld