We present World2Motion, a framework that generates scene-aware 3D human motion and corresponding video from a single image and a text prompt. While existing 3D motion generators learn from motion datasets, their generalization is constrained by limited coverage of environments. In contrast, video world models such as Cosmos 3 offer broader environmental priors but are not designed for full-body motion generation; recovering motion from their generated videos requires costly two-stage inference. To address these, we turn Cosmos 3 into a single-stage 3D motion generator. This adaptation has two challenges: the scarcity of paired video--motion data and temporal instability in the generated motion. First, we construct a training dataset combining synthetic video--motion pairs with real videos paired with estimated 3D motion. Second, we propose a shift-decoupled noise schedule that assigns different noise levels to video and motion through shared denoising progress. This design accommodates the different denoising requirements of the two modalities, reducing motion jitter. Experiments on a multi-source interaction benchmark show that World2Motion has better motion--text alignment and scene interaction compared with the evaluated 3D motion generators. It also matches the interaction success rate of the two-stage baseline while achieving approximately 3.3× faster inference. Our project page is available at https://fyantu.github.io/World2Motion/.
Figures & tables
Figure 1: World2Motion adapts a video world model to jointly generate scene-aware 3D human motion and corresponding video from a single image and a text prompt. The example shows a person approaching a stool, turning, and sitting down. The top row shows generated video frames, while the bottom row shows the corresponding 3D motion rendered in a scene for visualization.
Approach
Direct 3D motion output
No 3D object input
Scene conditioning
T2M
✓
✓
I2V + HMR
✓
✓
HOI
✓
✓
World2Motion
✓
✓
✓
Table 1: Motion generation methods comparison.
Figure 2: Overview of World2Motion. From an image and a text prompt, the adapted Cosmos 3 backbone jointly generates 3D human motion and video. CameraHMR estimates the initial pose from the image. The shift-decoupled schedule uses modality-specific noise shifts with shared denoising progress.
Figure 3: Shift-decoupled video–motion denoising. (a) Video and motion noise levels during inference with sv=10 and sm=1 . (b) The corresponding training noise pairs follow a shared curve, with lower motion noise than video noise at intermediate steps. (c) Intermediate video frames and motion states. Rows show denoising steps, and column groups show video times. The chair mesh is reference geometry for visualization.
Figure 4: Qualitative comparison of scene interactions: (a) lifting a plastic box, (b) walking to a chair and sitting, and (c) walking to a sofa and sitting. Each method shows successive motion states; video-generating methods also show corresponding video frames. Human–object interaction baselines include their predicted object trajectories.
Translation
Raw Jerk ↓ ( 103m/s3 )
Contact FSR ↓ (0–1)
Absolute
4.355
0.766
Relative
1.251
0.457
Table 3: Root translation ablation on 35 cases without post-processing. Raw Jerk uses 25 fps ( 103m/s3 ). Contact FSR is the sliding fraction among contacting foot-joint/frame pairs on a normalized skeleton (0–1); its definition differs from Table 2 .
Raw Jerk ( 103m/s3 ) ↓
Contact FSR (0–1) ↓
Training sm
1
10
15
1
10
15
1
1.251
8.003
9.703
0.457
0.986
1.000
10
7.021
1.635
2.557
0.985
0.635
0.830
15
8.413
2.127
1.809
0.981
0.700
0.649
Table 4: Noise-schedule ablation on the same 35 cases, using Raw Jerk and contact-conditioned FSR as defined in Table 3 . These unfiltered ablation scores are not directly comparable to Table 2 . Columns give the inference motion shift sm ; sv=10 is fixed. Bold marks the lowest value per metric within each row.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Qualitative results of World2Motion for box lifting and carrying, sitting on a stool or armchair, and two yoga movements. Each example shows five successive video frames (top) and corresponding 3D motion renderings (bottom), ordered from left to right. The input text prompt appears below each example.
Figure 6: Additional qualitative results of World2Motion for sofa sitting, picking up and carrying a vacuum cleaner, standing up from a stool, and two dance sequences. Each example shows five successive video frames (top) and corresponding 3D motion renderings (bottom), ordered from left to right. The input text prompt appears below each example.
Text-driven 3D human motion generation models face significant challenges in responding to diverse and unconstrained textual prompts, primarily due to the limited availability of 3D motion training data. To address this, we introduce MoVT, a novel framework that effectively leverages the extensive range of human action videos to enhance text-to-motion generation. At the core of our approach is the cross-modal augmented motion tokenizer, which projects discrete 3D motion tokens into the 2D domain. This projection allows us to enrich the motion codebook with complex, real-world motion patterns derived from videos. The enriched discrete tokens are then mapped back to the 3D domain, resulting in aligned 3D and 2D codebooks with an enhanced capacity to represent intricate motions. These enhanced codebooks are integrated into a generative masked transformer, which predicts masked motion token indices in a modality-agnostic manner. This enables the use of text-index pairs, generated from the 2D codebook and annotated motion videos, to further enhance the generator. Extensive empirical evaluations show that MoVT performs favorably against prior state-of-the-art methods across multiple key metrics.
Beibei Jing, Tianle Guo, Youjia Zhang +5
Huazhong University of Science and Technology School of Computer Science and Technology Wuhan, Hubei, China · Zhejiang University School of Software Technology Ningbo, China
We present ScaleMoGen, a scale-wise autoregressive framework for text-driven human motion generation. Unlike conventional autoregressive approaches that rely on standard next-token prediction, ScaleMoGen frames motion generation as a coarse-to-fine process. We quantize 3D motions into compositional discrete tokens across multiple skeletal-emporal scales of increasing granularity, learning to generate motion by autoregressively predicting next-scale token maps. To maintain structural integrity, our motion tokenizers and quantizers are explicitly designed so that discrete tokens at every scale strictly preserve the skeletal hierarchy. Additionally, we employ bitwise quantization and prediction, which efficiently scale up the tokenizer vocabulary to preserve motion details and stabilize optimization. Extensive experiments demonstrate that ScaleMoGen achieves state-of-the-art performance, establishing an FID of 0.030 (vs. 0.045 for MoMask) on HumanML3D and a CLIP Score of 0.693 (vs. 0.685 for MoMask++) on the SnapMoGen dataset. Furthermore, we demonstrate that our skeletal-temporal multi-scale representation naturally facilitates training-free, text-guided motion editing.
Inwoo Hwang, Hojun Jang, Bing Zhou +3
Seoul National University · Snap Inc. · Meta Reality Labs
We present ABot-3DWorld 0, a universal multimodal 3D world model that turns text, image, and video inputs into high-fidelity, explorable 3D worlds. At the heart of our framework is a unified Spatial Generative Primitive (SGP), a compact tuple of a high-quality panorama and a spatial point cloud that delivers an efficient description of any 3D space. Multimodal inputs are first lifted into this primitive; a 3D-consistent panoramic video generator then explores the primitive along a planned trajectory; finally, our panoramic video reconstruction engine converts the generated video into a clean, photorealistic 3D Gaussian Splatting (3DGS) world. This pipeline covers two regimes: rich inputs (multi-view sets, casual video) are lifted into the SGP through a geometry-rigorous recovery that mirrors the observed scene, while a single image or sentence is completed generatively into a creative world. The result is one low-barrier engine for general 3D content creation that further anchors generated worlds to geographic points of interest, enabling map-native spatial exploration at consumer scale. Experiments show that ABot-3DWorld 0 sets the state of the art among open-source methods and demonstrates stronger scene fidelity than Marble under rich multimodal inputs.