Recent 3D world models generate photorealistic, explorable scenes that remain frozen in time. OuroWorld is a mask-free framework that turns any static 3D Gaussian Splatting scene into a 3D cinemagraph: a dynamic scene with vivid, diverse motion looping seamlessly from any viewpoint. A vision-language model infers plausible dynamics and guides a video model to synthesize a reference video, which we lift and complete into multi-view videos. To learn from this imperfect supervision, we propose Inconsistency-Robust Periodic 4DGS: a Fourier-series deformation field guarantees looping by construction, while a Grounded Drift Field anchored at the reference view absorbs cross-view inconsistency. Unlike prior Eulerian methods limited to fluid-like motion, we capture general deformation, object motion, and illumination change. We introduce a ground-truth-free evaluation covering vividness, naturalness, loop seam coherence, and scene quality. On 39 reconstructed and generated scenes, OuroWorld outperforms all baselines and wins 70.8%-99.0% of user-study comparisons. Project page: https://ouroworld.userwei.com
Figures & tables
Figure 1 : OuroWorld brings static 3D worlds alive. Left : Given a static 3DGS scene from any source, e.g., the 3D world models HY-World 2.0, Marble, and Lyra 2.0, OuroWorld converts it into a 3D cinemagraph: a dynamic scene with vivid, diverse motion that loops endlessly. Right : free-viewpoint renders along a moving camera, with green boxes highlighting regions with noticeable dynamics. Smoke billows in the volcano (top), the galaxy scatters stardust (middle), and the ships bob up and down on the water (bottom). Motion stays consistent across viewpoints, and the scene at t=T returns to t=0 , so the loop is seamless. See the supplementary videos for full results.
Figure 2
Figure 3 : Overview of OuroWorld. (a) Looping video generation. A VLM selects a reference view rendered from the input 3DGS, infers plausible scene dynamics, and prompts a video generation model to synthesize an approximately looping reference video. (b) Multi-view video generation. A 3D foundation model (VGGT- Ω ) lifts the reference video into a dynamic point cloud, using static multi-view renders to improve depth. The point cloud is rendered from N surrounding viewpoints into incomplete videos with disoccluded holes, which a video inpainting model (TrajectoryCrafter) completes into multi-view videos. (c) Inconsistency-Robust Periodic 4DGS. The 3D cinemagraph is represented as a canonical 3DGS deformed by a Fourier-parameterized Periodic Deformation Field, which loops by construction. A Grounded Drift Field absorbs cross-view inconsistency of the generated videos during training and is discarded at inference.
Figure 4 : Learning deformation from inconsistent multi-view videos. Rows: canonical geometry, deformed scene at inference, per-view training renders, and generated training views. (a) w/o Drift Field: conflicting views are averaged, causing blur and erratic motion. (b) w/ Drift Field: drift on every view absorbs too much of the dynamics and is discarded at inference, reducing motion. (c) w/ Grounded Drift Field (ours): with no drift at the reference view, the deformation field must reproduce the reference dynamics, yielding a sharp scene with consistent motion.
Vividness
Naturalness
Loop Seam Coherence
Scene Quality
Method
Vividness Degree ↑
KVD ↓
Motion-Aware Loop Fidelity ↑
Seam SSIM ↑
Aesthetic Quality ↑
Overall Consist. ↑
Gaussians-to-Life
0.0374
127.47
0.0265
0.9388
0.6602
0.2109
3D Cinemagraphy
0.0066
142.42
0.0134
0.9995
0.6374
0.2110
LoopGaussian
0.0024
154.65
0.0025
0.9998
0.6718
0.2100
3D-MOM
0.0247
155.23
0.0317
0.9158
0.6245
0.2069
Static
Ours
0.0423
108.23
0.0809
0.9965
0.6748
0.2111
Table 2 : Quantitative result. Vividness Degree measures motion, illumination, and appearance change. KVD measures the distance to the natural video distribution. MALF measures whether the scene loops while it actually moves, and Seam SSIM measures the similarity at the seam of consecutive cycles; Aesthetic Quality and Overall Consistency follow VBench to assess scene quality. Bold and underline denote the best and second-best results. Ours achieves the best Vividness, Naturalness, and Aesthetic Quality. Near-static methods (LoopGaussian, 3D Cinemagraphy) trivially reach a high Seam SSIM but near-zero MALF, whereas ours achieves the best MALF with a high Seam SSIM.
Figure 5 : Qualitative comparison under the orbit camera , one scene per source. The leftmost column shows the input 3DGS at t1 ; the others show each method at later time steps t2 and t3 . 3D Cinemagraphy and 3D-MOM produce disocclusion holes and color artifacts, while Gaussians-to-Life and LoopGaussian remain nearly static. OuroWorld produces vivid, view-consistent dynamics such as changing illumination, rippling water, drifting fog, and curtains swaying in the wind. See the supplementary videos for full loops.
Figure 6 : User study. In a two-alternative forced-choice study, participants compare OuroWorld with each baseline under four criteria. Bars show the percentage of votes preferring ours (left) versus the baseline (right). Ours is preferred in 70.8%–99.0% of comparisons across all baselines and criteria, with the largest margins in Vividness. 30 participants took part; see Sec. A.6 for details.
Camera
Method
Subject Consistency ↑
Background Consistency ↑
VoL ↑
Vividness Degree ↑
Static camera
w/o Drift
0.9919
0.9818
883.55
0.0424
w/o Grounding
0.9918
0.9833
882.44
0.0392
w/o Scene-View Consistent Optimization
0.9881
0.9817
876.57
0.0450
Ours
0.9924
0.9831
892.43
0.0423
Orbit camera
w/o Drift
0.9666
0.9615
704.34
0.5113
w/o Grounding
0.9649
0.9604
681.16
0.5080
Table 3 : Ablation on consistency components. As illustrated in Fig. 4 , removing the drift field blurs the scene (low VoL) and lowers both subject and background consistency; removing grounding lets drift absorb the dynamics, reducing the scene motion (lowest Vividness); removing scene-view consistent optimization breaks scene-view consistency and blurs the scene most severely (lowest subject consistency, VoL −22% under orbit). Our full model is best on all orbit-camera metrics.
Figure 7 : Effect of drift. Without drift, fine details degrade over time: the clouds in Iceland Canyon blur and fade away (top row), the galaxy in World becomes blurry (middle row), and the rocks on the ground vanish (bottom row). In contrast, our full model preserves these details while maintaining vivid and coherent motion. The first column shows the full frame at t1 with the zoomed region outlined; the remaining columns zoom into that region at t1 , and at t2 and t3 for each method.
Stage
Hyperparameter
Value
Looping video
VLM / video generator
GPT-5.5 / Seedance 2.0
Conditioning image resolution
1280×720
Duration / frame rate
10 s / 24 FPS
Lifting
Static renders / temporal frames
21 / 24
Input resolution
512×288
Discarded low-confidence points
10% per frame
Table A.1 : Hyperparameters of OuroWorld.
Figure A.2 : Motion masks. (a) The 2D motion mask obtained by segmenting the moving entities in the reference image with SAM 3. (b) The corresponding 3D motion mask.
Method
Scene input
Motion input
Native period
Gaussians-to-Life
3DGS
3D mask + text
0.5 s
3D Cinemagraphy
Single image
2D mask
2.4 s
3D-MOM
Single image
2D mask
2.0 s
LoopGaussian
3DGS
2D mask
2.0 s
Ours
3DGS
None
10 s
Table A.2 : Inputs and native loop periods of all methods. For Gaussians-to-Life, which does not loop, the period is the length of its generated motion.
Static camera
Orbit camera
Method
MV ↑
IV ↑
VV ↑
Vividness ↑
MV ↑
IV ↑
VV ↑
Vividness ↑
Gaussians-to-Life
0.0588
0.0170
0.0364
0.0374
0.9412
0.0837
0.5044
0.5097
3D Cinemagraphy
0.0045
0.0069
0.0086
0.0066
0.9348
0.0765
0.4722
0.4945
LoopGaussian
0.0021
0.0025
0.0027
0.0024
0.9390
0.0823
0.5047
0.5087
3D-MOM
0.0234
0.0248
0.0259
0.0247
0.9114
0.0970
0.4089
0.4724
Ours
0.0413
0.0518
0.0338
0.0423
0.9421
0.1146
0.4784
0.5117
Table A.3 : Vividness breakdown against baselines. MV, IV, and VV denote Motion, Illumination, and Visual Variation; Vividness is their average. Bold denotes the best result.
Static camera
Orbit camera
Method
MV ↑
IV ↑
VV ↑
Vividness ↑
MV ↑
IV ↑
VV ↑
Vividness ↑
w/o Drift
0.0414
0.0519
0.0340
0.0424
0.9420
0.1151
0.4768
0.5113
w/o Grounding
0.0324
0.0528
0.0323
0.0392
0.9421
0.1144
0.4676
0.5080
w/o SVCO
0.0442
0.0450
0.0454
0.0450
0.9409
0.1039
0.4795
0.5081
Ours
0.0413
0.0518
0.0338
0.0423
0.9421
0.1146
0.4784
0.5117
Table A.4 : Vividness breakdown of the ablation. SVCO denotes scene-view consistent optimization. Bold denotes the best result.
Figure A.3 : User study interface. Two videos of the same scene are shown side by side in random order (top), and participants choose one for each of the four criteria (bottom).
We present WorldDirector, a highly controllable video world model framework designed for persistent dynamic object memory and unrestricted viewpoint exploration. Unlike existing world models that entangle physical dynamics with pixel rendering and rely on continuous visual observation to sustain motion, our framework explicitly decouples semantic motion orchestration from visual generation. By leveraging an LLM to coordinate 3D trajectories with camera movements and subsequently employing these orchestrated trajectories as control signals for video generation, our approach ensures strict physical logic and appearance stability, successfully preserving the exact visual identities of dynamic entities even when they re-enter the scene after prolonged periods out of view. Experimental results demonstrate that our method supports the synthesis of complex and extended events with unprecedented controllability and persistent dynamic object memory. Project Page: https://worlddirector.github.io/
We present World2Motion, a framework that generates scene-aware 3D human motion and corresponding video from a single image and a text prompt. While existing 3D motion generators learn from motion datasets, their generalization is constrained by limited coverage of environments. In contrast, video world models such as Cosmos 3 offer broader environmental priors but are not designed for full-body motion generation; recovering motion from their generated videos requires costly two-stage inference. To address these, we turn Cosmos 3 into a single-stage 3D motion generator. This adaptation has two challenges: the scarcity of paired video--motion data and temporal instability in the generated motion. First, we construct a training dataset combining synthetic video--motion pairs with real videos paired with estimated 3D motion. Second, we propose a shift-decoupled noise schedule that assigns different noise levels to video and motion through shared denoising progress. This design accommodates the different denoising requirements of the two modalities, reducing motion jitter. Experiments on a multi-source interaction benchmark show that World2Motion has better motion--text alignment and scene interaction compared with the evaluated 3D motion generators. It also matches the interaction success rate of the two-stage baseline while achieving approximately 3.3× faster inference. Our project page is available at https://fyantu.github.io/World2Motion/.
Fangyuan Tu, Xiangyue Zhang, Yiyi Cai +8
Japan Advanced Institute of Science and Technology · The University of Tokyo · Institute of Science Tokyo +1
Previous works that leverage video models for image-to-3D scene generation often suffer from geometric distortions and blurry content. Using video generation models to implicitly maintain geometric consistency according to a single-frame input is ineffective. In this paper, we present a two-stage method, named GeoWorld, that renovates the image-to-3D scene generation pipeline by providing full-frame geometry features. The first-stage video generation model, followed by a multi-view geometry model, produces full-frame geometry features, which are then used as a mental draft of geometric conditions to aid the second-stage video-generation model. A geometric loss is proposed to impose real-world geometric constraints, and a geometry adaptation module is introduced to ensure the effective utilization of geometry features. Thanks to full-frame geometric modeling, the two smaller video models in our two-stage method can generate higher-fidelity 3D scenes than SOTA methods, while being even faster, e.g. 7.5× faster than Hunyuan-Voyager. Project page: https://peaes.github.io/GeoWorld.
Yuhao Wan, Lijuan Liu, Jingzhi Zhou +6
1VCIP & AAIS, Nankai University · 2ByteDance Inc. · 3Renmin University of China +2