Despite rapid progress, world models still lack explicit, persistent structural memory, making it difficult to preserve consistent world structure during continual scene expansion and cross-view revisits. To address this limitation, we present WorldWeave, a world generation framework that decouples world-state maintenance from visual rendering. Specifically, WorldWeave combines continual elevation-map generation with agent-guided scene organization and stitching to build an expandable explicit 3D world that incrementally extends structural memory while preserving existing structure. First, its terrain module uses diffusion-based image outpainting to generate continuous metric elevation maps under neighborhood conditioning and boundary constraints. Next, an agent integrates user intent, terrain evidence, and cross-region connectivity constraints to construct scenes through hierarchical semantic planning, deterministic geometry compilation, and local revision. Finally, during visual generation, planned camera trajectories query world geometry through a read-only interface, producing depth sequences that guide video synthesis without writing the generated results back into the world state. As a result, structural memory remains independent of short-window video generation, enabling continual expansion without predefined map boundaries and providing a consistent geometric basis for observations across trajectories and repeated visits.
Figures & tables
Figure 1: Overview of WorldWeave. Terrain completion extends metric terrain, agent planning organizes and validates coarse-to-fine world layouts, and rendering queries the persistent world state to guide video generation. Previously generated regions are preserved while new regions are progressively added.
Figure 2: WorldWeave pipeline. Neighbor-conditioned terrain generation uses residual HDR encoding and new-side joining. The agent plans regions, districts and asset relations, while deterministic tools compile geometry and provide structural and visual evidence for revision. Validated additions extend persistent world state; camera trajectories query its geometry to condition video generation without writing RGB back into the world.
Figure 3: New-side terrain joining. Old samples remain fixed; newly owned connector cells inherit boundary height and normal derivative. The correction vanishes at the interior support boundary. Schematic.
Figure 4: World-grounded video generation across four scenes. Each row follows a viewing trajectory from left to right, showing revisitation or continued exploration. Insets show depth guidance, and numbered markers identify corresponding scene elements across views.
Method
Visual quality
Memory and camera control
Structural consistency
Imaging quality ↑
Aesthetic quality ↑
Structural memory ↑
Camera compliance ↑
MN-MS ↑
MC-GeCo ↓
MC-MEt3R ↓
MC-GeoCon ↓
SANA-WM
0.739
0.619
0.328
55.5
0.973
0.113
0.203
0.189
Zing
0.763
0.685
0.989
55.8
0.974
0.0698
0.129
0.142
SolarWM-5B
0.755
0.604
0.370
54.7
0.956
0.0786
0.166
0.217
AlayaWorld
0.784
0.641
1.51
56.4
0.896
0.926
0.239
0.295
EVOKE-Turbo
0.784
0.718
1.12
61.6
0.961
0.105
0.182
0.249
Table 1: World-video comparison. The first six methods are open-source world models; the remaining baselines are video models. Bold/underline indicate best/second-best scores; yellow highlights the three largest relative gains over MiniMax-H3 (Base), with signed percentage changes.
Figure 5: Opposite-neighbor completion with canvas (top) and multi-image inputs (bottom). Additional cases: Appendix A.2 .
Table 2: Terrain ablations of (a) elevation encoding, (b) input organization and (c) terrain joining. Δh denotes whole-target RMS modification.
Table 3: User ratings (1–5; higher is better).
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Terrain completion with spatial canvas (left) and separate images (right). Each pair shares known terrain, elevation colors and illumination. Dashed lines mark context–target interfaces; gray regions are unavailable.
Figure 7: Full planning-model comparison: (a) alternative and (b) default planner. Within each map, the shared existing region is on the left and the new region on the right. Terrain and drawing conventions are identical.
Version
Seed
IQ ↑
AQ ↑
MN-MS ↑
MC-GeCo ↓
SMC ↑
Cam ↑
MC-MEt3R ↓
MC-GeoCon ↓
Forest
Reference
–
0.7927
0.5921
0.9853
0.0468
1.0000
97.62
0.1080
0.0768
Seed
0
0.7943
0.5928
0.9834
0.0463
1.0000
97.62
0.1075
0.1002
Seed
1
0.7917
0.5929
0.9825
0.0448
1.0000
97.62
0.1053
0.0957
Seed
2
0.7916
0.5993
0.9828
0.0431
1.0000
97.62
0.1095
0.0902
Seed
42
0.7896
0.5929
0.9847
0.0423
1.0000
97.62
0.1078
0.0860
Appendix
Table 4: Per-task seed comparison. Each scene contains one main-experiment reference and five controlled seed variants. All values are retained; † marks the observed black-sky case.
Seed
IQ
AQ
MN-MS
MC-GeCo
SMC
Cam
MC-MEt3R
MC-GeoCon
0
0.7945
0.6851
0.9811
0.0521
1.7500
97.82
0.1286
0.1369
1
0.7937
0.6883
0.9808
0.0548
0.7000
62.52
0.1248
0.1177
2
0.7932
0.6825
0.9814
0.0512
1.4500
85.17
0.1263
0.1293
42
0.7952
0.6903
0.9820
0.0486
1.1750
73.74
0.1253
0.1220
100
0.7866
0.6850
0.9821
0.0472
1.2000
71.60
0.1356
0.1155
Mean
0.7926
0.6863
0.9815
0.0508
1.2550
78.17
0.1281
0.1243
Appendix
Table 5: Seed-wise means over five tasks and their across-seed dispersion. The five main-experiment references are excluded. SD is the population standard deviation and CV is SD divided by the mean, expressed as a percentage.
Figure 8: Motion and consistency are complementary. AlayaWorld’s low raw MEt3R error coexists with distorted snow texture; Wan3.0’s low raw GeCo error accompanies little camera movement. Times are shown; scores refer to whole videos.
Figure 9: Occlusion-and-return comparison in a farmstead. Rows show different methods and columns follow the video sequence. Numbered markers identify the farmhouse, well and barn for comparison across disappearance and return.
Figure 10: Occlusion-and-return comparison in a town. Rows show different methods and columns follow the video sequence. Markers track the workshop, chimney, crates and loading barn as the camera passes behind the warehouse.
Figure 11: Off-screen-return comparison in a woodland scene. Rows show different methods and columns follow the video sequence. Markers track the ranger cabin, log crossing and wood rack during the turn-away and return sequence.
We present WorldDirector, a highly controllable video world model framework designed for persistent dynamic object memory and unrestricted viewpoint exploration. Unlike existing world models that entangle physical dynamics with pixel rendering and rely on continuous visual observation to sustain motion, our framework explicitly decouples semantic motion orchestration from visual generation. By leveraging an LLM to coordinate 3D trajectories with camera movements and subsequently employing these orchestrated trajectories as control signals for video generation, our approach ensures strict physical logic and appearance stability, successfully preserving the exact visual identities of dynamic entities even when they re-enter the scene after prolonged periods out of view. Experimental results demonstrate that our method supports the synthesis of complex and extended events with unprecedented controllability and persistent dynamic object memory. Project Page: https://worlddirector.github.io/
We present OctWorld, a video diffusion framework with persistent 3D memory for generating explorable, world-consistent, and high-fidelity visual scenes. Given a single image, OctWorld performs stable autoregressive world generation along user-specified camera trajectories. We focus on long-range generation, characterized by extended camera paths and wide viewpoint coverage, where preserving spatial consistency is particularly challenging when previously generated regions are revisited. To address this problem, we introduce OctMap, an extensible and spatially adaptive 3D memory that progressively fuses generated visual observations and their corresponding depth maps into a global representation. OctMap employs TSDF fusion within a dynamic sparse octree whose spatial resolution adapts to image evidence. This design preserves geometric and appearance details across diverse scene scales while maintaining low memory overhead. Experiments demonstrate that OctWorld generates long-range, spatially consistent videos and outperforms prior methods on both existing benchmarks and challenging long-range generation settings. OctMap also provides clear advantages over point-based caches and fixed-resolution TSDF volumes. Project page: https://maxtirerror.github.io/octworldpage/
Zelong Lv, Sicheng Xu, Jianfeng Xiang +5
University of Science and Technology of China · Microsoft Research Asia · Tsinghua University
Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator's limited token budget. Trained jointly with the video generator, a memory encoder and pose-conditioned readout module integrate historical observations into a fixed set of target view-specific tokens before denoising, without explicit depth-based correspondences. By combining this memory with recent temporal context and few-step distillation, WorldCrafter enables streaming scene exploration from a single input image or text prompt. Experiments across static and dynamic scenes show substantial gains in long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploration.