Despite rapid progress, world models still lack explicit, persistent structural memory, making it difficult to preserve consistent world structure during continual scene expansion and cross-view revisits. To address this limitation, we present WorldWeave, a world generation framework that decouples world-state maintenance from visual rendering. Specifically, WorldWeave combines continual elevation-map generation with agent-guided scene organization and stitching to build an expandable explicit 3D world that incrementally extends structural memory while preserving existing structure. First, its terrain module uses diffusion-based image outpainting to generate continuous metric elevation maps under neighborhood conditioning and boundary constraints. Next, an agent integrates user intent, terrain evidence, and cross-region connectivity constraints to construct scenes through hierarchical semantic planning, deterministic geometry compilation, and local revision. Finally, during visual generation, planned camera trajectories query world geometry through a read-only interface, producing depth sequences that guide video synthesis without writing the generated results back into the world state. As a result, structural memory remains independent of short-window video generation, enabling continual expansion without predefined map boundaries and providing a consistent geometric basis for observations across trajectories and repeated visits.
Figures & tables
Figure 1: Overview of WorldWeave. Terrain completion extends metric terrain, agent planning organizes and validates coarse-to-fine world layouts, and rendering queries the persistent world state to guide video generation. Previously generated regions are preserved while new regions are progressively added.
Figure 2: WorldWeave pipeline. Neighbor-conditioned terrain generation uses residual HDR encoding and new-side joining. The agent plans regions, districts and asset relations, while deterministic tools compile geometry and provide structural and visual evidence for revision. Validated additions extend persistent world state; camera trajectories query its geometry to condition video generation without writing RGB back into the world.
Figure 3: New-side terrain joining. Old samples remain fixed; newly owned connector cells inherit boundary height and normal derivative. The correction vanishes at the interior support boundary. Schematic.
Figure 4: World-grounded video generation across four scenes. Each row follows a viewing trajectory from left to right, showing revisitation or continued exploration. Insets show depth guidance, and numbered markers identify corresponding scene elements across views.
Method
Visual quality
Memory and camera control
Structural consistency
Imaging quality ↑
Aesthetic quality ↑
Structural memory ↑
Camera compliance ↑
MN-MS ↑
MC-GeCo ↓
MC-MEt3R ↓
MC-GeoCon ↓
SANA-WM
0.739
0.619
0.328
55.5
0.973
0.113
0.203
0.189
Zing
0.763
0.685
0.989
55.8
0.974
0.0698
0.129
0.142
SolarWM-5B
0.755
0.604
0.370
54.7
0.956
0.0786
0.166
0.217
AlayaWorld
0.784
0.641
1.51
56.4
0.896
0.926
0.239
0.295
EVOKE-Turbo
0.784
0.718
1.12
61.6
0.961
0.105
0.182
0.249
Table 1: World-video comparison. The first six methods are open-source world models; the remaining baselines are video models. Bold/underline indicate best/second-best scores; yellow highlights the three largest relative gains over MiniMax-H3 (Base), with signed percentage changes.
Figure 5: Opposite-neighbor completion with canvas (top) and multi-image inputs (bottom). Additional cases: Appendix A.2 .
Table 2: Terrain ablations of (a) elevation encoding, (b) input organization and (c) terrain joining. Δh denotes whole-target RMS modification.
Table 3: User ratings (1–5; higher is better).
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Terrain completion with spatial canvas (left) and separate images (right). Each pair shares known terrain, elevation colors and illumination. Dashed lines mark context–target interfaces; gray regions are unavailable.
Figure 7: Full planning-model comparison: (a) alternative and (b) default planner. Within each map, the shared existing region is on the left and the new region on the right. Terrain and drawing conventions are identical.
Version
Seed
IQ ↑
AQ ↑
MN-MS ↑
MC-GeCo ↓
SMC ↑
Cam ↑
MC-MEt3R ↓
MC-GeoCon ↓
Forest
Reference
–
0.7927
0.5921
0.9853
0.0468
1.0000
97.62
0.1080
0.0768
Seed
0
0.7943
0.5928
0.9834
0.0463
1.0000
97.62
0.1075
0.1002
Seed
1
0.7917
0.5929
0.9825
0.0448
1.0000
97.62
0.1053
0.0957
Seed
2
0.7916
0.5993
0.9828
0.0431
1.0000
97.62
0.1095
0.0902
Seed
42
0.7896
0.5929
0.9847
0.0423
1.0000
97.62
0.1078
0.0860
Appendix
Table 4: Per-task seed comparison. Each scene contains one main-experiment reference and five controlled seed variants. All values are retained; † marks the observed black-sky case.
Seed
IQ
AQ
MN-MS
MC-GeCo
SMC
Cam
MC-MEt3R
MC-GeoCon
0
0.7945
0.6851
0.9811
0.0521
1.7500
97.82
0.1286
0.1369
1
0.7937
0.6883
0.9808
0.0548
0.7000
62.52
0.1248
0.1177
2
0.7932
0.6825
0.9814
0.0512
1.4500
85.17
0.1263
0.1293
42
0.7952
0.6903
0.9820
0.0486
1.1750
73.74
0.1253
0.1220
100
0.7866
0.6850
0.9821
0.0472
1.2000
71.60
0.1356
0.1155
Mean
0.7926
0.6863
0.9815
0.0508
1.2550
78.17
0.1281
0.1243
Appendix
Table 5: Seed-wise means over five tasks and their across-seed dispersion. The five main-experiment references are excluded. SD is the population standard deviation and CV is SD divided by the mean, expressed as a percentage.
Figure 8: Motion and consistency are complementary. AlayaWorld’s low raw MEt3R error coexists with distorted snow texture; Wan3.0’s low raw GeCo error accompanies little camera movement. Times are shown; scores refer to whole videos.
Figure 9: Occlusion-and-return comparison in a farmstead. Rows show different methods and columns follow the video sequence. Numbered markers identify the farmhouse, well and barn for comparison across disappearance and return.
Figure 10: Occlusion-and-return comparison in a town. Rows show different methods and columns follow the video sequence. Markers track the workshop, chimney, crates and loading barn as the camera passes behind the warehouse.
Figure 11: Off-screen-return comparison in a woodland scene. Rows show different methods and columns follow the video sequence. Markers track the ranger cabin, log crossing and wood rack during the turn-away and return sequence.