Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memories accumulate RGB observations or latent features, increasing storage requirements as generation proceeds. We introduce Honeycomb, a video world model built on HexMemory, our proposed low-rank representation for storing scene features in a fixed-size memory with a total of six spatial and spatiotemporal planes. A feed-forward writer maps each generated chunk into new plane features. As the spatial coverage or temporal range expands, we warp the previous planes while preserving their dimensions, then fuse them with the new features through confidence-weighted pooling and a learned residual correction. A reader retrieves latents from HexMemory to condition subsequent video generation. The writer processes only observations from the new chunk, avoiding per-scene optimization and repeated processing of the full history. Experiments on WorldScore and RealEstate10K demonstrate strong video generation quality and robust revisit consistency while keeping HexMemory feature storage constant throughout generation. Code and additional visualizations are available on our project page at https://jackswl.github.io/honeycomb/.
Figures & tables
Figure 1: Revisiting a scene with constant-size memory. Given a single input frame and a camera trajectory that moves away and returns to the initial pose, Honeycomb generates a video whose final frame is consistent with the input frame.
Figure 2: Memory design comparison. Spatia backprojects RGB observations into a point cloud, while LSM-World accumulates latent points. Honeycomb instead writes to HexMemory with fixed-size feature storage.
Figure 3: Overview of the Honeycomb pipeline. For scene-consistent video generation, we a) initialize a fixed-size HexMemory, b) read it to condition a generated chunk, and then c) write new observations back into the same tensors.
Figure 4: Bilinear splatting within the feed-forward writer. (a) Point pi maps to location qi on plane P and (b) its learned contribution ciP is distributed to four neighboring cells using bilinear weights wim . This operation is repeated for all six planes.
Figure 5: Closed-loop comparison. Honeycomb preserves object appearance and scene layout, while Spatia and LSM-World show changes in furnishings, geometry, and texture.
Method
Average Score
Static Score
Dynamic Score
3D Const
Photo Const
Style Const
Subject Quality
Models with 3D cache
WonderJourney
54.19
63.75
44.63
80.60
79.03
62.82
66.56
WonderWorld
61.79
72.69
50.88
86.87
85.56
70.57
49.81
Spatia
63.21
64.88
61.54
83.26
89.09
83.33
46.66
LSM-World
61.20
62.69
59.70
80.88
76.10
–
–
General video models
Table 1: Evaluation results on WorldScore. The Average Score is the mean of the Static and Dynamic Scores; all remaining metrics are computed by the WorldScore benchmark.
RE10K NVS
WorldScore closed-loop
Method
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR C ↑
SSIM C ↑
LPIPS C ↓
Flow C ↓
ViewCrafter
12.28
0.512
0.571
12.32
0.369
0.574
30.78
FlexWorld
13.17
0.567
0.544
12.86
0.430
0.602
55.77
Voyager
14.67
0.577
0.493
15.99
0.459
0.423
7.11
Spatia
15.58
0.616
0.390
15.67
0.488
0.353
6.64
LSM-World
17.46
0.636
0.452
15.12
0.460
0.463
27.05
Table 2: Novel-view synthesis on RealEstate10K and closed-loop on WorldScore. We report evaluation results of all baseline methods using their default settings.
Figure 6: Visualization of HexMemory. Three spatial and three spatiotemporal planes are shown at initialization and after successive chunks. Dashed boundaries mark the extent of the previous memory after warping into the updated coordinate range.
Resolution
HexMemory (MB) ↓
PSNR C ↑
SSIM C ↑
LPIPS C ↓
512
73.9
17.22
0.504
0.311
384
42.6
17.14
0.503
0.313
256
19.8
17.10
0.500
0.319
128
5.8
16.77
0.490
0.338
Table 3: HexMemory resolution ablation. We vary the plane resolution and evaluate on WorldScore closed-loop samples.
Write time (ms) ↓
WorldScore closed-loop
Writer
Chunk 2
Chunk 5
Chunk 9
PSNR C ↑
SSIM C ↑
LPIPS C ↓
Direct optimization
3217
3228
3130
17.44
0.510
0.299
Replacement
11.2
23.3
40.1
17.16
0.502
0.313
Recurrent
13.1
13.2
13.2
17.22
0.504
0.311
Table 4: Memory writer ablation. We hold the plane resolution fixed and compare write time as well as the WorldScore closed-loop performance.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Extended closed-loop comparisons on WorldScore. For each of the three examples, rows show Spatia, LSM-World, and Honeycomb, respectively.
Figure 8: Extended closed-loop comparisons on RealEstate10K. For each of the four examples, rows show Spatia, LSM-World, and Honeycomb, respectively.
Figure 9: Long-horizon novel-view synthesis on RealEstate10K. Each row shows the camera trajectory, input image, and five generated views for one example. Camera markers progress from gray to black over time and use a fixed viewing orientation relative to each input camera.
Figure 10: Temporal organization of the XT feature plane. The input image initializes memory at τ=0 , and the first chunk adds eight new latent-time samples. Subsequent updates expand the temporal range and resample the accumulated memory within a fixed-resolution grid. Dashed lines mark the previous temporal extent. The annotated RGB views illustrate changes in lateral visibility as the camera advances through the alley.
Figure 11: Progression of HexMemory confidence maps. Columns show initialization from the input frame and the memory state after each of four generated chunks. Colors represent log(1+N) , where N is the accumulated interpolation weight, with a shared color scale across updates for each plane.
Figure 12: Spatial feature organization across memory updates. Columns show initialization from the input frame and the memory state after each of four generated chunks. The first three principal components of each plane’s features are mapped to RGB, using a fixed PCA basis and color normalization derived from its final state. This allows feature patterns to be followed across updates within each plane.