Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memories accumulate RGB observations or latent features, increasing storage requirements as generation proceeds. We introduce Honeycomb, a video world model built on HexMemory, our proposed low-rank representation for storing scene features in a fixed-size memory with a total of six spatial and spatiotemporal planes. A feed-forward writer maps each generated chunk into new plane features. As the spatial coverage or temporal range expands, we warp the previous planes while preserving their dimensions, then fuse them with the new features through confidence-weighted pooling and a learned residual correction. A reader retrieves latents from HexMemory to condition subsequent video generation. The writer processes only observations from the new chunk, avoiding per-scene optimization and repeated processing of the full history. Experiments on WorldScore and RealEstate10K demonstrate strong video generation quality and robust revisit consistency while keeping HexMemory feature storage constant throughout generation. Code and additional visualizations are available on our project page at https://jackswl.github.io/honeycomb/.
Figures & tables
Figure 1: Revisiting a scene with constant-size memory. Given a single input frame and a camera trajectory that moves away and returns to the initial pose, Honeycomb generates a video whose final frame is consistent with the input frame.
Figure 2: Memory design comparison. Spatia backprojects RGB observations into a point cloud, while LSM-World accumulates latent points. Honeycomb instead writes to HexMemory with fixed-size feature storage.
Figure 3: Overview of the Honeycomb pipeline. For scene-consistent video generation, we a) initialize a fixed-size HexMemory, b) read it to condition a generated chunk, and then c) write new observations back into the same tensors.
Figure 4: Bilinear splatting within the feed-forward writer. (a) Point pi maps to location qi on plane P and (b) its learned contribution ciP is distributed to four neighboring cells using bilinear weights wim . This operation is repeated for all six planes.
Figure 5: Closed-loop comparison. Honeycomb preserves object appearance and scene layout, while Spatia and LSM-World show changes in furnishings, geometry, and texture.
Method
Average Score
Static Score
Dynamic Score
3D Const
Photo Const
Style Const
Subject Quality
Models with 3D cache
WonderJourney
54.19
63.75
44.63
80.60
79.03
62.82
66.56
WonderWorld
61.79
72.69
50.88
86.87
85.56
70.57
49.81
Spatia
63.21
64.88
61.54
83.26
89.09
83.33
46.66
LSM-World
61.20
62.69
59.70
80.88
76.10
–
–
General video models
Table 1: Evaluation results on WorldScore. The Average Score is the mean of the Static and Dynamic Scores; all remaining metrics are computed by the WorldScore benchmark.
RE10K NVS
WorldScore closed-loop
Method
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR C ↑
SSIM C ↑
LPIPS C ↓
Flow C ↓
ViewCrafter
12.28
0.512
0.571
12.32
0.369
0.574
30.78
FlexWorld
13.17
0.567
0.544
12.86
0.430
0.602
55.77
Voyager
14.67
0.577
0.493
15.99
0.459
0.423
7.11
Spatia
15.58
0.616
0.390
15.67
0.488
0.353
6.64
LSM-World
17.46
0.636
0.452
15.12
0.460
0.463
27.05
Table 2: Novel-view synthesis on RealEstate10K and closed-loop on WorldScore. We report evaluation results of all baseline methods using their default settings.
Figure 6: Visualization of HexMemory. Three spatial and three spatiotemporal planes are shown at initialization and after successive chunks. Dashed boundaries mark the extent of the previous memory after warping into the updated coordinate range.
Resolution
HexMemory (MB) ↓
PSNR C ↑
SSIM C ↑
LPIPS C ↓
512
73.9
17.22
0.504
0.311
384
42.6
17.14
0.503
0.313
256
19.8
17.10
0.500
0.319
128
5.8
16.77
0.490
0.338
Table 3: HexMemory resolution ablation. We vary the plane resolution and evaluate on WorldScore closed-loop samples.
Write time (ms) ↓
WorldScore closed-loop
Writer
Chunk 2
Chunk 5
Chunk 9
PSNR C ↑
SSIM C ↑
LPIPS C ↓
Direct optimization
3217
3228
3130
17.44
0.510
0.299
Replacement
11.2
23.3
40.1
17.16
0.502
0.313
Recurrent
13.1
13.2
13.2
17.22
0.504
0.311
Table 4: Memory writer ablation. We hold the plane resolution fixed and compare write time as well as the WorldScore closed-loop performance.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Extended closed-loop comparisons on WorldScore. For each of the three examples, rows show Spatia, LSM-World, and Honeycomb, respectively.
Figure 8: Extended closed-loop comparisons on RealEstate10K. For each of the four examples, rows show Spatia, LSM-World, and Honeycomb, respectively.
Figure 9: Long-horizon novel-view synthesis on RealEstate10K. Each row shows the camera trajectory, input image, and five generated views for one example. Camera markers progress from gray to black over time and use a fixed viewing orientation relative to each input camera.
Figure 10: Temporal organization of the XT feature plane. The input image initializes memory at τ=0 , and the first chunk adds eight new latent-time samples. Subsequent updates expand the temporal range and resample the accumulated memory within a fixed-resolution grid. Dashed lines mark the previous temporal extent. The annotated RGB views illustrate changes in lateral visibility as the camera advances through the alley.
Figure 11: Progression of HexMemory confidence maps. Columns show initialization from the input frame and the memory state after each of four generated chunks. Colors represent log(1+N) , where N is the accumulated interpolation weight, with a shared color scale across updates for each plane.
Figure 12: Spatial feature organization across memory updates. Columns show initialization from the input frame and the memory state after each of four generated chunks. The first three principal components of each plane’s features are mapped to RGB, using a fixed PCA basis and color normalization derived from its final state. This allows feature patterns to be followed across updates within each plane.
Video world models that maintain 3D spatial consistency across generated frames typically rely on explicit point cloud memory constructed in RGB space. This design is both computationally expensive, requiring repeated rendering and VAE encoding, and inherently lossy, as the round trip through pixel space discards rich features of the learned latent representation. In this paper, we introduce \emph{latent spatial memory} for video world models, a persistent 3D cache that stores scene information directly in the diffusion latent space, avoiding pixel-space reconstruction. Building on this, we propose Mirage, a latent-space spatial memory framework that constructs the memory by lifting latent tokens into 3D via depth-guided back-projection and queries it by synthesizing novel views through direct latent-space warping. This unified formulation eliminates both the information loss of pixel-space reconstruction and the computational burden of repeated encoding and rendering. Experiments show that latent spatial memory achieves up to \textbf{10.57}× faster end-to-end video generation and \textbf{55}× reduction in memory footprint relative to explicit 3D baselines. Leveraging the geometric prior of the diffusion model, Mirage attains state-of-the-art performance on WorldScore and strong reconstruction quality on RealEstate10K.
Weijie Wang, Haoyu Zhao, Yifan Yang +7
Zhejiang University · Microsoft Research · Adelaide University +1
Recent advances in video generative models have promoted rapid progress in controllable world models. However, maintaining fine-grained spatio-temporal consistency under long-horizon reasoning remains a key challenge. In this work, we move beyond explicit 3D memory and coarse frame-level implicit modeling, and propose a fine-grained, learnable, and scalable memory for consistent world generation. We first identify two fundamental limitations of naïve learnable memory architectures in long-horizon extrapolation, namely computational inefficiency and attention dispersion. Through a systematic analysis of attention dispersion, we propose DecMem, a decoupled memory architecture that employs Sparse Global Memory for efficient fine-grained access to global history and Anchored Local Memory for stable and high-quality extrapolation. Extensive experiments demonstrate that DecMem significantly outperforms current state-of-the-art methods. By ensuring precise and efficient long-term memory and achieving superior extrapolation capabilities, DecMem enables minute-level controllable long video generation with high fidelity and consistency.
Zhenhao Yang, Xiaoshi Wu, Zhengyao Lv +5
1The University of Hong Kong · *Work done during an internship at Kling Team, Kuaishou Tech. · 2Kling Team, Kuaishou Technology
Recent advancements in video-based world models have demonstrated an unprecedented ability to synthesize high-fidelity visual sequences. However, a fundamental gap persists between visually plausible video generation and the functional requirements of a world model, particularly in maintaining a stable and reasonable internal state over extended temporal horizons. While existing benchmarks primarily emphasize visual quality, motion coherence, and text-video alignment, they largely overlook memory, the core capability of a world model to preserve consistency across long-term horizons and complex interactions. To address this gap, we present \textbf{MBench}, a comprehensive benchmark dedicated to quantifying and evaluating the memory capability of video world models. We systematically decompose the memory capability of video world models into three hierarchical and complementary core dimensions: entity consistency, environment consistency, and causal consistency, which are further refined into 12 quantifiable sub-dimensions for comprehensive characterization of long-term memory. Our benchmark is built upon rigorously curated real-captured long videos, and evaluated by rule-based quantitative matrices and VLM to enable objective and comprehensive consistency assessment. Extensive evaluations of mainstream state-of-the-art video world models reveal critical systemic limitations of existing methods in long-term state retention, providing a standardized benchmark and clear research direction to advance the field.
Shengjun Zhang, Zhang Zhang, Simin Huang +11
1Tsinghua University · 2WeChat Vision, Tecent Inc. · 3Peking University