cs.CVSep 29, 2026

Honeycomb: Constant-Size Scene Memory Representation for Video World Models

Authors: Jack Wei Lun Shi, Kaichen Zhou, Haoyu Chen, Yufeng Weng, Keane Ong, Ruojin Cai, Hang Hua, Justin K. W. Yeoh, +1 more

Organizations: National University of Singapore · Harvard University · MIT · MIT-IBM Watson AI Lab

Abstract

Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memories accumulate RGB observations or latent features, increasing storage requirements as generation proceeds. We introduce Honeycomb, a video world model built on HexMemory, our proposed low-rank representation for storing scene features in a fixed-size memory with a total of six spatial and spatiotemporal planes. A feed-forward writer maps each generated chunk into new plane features. As the spatial coverage or temporal range expands, we warp the previous planes while preserving their dimensions, then fuse them with the new features through confidence-weighted pooling and a learned residual correction. A reader retrieves latents from HexMemory to condition subsequent video generation. The writer processes only observations from the new chunk, avoiding per-scene optimization and repeated processing of the full history. Experiments on WorldScore and RealEstate10K demonstrate strong video generation quality and robust revisit consistency while keeping HexMemory feature storage constant throughout generation. Code and additional visualizations are available on our project page at https://jackswl.github.io/honeycomb/.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Latent Spatial Memory for Video World Models

    Jun 8, 2026Weijie Wang, Haoyu Zhao, Yifan Yang +7Video World ModelsVideo Latents

  2. DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory

    May 29, 2026Zhenhao Yang, Xiaoshi Wu, Zhengyao Lv +5Interactive Video GenerationPlayable Video World Generation

  3. MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

    May 30, 2026Shengjun Zhang, Zhang Zhang, Simin Huang +11Video World ModelsLong-Video Benchmarks