cs.CVOct 8, 2026

Learning to Retrieve: Internalizing Memory Retrieval for Video World Models

Authors: JiaKui Hu, Tailai Chen, Yuqi Pan, Xuerui Qiu, Jialun Liu, Xiao Cao, Zhenxin Zhu, Guang Chen, +3 more

Organizations: Peking University · Xiaomi EV · CASIA · UQ · NUS

Abstract

Video world models aim to generate explorable, 3D-consistent scene videos conditioned on camera trajectories. Existing approaches often rely on external memory systems that explicitly retrieve previously observed content to mitigate scene drift during long-horizon generation. However, these auxiliary memory pathways operate outside the model's internal generative dynamics, preventing the model from intrinsically learning when and what historical information should be retrieved. We propose to internalize memory retrieval into the generation process, allowing retrieval to emerge as an intrinsic behavior of the video world model rather than relying on an external memory system. Based on this principle, we introduce \textbf{Learning-to-Retrieve (L2R)}, which repurposes the model's persistent internal state as a memory for historical context. A camera-conditioned retrieval gate selectively accesses relevant historical information from this state, determining \textit{what to retrieve}, while a retrieval trigger determines \textit{when to retrieve}. We further supervise the trigger with a 3D re-visibility signal, activating retrieval when previously observed content re-enters the current view while otherwise preserving the existing context. Together, these components enable the model to intrinsically acquire memory retrieval behavior and incorporate relevant historical observations into generation without a separate retrieval pathway. Across multiple base models and camera-revisit benchmarks, L2R improves long-term scene consistency while eliminating the need for an external memory bank or 3D conditions. https://jkhu29.github.io/l2r

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Compression and Retrieval: Implicit Memory Retrieval for Video World Models

    Jun 22, 2026Zhan Peng, Jie Ma, Huiqiang Sun +6Long-Horizon Video GenerationMemory-Augmented Video Generation

  2. MemLearner: Learning to Query Context memory for Video World Models

    Jun 30, 2026Jiwen Yu, Jianxiong Gao, Jianhong Bai +7Action-Conditioned Video GenerationLong-Horizon Video Generation