Video world models aim to generate explorable, 3D-consistent scene videos conditioned on camera trajectories. Existing approaches often rely on external memory systems that explicitly retrieve previously observed content to mitigate scene drift during long-horizon generation. However, these auxiliary memory pathways operate outside the model's internal generative dynamics, preventing the model from intrinsically learning when and what historical information should be retrieved. We propose to internalize memory retrieval into the generation process, allowing retrieval to emerge as an intrinsic behavior of the video world model rather than relying on an external memory system. Based on this principle, we introduce \textbf{Learning-to-Retrieve (L2R)}, which repurposes the model's persistent internal state as a memory for historical context. A camera-conditioned retrieval gate selectively accesses relevant historical information from this state, determining \textit{what to retrieve}, while a retrieval trigger determines \textit{when to retrieve}. We further supervise the trigger with a 3D re-visibility signal, activating retrieval when previously observed content re-enters the current view while otherwise preserving the existing context. Together, these components enable the model to intrinsically acquire memory retrieval behavior and incorporate relevant historical observations into generation without a separate retrieval pathway. Across multiple base models and camera-revisit benchmarks, L2R improves long-term scene consistency while eliminating the need for an external memory bank or 3D conditions. https://jkhu29.github.io/l2r
Figures & tables
Figure 2: Three stages to compute re-visibility supervision. (1) estimate per-frame depth; (2) unproject each token to 3D points; and (3) re-project previous frames’ 3D points into the current frame. If a previously observed 3D point leaves the frustum (or is occluded) and later reappears, assign yt=1 to the corresponding token.
Method
Replay
Loop closure
O-V ↑
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
I2V baseline (Wan2.1)
10.03
0.398
0.534
10.32
0.291
0.643
12.25
Spatial Memory
13.60
0.411
0.554
10.04
0.260
0.617
6.00
VideoSSM
12.69
0.344
0.581
12.23
0.298
0.584
34.75
ARL 2∗
12.73
0.326
0.594
12.13
0.309
0.633
43.06
Context as Memory, K=5
11.92
0.408
0.501
10.72
0.307
0.596
50.75
Table 1: Quantitative comparison on memory capacity. Best in bold , second best underlined . ∗ denotes methods are retrained because their released weights were not trained on dataset ( Yu et al., 2025a ) . The results for the other methods are taken from Echo-Memory ( King et al., 2026 ) .
Figure 3: Qualitative comparison on memory capacity. Red boxes highlight the revisited regions. L2R-GLA has the best scene consistency.
Method
Visual quality ( ↑ )
Camera control ( ↓ )
Consistency ( ↑ )
Subject
Bg
Motion
Temporal
Aesthetic
Imaging
Rerr
Terr
PSNR
SSIM
Consist
Consist
Smooth
Flicker
Quality
Quality
ViewCrafter
78.51
88.30
95.12
92.43
46.28
63.42
8.86
0.482
13.81
0.5217
SEVA
90.53
93.79
96.68
93.41
46.01
64.08
5.26
0.249
16.62
0.5427
VMem
87.71
92.50
94.16
90.41
42.05
61.50
10.31
0.518
16.85
0.5505
SANA-WM
91.73
93.67
96.81
93.99
46.58
66.47
5.40
0.383
15.52
0.5076
Table 2: Quantitative results of video quality, camera controllability, and re-visit consistency on RealEstate10K and DL3DV samples. Best in bold , second best underlined . ∗ marks methods we retrained on our dataset because checkpoints weren’t released; others use official checkpoints.
Variant
Replay
Loop closure
O-V ↑
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
No Memory
10.03
0.398
0.534
10.32
0.291
0.643
12.25
VideoSSM
12.69
0.344
0.581
12.23
0.298
0.584
34.75
Context as Memory, K=5
11.92
0.408
0.501
10.72
0.307
0.596
50.75
GLA
10.55
0.199
0.763
12.74
0.273
0.584
66.13
+ Camera-guided retrieval ( γt )
13.88
0.421
0.560
12.76
0.399
0.687
68.74
Table 3: Ablation on the memory capacity. GLA (baseline) is our starting point. The last three rows progressively add camera-guided retrieval γt , retrieval trigger rt and re-visibility supervision (Eq. 5 ). Best in bold , second best underlined .
Figure 4: Layer-ratio ablation and inference efficiency. Left: L2R-GDN and L2R-GLA with layer ratios of 1:1 , 3:1 , and 7:1 on Echo-Memory benchmark, evaluated by replay and loop-closure PSNR. Right: latency comparison on one H100, 5 s chunks.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Backbone
Wan2.1
Resolution
352×640
Chunk length T
81 frames
Optimizer
AdamW
Learning rate
5×10−5
Per-device batch / grad accumulation
1 / 1
Appendix
Table 7
Setting
Value
Backbone
Wan2.2 / SANA-WM
Resolution
704×1280
Chunk length T
81 frames
Optimizer
AdamW
Learning rate
8×10−5
Per-device batch / grad accumulation
2 / 1
Appendix
Table 8
Figure 5: Failure mode at a high memory-layer ratio. Loop closure on two scenes with memory-layer ratios of 3:1 , 1:1 and 7:1 . With too few memory layers the rollout loses the observed structure on the return.
Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator's limited token budget. Trained jointly with the video generator, a memory encoder and pose-conditioned readout module integrate historical observations into a fixed set of target view-specific tokens before denoising, without explicit depth-based correspondences. By combining this memory with recent temporal context and few-step distillation, WorldCrafter enables streaming scene exploration from a single input image or text prompt. Experiments across static and dynamic scenes show substantial gains in long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploration.
Video world models hold promise for simulating interactive environments, yet maintaining consistent long-term memory across complex camera trajectories remains a critical challenge. Existing methods typically rely on computationally expensive context scaling or rigid heuristic retrieval mechanisms, which lacks generalization to varying camera trajectories and environments. In this paper, we propose Compression and Retrieval (CaR), an attention-driven implicit memory retrieval mechanism to overcome these limitations. By injecting viewpoint information via positional encoding, our method performs flexible memory retrieval through attention computation. To efficiently process extended contexts with minimal computational overhead, we further introduce a lightweight context compression network. Furthermore, we construct SceneFly, a large-scale synthetic dataset featuring realistic camera trajectories and frame-level annotations to train and evaluate long-horizon video world models. Extensive experiments demonstrate that our approach achieves state-of-the-art results on established benchmarks and exhibits strong generalization to open-domain scenes.
Zhan Peng, Jie Ma, Huiqiang Sun +6
Huazhong University of Science and Technology, China · HUJING Digital Media & Entertainment Group, China · Sun Yat-sen University, China
Video World Models are interactive video generation models that predict future world states based on user actions and history video frames. A critical challenge in video world models is the lack of memory, causing inconsistent generated scenes over extended durations. Previous methods explored rule-based context frame retrieval as memory, but they fail to generalize in scenarios with scene occlusions and dynamic objects. We propose MemLearner, a learning-based adaptive context query method using query tokens to bridge context and predicted tokens. By leveraging the video generation model itself for context querying, MemLearner exploits pre-trained visual priors without training additional modules from scratch, and incorporates efficient strategies for training and inference. We collect a dataset of long videos with scene occlusions and dynamic objects, paired with camera pose annotations, and propose a multi-dataset training strategy leveraging both annotated rendered and unannotated real-world videos. Extensive experiments demonstrate that MemLearner significantly outperforms prior video world models in terms of scene consistency and memory, particularly under challenging occlusion and dynamic scenarios.
Jiwen Yu, Jianxiong Gao, Jianhong Bai +7
The University of Hong Kong · ‡Work done during an internship at Kling Team, Kuaishou Technology. · Fudan University +2