Video world models aim to generate explorable, 3D-consistent scene videos conditioned on camera trajectories. Existing approaches often rely on external memory systems that explicitly retrieve previously observed content to mitigate scene drift during long-horizon generation. However, these auxiliary memory pathways operate outside the model's internal generative dynamics, preventing the model from intrinsically learning when and what historical information should be retrieved. We propose to internalize memory retrieval into the generation process, allowing retrieval to emerge as an intrinsic behavior of the video world model rather than relying on an external memory system. Based on this principle, we introduce \textbf{Learning-to-Retrieve (L2R)}, which repurposes the model's persistent internal state as a memory for historical context. A camera-conditioned retrieval gate selectively accesses relevant historical information from this state, determining \textit{what to retrieve}, while a retrieval trigger determines \textit{when to retrieve}. We further supervise the trigger with a 3D re-visibility signal, activating retrieval when previously observed content re-enters the current view while otherwise preserving the existing context. Together, these components enable the model to intrinsically acquire memory retrieval behavior and incorporate relevant historical observations into generation without a separate retrieval pathway. Across multiple base models and camera-revisit benchmarks, L2R improves long-term scene consistency while eliminating the need for an external memory bank or 3D conditions. https://jkhu29.github.io/l2r
Figures & tables
Figure 2: Three stages to compute re-visibility supervision. (1) estimate per-frame depth; (2) unproject each token to 3D points; and (3) re-project previous frames’ 3D points into the current frame. If a previously observed 3D point leaves the frustum (or is occluded) and later reappears, assign yt=1 to the corresponding token.
Method
Replay
Loop closure
O-V ↑
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
I2V baseline (Wan2.1)
10.03
0.398
0.534
10.32
0.291
0.643
12.25
Spatial Memory
13.60
0.411
0.554
10.04
0.260
0.617
6.00
VideoSSM
12.69
0.344
0.581
12.23
0.298
0.584
34.75
ARL 2∗
12.73
0.326
0.594
12.13
0.309
0.633
43.06
Context as Memory, K=5
11.92
0.408
0.501
10.72
0.307
0.596
50.75
Table 1: Quantitative comparison on memory capacity. Best in bold , second best underlined . ∗ denotes methods are retrained because their released weights were not trained on dataset ( Yu et al., 2025a ) . The results for the other methods are taken from Echo-Memory ( King et al., 2026 ) .
Figure 3: Qualitative comparison on memory capacity. Red boxes highlight the revisited regions. L2R-GLA has the best scene consistency.
Method
Visual quality ( ↑ )
Camera control ( ↓ )
Consistency ( ↑ )
Subject
Bg
Motion
Temporal
Aesthetic
Imaging
Rerr
Terr
PSNR
SSIM
Consist
Consist
Smooth
Flicker
Quality
Quality
ViewCrafter
78.51
88.30
95.12
92.43
46.28
63.42
8.86
0.482
13.81
0.5217
SEVA
90.53
93.79
96.68
93.41
46.01
64.08
5.26
0.249
16.62
0.5427
VMem
87.71
92.50
94.16
90.41
42.05
61.50
10.31
0.518
16.85
0.5505
SANA-WM
91.73
93.67
96.81
93.99
46.58
66.47
5.40
0.383
15.52
0.5076
Table 2: Quantitative results of video quality, camera controllability, and re-visit consistency on RealEstate10K and DL3DV samples. Best in bold , second best underlined . ∗ marks methods we retrained on our dataset because checkpoints weren’t released; others use official checkpoints.
Variant
Replay
Loop closure
O-V ↑
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
No Memory
10.03
0.398
0.534
10.32
0.291
0.643
12.25
VideoSSM
12.69
0.344
0.581
12.23
0.298
0.584
34.75
Context as Memory, K=5
11.92
0.408
0.501
10.72
0.307
0.596
50.75
GLA
10.55
0.199
0.763
12.74
0.273
0.584
66.13
+ Camera-guided retrieval ( γt )
13.88
0.421
0.560
12.76
0.399
0.687
68.74
Table 3: Ablation on the memory capacity. GLA (baseline) is our starting point. The last three rows progressively add camera-guided retrieval γt , retrieval trigger rt and re-visibility supervision (Eq. 5 ). Best in bold , second best underlined .
Figure 4: Layer-ratio ablation and inference efficiency. Left: L2R-GDN and L2R-GLA with layer ratios of 1:1 , 3:1 , and 7:1 on Echo-Memory benchmark, evaluated by replay and loop-closure PSNR. Right: latency comparison on one H100, 5 s chunks.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Backbone
Wan2.1
Resolution
352×640
Chunk length T
81 frames
Optimizer
AdamW
Learning rate
5×10−5
Per-device batch / grad accumulation
1 / 1
Appendix
Table 7
Setting
Value
Backbone
Wan2.2 / SANA-WM
Resolution
704×1280
Chunk length T
81 frames
Optimizer
AdamW
Learning rate
8×10−5
Per-device batch / grad accumulation
2 / 1
Appendix
Table 8
Figure 5: Failure mode at a high memory-layer ratio. Loop closure on two scenes with memory-layer ratios of 3:1 , 1:1 and 7:1 . With too few memory layers the rollout loses the observed structure on the return.