Streaming world models should render a place consistently across repeated visits. Directly supervising such revisits requires training samples that capture both visits, often spanning minutes. Yet dense attention over the full span incurs quadratic costs, making long-span supervision expensive. Memorizon breaks this coupling: long spans are needed for supervision, but not for attention, since the two visits can share a forward pass without including every intervening frame. A training sample covers a span of any length but is scored only on its last k chunks. Instead of tokenizing the history before them, each scored chunk retrieves its own top-K latents by camera co-visibility, and the union of these requests forms a shared bank. The bank is bounded by kK, so the sequence stays bounded however long the span; at the shortest span the recipe is exactly conventional training. Adding the bank raises the cost of a step once; beyond that, a longer span costs little, and going from 100 to 400 s adds 12% to the step time. Against a sliding-window baseline, retrieval raises revisit consistency on every split, and a span long enough to reach the first visit of each return adds a further 24% to 30%, at some cost in image quality; beyond that span, more length no longer helps. Filling the bank from another episode lowers revisit correlation by 83%, so the model uses what it retrieves. Project page: https://tingtingliao.github.io/memorizon
Figures & tables
Figure 1: Long-horizon training at bounded cost. (a) Ordinary training covers only a short tail of the episode, so the first visit falls outside it; Memorizon samples a span that holds both visits, scores only its tail and supplies the history through a bank B of retrieved latents. (b) Bank size against training span: the bank levels off at every K while the candidate pool grows; at 41 latents there is no history. Shading marks spans beyond 100 s.
Figure 2: The training sequence and what each scored chunk reads. Top: the packed sequence and its RoPE index; bank entries share one index. Grid: the attention mask, one row per scored chunk. Each chunk attends to the first frame, its window of r+1 chunks and its own top- K : 1+K+(r+1)c=15 latents, whatever the span. Drawn for K=6 , r=1 , c=4 .
Figure 3: Retrieval criterion graded against ground truth. Left: fraction of achievable gain over 960 scored chunks, from random ( 0% ) to the best candidates ( 100% ); λ=0.1 and 0.2 are level within the standard error; bars show ±1 standard error over chunks. Right: the frame each rule ranks first for one chunk.
Starting pose
Mid-path
VBench
Method
Revisit ↑
Gain ↑
DINO ↑
D. Gain ↑
Revisit ↑
Gain ↑
DINO ↑
D. Gain ↑
Subj. ↑
Backg. ↑
Imag. ↑
LingBot-World 2.0
0.143
− 0.006
0.526
− 0.001
0.147
0.011
0.546
0.017
0.765
0.852
0.741
DreamX-World 1.0
0.146
0.033
0.547
− 0.045
0.109
0.009
0.591
0.007
0.781
0.868
0.684
HY-WorldPlay 1.5 †
0.243
0.040
0.679
0.025
0.194
0.037
0.674
0.027
0.839
0.908
0.702
Matrix-Game 3.0
0.254
0.075
0.717
0.042
0.283
0.089
0.760
0.100
0.813
0.895
0.761
Infinite-World
0.109
0.019
0.610
0.013
0.168
0.070
0.646
0.052
0.789
0.856
0.746
Table 1: Comparison with open world models on the web photographs, with returns to the starting pose and to mid-path poses scored apart. Mean over five seeds; Table 8 gives the spreads. Memorizon : the 100 s run of Table 2 . DINO and D. Gain use DINO ViT-B/16 cosine in place of pixel correlation. † Tracks our walks only to within 1.1 m and 34∘ .
Figure 4: Comparison with SOTA on a web photograph. Each system is shown late in the walk, when the camera is back near the input view; the input is the reference.
Fidelity to GT
Memory
VBench
Training
PSNR ↑
LPIPS ↓
Revisit ↑
Gain ↑
DINO ↑
Subj. ↑
Backg. ↑
Imag. ↑
Unseen scenes
10 s, no retrieval
10.60 ± .06
0.646 ± .005
0.311 ± .018
0.049 ± .020
0.557 ± .020
0.751 ± .006
0.873 ± .001
0.666 ± .014
10 s, 10 -chunk window
10.47 ± .32
0.662 ± .006
0.363 ± .033
0.114 ± .033
0.572 ± .020
0.726 ± .008
0.863 ± .003
0.643 ± .007
10 s, top- K , no bank
10.71 ± .20
0.630 ± .005
0.503 ± .022
0.255 ± .026
0.735 ± .011
0.757 ± .003
0.873 ± .003
0.675 ± .016
100 s, per-segment retrieval, bank
11.27 ± .22
0.636 ± .013
0.511 ± .040
0.259 ± .023
0.743 ± .032
0.763 ± .003
0.873 ± .002
0.671 ± .022
Table 2: Ablation on training span, bank and retrieval. Rows add one ingredient at a time ( Sec. 4.3 ); spans are drawn uniformly from 10 s to the length given. PSNR and LPIPS need rendered ground truth, which web photographs lack. All runs at step 6,000 ; mean over five rollout seeds, standard deviation of the per-seed means in small type. DINO is Revisit with the cosine of DINO ViT-B/16 features in place of pixel correlation. Best per split in bold, second underlined. Seen scenes are in Table 7 ( App. C.2 ).
Figure 5: Ablation study on five models of Table 2 on a web photograph. Top: back near the input view. Middle: a place a little further along. Bottom: the return to it.
Figure 6: Revisit against return interval over 400 s rollouts on the web photographs; error bars are the standard deviation over five rollout seeds. (a) Revisit per model, grouped by the interval between the two visits; n counts returns per seed. (b) Difference from 10 s top- K on the same returns. Shading marks the first 60 s, which the tables score.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Translations rescaled by
Pose distance D ( w=4 )
Frustum overlap O
0.1×
57.1%
39.5%
1× (as posed)
55.9%
47.4%
10×
46.9%
53.0%
Appendix
Table 3: Sensitivity to pose scale. Achievable gain when translations are rescaled with each cue’s constants held fixed.
Per-chunk retrieval K
1
2
3
4
6
8
12
16
24
Achievable gain (%) ↑
59.6
60.0
58.9
57.4
55.4
54.2
51.1
49.1
45.9
Bank capacity
4
6
10
16
24
40
64
∞
Achievable gain (%) ↑
46.5
46.9
49.1
49.7
50.4
51.4
55.0
55.6
Appendix
Table 4: Retrieval width and bank capacity. Top: each width K against its own ceiling, the K most similar candidates. Bottom: six frames read from a bank that admits frames by coverage, dropping near-duplicates, and is capped at the given size, together with the latents already completed in the chunk’s window, which supply the rest when the cap is below six; ∞ lifts the cap but not the de-duplication, so it differs slightly from K=6 above.
Group
Setting
Value
Model and data
Initialization
Wan2.2-TI2V-5B ( Wan Team, 2025 )
Camera conditioning
PRoPE ( Li et al., 2025a )
Video
864×480 at 16 fps
Training sample
Latents per chunk c
4
Scored chunks k , recent chunks r
10 , 1
Span m (chunks)
U[10,mmax] , mmax=100 for Memorizon ( 41 – 401 latents)
Appendix
Table 5: Settings. Everything the runs of Table 2 share; the ablation rows differ only as Sec. 4.3 describes.
Definition
Position
Heading
Also required
Used for
Corpus return ( App. B.2 )
≤τ=0.5
—
camera away >ε=1.5 in between
could one window hold both visits
Evaluation return ( App. C.1 )
≤0.5
≤15∘
≥8 s apart; away >1.5 or 90∘
Revisit and Gain
Fold-back ( App. C.1 )
an evaluation return whose path back retraces the path out within 0.5
camera-following check
Appendix
Table 6: Three definitions of a return. Distances in the units of the corpus (one unit is 4 m).
Figure 7: Revisits in the training corpus. (a) Shortest return interval for three definitions. (b) Share of training windows of each length that contain a return. (c) Per scene, sorted by median, at τ=0.5 , ε=1.5 . Dashed lines mark spans of 10 , 100 , 200 and 400 s.
Fidelity to GT
Memory
VBench
Training
PSNR ↑
LPIPS ↓
Revisit ↑
Gain ↑
DINO ↑
Subj. ↑
Backg. ↑
Imag. ↑
Seen scenes, new trajectories
10 s, no retrieval
11.48 ± .18
0.610 ± .008
0.333 ± .017
0.074 ± .021
0.633 ± .013
0.801 ± .002
0.882 ± .003
0.648 ± .005
10 s, 10 -chunk window
11.08 ± .14
0.641 ± .008
0.321 ± .019
0.081 ± .017
0.599 ± .008
0.774 ± .004
0.875 ± .002
0.635 ± .005
10 s, top- K , no bank
11.41 ± .28
0.609 ± .008
0.427 ± .026
0.186 ± .026
0.744 ± .011
0.799 ± .006
0.883 ± .004
0.654 ± .008
100 s, per-segment retrieval, bank
11.92 ± .10
0.609 ± .006
0.454 ± .025
0.226 ± .031
0.761 ± .020
0.797 ± .005
0.879 ± .003
0.658 ± .008
Appendix
Table 7: Ablation of Table 2 on seen scenes , held-out episodes of the training scenes along new trajectories. Same runs, protocol and columns; best in bold, second underlined.
Starting pose
Mid-path
VBench
Method
Revisit ↑
Gain ↑
DINO ↑
D. Gain ↑
Revisit ↑
Gain ↑
DINO ↑
D. Gain ↑
Subj. ↑
Backg. ↑
Imag. ↑
LingBot-World 2.0
0.143 ± .026
− 0.006 ± .032
0.526 ± .013
− 0.001 ± .029
0.147 ± .017
0.011 ± .015
0.546 ± .020
0.017 ± .019
0.765 ± .004
0.852 ± .004
0.741 ± .004
DreamX-World 1.0
0.146 ± .009
0.033 ± .021
0.547 ± .016
− 0.045 ± .013
0.109 ± .006
0.009 ± .007
0.591 ± .013
0.007 ± .013
0.781 ± .001
0.868 ± .001
0.684 ± .006
HY-WorldPlay 1.5 †
0.243 ± .014
0.040 ± .034
0.679 ± .022
0.025 ± .031
0.194 ± .011
0.037 ± .010
0.674 ± .004
0.027 ± .004
0.839 ± .002
0.908 ± .001
0.702 ± .008
Matrix-Game 3.0
0.254 ± .018
0.075 ± .017
0.717 ± .013
0.042 ± .019
0.283 ± .012
0.089 ± .020
0.760 ± .016
0.100 ± .016
0.813 ± .004
0.895 ± .002
0.761 ± .003
Infinite-World
0.109 ± .022
0.019 ± .031
0.610 ± .026
0.013 ± .022
0.168 ± .017
0.070 ± .012
0.646 ± .013
0.052 ± .007
0.789 ± .005
0.856 ± .003
0.746 ± .006
Appendix
Table 8: Table 1 with its spreads. Mean over five rollout seeds, with the standard deviation of the per-seed means in small type.
Interval between the visits
Type
Training
8 – 10 s
10 – 20 s
20 – 60 s
Start
Mid
All
Seen scenes
10 s, no retrieval
0.230
0.064
0.073
0.160
0.067
0.078
10 s, top- K
0.397
0.174
0.185
0.130
0.199
0.191
100 s, top- K , bank
0.337
0.216
0.253
0.213
0.252
0.248
Unseen scenes
Appendix
Table 9: Gain by return interval and type. Pixel Gain with the five rollout seeds pooled; seen and unseen scenes as in Table 2 , web photographs as in Table 1 . Start : the earlier frame lies in the first second; mid : every other return. Pooling pairs rather than clips makes all differ slightly from Table 2 . For the 100 s run the three interval bins hold 25 , 140 and 470 returns on seen scenes, 30 , 105 and 230 on unseen ones and 20 , 100 and 605 on web photographs.
Training
Sequence slots
Step time (s)
Wall clock (h)
Peak mem. (GB)
10 s, no retrieval
41 / 41
3.49
5.8
16.05
10 s, 10 -chunk window
41 / 41
4.02
6.7
16.05
10 s, top- K , no bank
41 / 41
3.68
6.1
16.05
100 s, per-segment retrieval, bank
51 / 51
4.17
6.9
17.99
100 s, top- K , ordered bank
65 / 105
8.18
13.6
28.45
100 s, top- K , bank
65 / 105
8.02
13.4
28.45
Appendix
Table 10: Training cost of the runs in Table 2 , from their logs. Slots: median and maximum over logged steps. Wall clock is the step time over the 6,000 steps every run trains for.
System
Correlation ↑
Direction ↑
Implied hfov †
Calibration
Rendered frames, engine poses
0.906
0.974
89.1∘
Open world models
LingBot-World 2.0
0.855
0.966
73.3∘
DreamX-World 1.0
0.825
0.955
132.5∘
HY-WorldPlay 1.5
0.697
0.909
123.6∘
Appendix
Table 11: Camera following : yaw only, web photographs, first 60 s, rollout seed 0 . † Field of view implied by image shift per degree of prescribed yaw; the paths assume 90∘ . Memorizon : mean over six runs of Table 2 , the three trained on 10 s clips and the 100 , 200 and 400 s runs.
Starting pose
Mid-path
Bank contents
Revisit ↑
Gain ↑
DINO ↑
D. Gain ↑
Revisit ↑
Gain ↑
DINO ↑
D. Gain ↑
Top- K (ours)
0.377
0.213
0.794
0.218
0.440
0.319
0.826
0.256
Empty
0.176
0.053
0.553
0.085
0.105
0.019
0.504
0.035
K random frames
0.153
0.055
0.583
0.044
0.207
0.109
0.634
0.087
K frames, another episode
0.071
0.012
0.445
− 0.020
0.073
0.016
0.459
− 0.014
Appendix
Table 12: Bank contents at inference. The 100 s run with only the bank’s frames replaced: its top- K (ours), nothing, K random frames from the same history, or K frames from another episode. Web photographs, first 60 s, returns split as in Table 1 , DINO columns as there; rollout seed 0 .
Figure 8: More returns from open world models , extending Figure 4 to three more web photographs. Each block follows one photograph’s walk. Blue tags mark returns to the starting pose; orange tags the two ends of a return to a pose met mid-path. Rollout seed 0 .
Imaging ↑
Subject ↑
Training
Generated
GT bank
GT context
Generated
GT bank
GT context
Seen scenes
10 s, no retrieval
0.648 ± .005
–
0.671 ± .001
0.801 ± .002
–
0.820 ± .000
10 s, top- K , no bank in training
0.654 ± .008
0.663 ± .004
0.673 ± .000
0.799 ± .006
0.811 ± .002
0.824 ± .000
100 s, top- K , bank
0.646 ± .015
0.666 ± .004
0.668 ± .000
0.804 ± .001
0.818 ± .002
0.820 ± .000
200 s, top- K , bank
0.637 ± .012
0.656 ± .003
0.666 ± .000
0.806 ± .006
0.820 ± .001
0.821 ± .000
Appendix
Table 13: Quality against what fills the context on seen and unseen scenes, first 60 s, five seeds. Generated : ordinary rollout. GT bank : the bank holds ground-truth latents. GT context : bank and recent window both ground truth. The ground-truth videos themselves score 0.678 and 0.803 on seen scenes and 0.719 and 0.774 on unseen ones. With a clean history each chunk generates only four latents, so the seed spread rounds to zero.
Figure 9: More returns from Memorizon on six web photographs. Blue: a pose close to the first frame’s. Orange and green: two returns, each shown at its first visit and on coming back; they meet the Revisit rule of Sec. 4.1 , or, where a path has only one such return, the same rule at twice its tolerance.
When a camera revisits a previously observed region, a video world model should reproduce what was there before. This requires both remembering past observations and retrieving the right one for the current viewpoint. Key-value caches preserve visual detail but grow with video length; recurrent memory is compact but compresses history into a fixed-size state, so individual past observations are no longer directly accessible. We introduce LOCI, a hybrid spatial-memory architecture that keeps both representations. In half of the transformer blocks, main attention keeps a key-value cache of past observations; in the other half, it is restricted to the current chunk and complemented by a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry, so viewpoint enters both memory addressing and stored content. Recurrent readouts flow into subsequent cache-backed blocks and supply their queries with accumulated scene context. On the public MIND memory benchmark and on held-out recorded trajectories, LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model; with full history, it lowers peak memory at equal length by about 30% relative to full softmax. With a bounded bank of retained observations, it streams long videos at constant memory and remains more faithful than full softmax under the same budget.
Ji Xia, Tingting Liao, Xuezhi Liang +2
Institute of Foundation Models, Mohamed bin Zayed University of Artificial Intelligence · Mohamed bin Zayed University of Artificial Intelligence · Pinscreen
We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward previously generated frames. However, we find that models can no longer reliably address stored content once rollouts extend beyond the training horizon, because temporal Rotary Positional Embeddings (RoPE) offsets then fall outside the range seen during training and the model struggles to retrieve the relevant visual information through attention. Moreover, naively compressing the cache in the RoPE-rotated space corrupts memory by averaging together incompatible positional phases. To address this, we propose WorldTrace, a training-free memory framework for long-horizon visual persistence. WorldTrace keeps compressed memory addressable by assigning each summary slot a distinct, in-distribution virtual position. Within this addressable cache, we study two memory compression approaches: WorldTrace-Field compresses history for temporal coherence, while WorldTrace-Landmark stores verbatim scene traces at detected transitions for episodic recall. We further introduce LoopBench, a benchmark evaluating whether a compressed cache can reconstruct a previously visited scene after a long detour. WorldTrace-Field improves temporal consistency by +15.5%, and WorldTrace-Landmark improves episodic recall by +19.5% on LoopBench, extending visually persistent generation without retraining.
Xindi Wu, Sven Elflein, James Lucas +5
1NVIDIA · 2Princeton University · University of Toronto +1
Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator's limited token budget. Trained jointly with the video generator, a memory encoder and pose-conditioned readout module integrate historical observations into a fixed set of target view-specific tokens before denoising, without explicit depth-based correspondences. By combining this memory with recent temporal context and few-step distillation, WorldCrafter enables streaming scene exploration from a single input image or text prompt. Experiments across static and dynamic scenes show substantial gains in long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploration.