Streaming world models should render a place consistently across repeated visits. Directly supervising such revisits requires training samples that capture both visits, often spanning minutes. Yet dense attention over the full span incurs quadratic costs, making long-span supervision expensive. Memorizon breaks this coupling: long spans are needed for supervision, but not for attention, since the two visits can share a forward pass without including every intervening frame. A training sample covers a span of any length but is scored only on its last k chunks. Instead of tokenizing the history before them, each scored chunk retrieves its own top-K latents by camera co-visibility, and the union of these requests forms a shared bank. The bank is bounded by kK, so the sequence stays bounded however long the span; at the shortest span the recipe is exactly conventional training. Adding the bank raises the cost of a step once; beyond that, a longer span costs little, and going from 100 to 400 s adds 12% to the step time. Against a sliding-window baseline, retrieval raises revisit consistency on every split, and a span long enough to reach the first visit of each return adds a further 24% to 30%, at some cost in image quality; beyond that span, more length no longer helps. Filling the bank from another episode lowers revisit correlation by 83%, so the model uses what it retrieves. Project page: https://tingtingliao.github.io/memorizon
Figures & tables
Figure 1: Long-horizon training at bounded cost. (a) Ordinary training covers only a short tail of the episode, so the first visit falls outside it; Memorizon samples a span that holds both visits, scores only its tail and supplies the history through a bank B of retrieved latents. (b) Bank size against training span: the bank levels off at every K while the candidate pool grows; at 41 latents there is no history. Shading marks spans beyond 100 s.
Figure 2: The training sequence and what each scored chunk reads. Top: the packed sequence and its RoPE index; bank entries share one index. Grid: the attention mask, one row per scored chunk. Each chunk attends to the first frame, its window of r+1 chunks and its own top- K : 1+K+(r+1)c=15 latents, whatever the span. Drawn for K=6 , r=1 , c=4 .
Figure 3: Retrieval criterion graded against ground truth. Left: fraction of achievable gain over 960 scored chunks, from random ( 0% ) to the best candidates ( 100% ); λ=0.1 and 0.2 are level within the standard error; bars show ±1 standard error over chunks. Right: the frame each rule ranks first for one chunk.
Starting pose
Mid-path
VBench
Method
Revisit ↑
Gain ↑
DINO ↑
D. Gain ↑
Revisit ↑
Gain ↑
DINO ↑
D. Gain ↑
Subj. ↑
Backg. ↑
Imag. ↑
LingBot-World 2.0
0.143
− 0.006
0.526
− 0.001
0.147
0.011
0.546
0.017
0.765
0.852
0.741
DreamX-World 1.0
0.146
0.033
0.547
− 0.045
0.109
0.009
0.591
0.007
0.781
0.868
0.684
HY-WorldPlay 1.5 †
0.243
0.040
0.679
0.025
0.194
0.037
0.674
0.027
0.839
0.908
0.702
Matrix-Game 3.0
0.254
0.075
0.717
0.042
0.283
0.089
0.760
0.100
0.813
0.895
0.761
Infinite-World
0.109
0.019
0.610
0.013
0.168
0.070
0.646
0.052
0.789
0.856
0.746
Table 1: Comparison with open world models on the web photographs, with returns to the starting pose and to mid-path poses scored apart. Mean over five seeds; Table 8 gives the spreads. Memorizon : the 100 s run of Table 2 . DINO and D. Gain use DINO ViT-B/16 cosine in place of pixel correlation. † Tracks our walks only to within 1.1 m and 34∘ .
Figure 4: Comparison with SOTA on a web photograph. Each system is shown late in the walk, when the camera is back near the input view; the input is the reference.
Fidelity to GT
Memory
VBench
Training
PSNR ↑
LPIPS ↓
Revisit ↑
Gain ↑
DINO ↑
Subj. ↑
Backg. ↑
Imag. ↑
Unseen scenes
10 s, no retrieval
10.60 ± .06
0.646 ± .005
0.311 ± .018
0.049 ± .020
0.557 ± .020
0.751 ± .006
0.873 ± .001
0.666 ± .014
10 s, 10 -chunk window
10.47 ± .32
0.662 ± .006
0.363 ± .033
0.114 ± .033
0.572 ± .020
0.726 ± .008
0.863 ± .003
0.643 ± .007
10 s, top- K , no bank
10.71 ± .20
0.630 ± .005
0.503 ± .022
0.255 ± .026
0.735 ± .011
0.757 ± .003
0.873 ± .003
0.675 ± .016
100 s, per-segment retrieval, bank
11.27 ± .22
0.636 ± .013
0.511 ± .040
0.259 ± .023
0.743 ± .032
0.763 ± .003
0.873 ± .002
0.671 ± .022
Table 2: Ablation on training span, bank and retrieval. Rows add one ingredient at a time ( Sec. 4.3 ); spans are drawn uniformly from 10 s to the length given. PSNR and LPIPS need rendered ground truth, which web photographs lack. All runs at step 6,000 ; mean over five rollout seeds, standard deviation of the per-seed means in small type. DINO is Revisit with the cosine of DINO ViT-B/16 features in place of pixel correlation. Best per split in bold, second underlined. Seen scenes are in Table 7 ( App. C.2 ).
Figure 5: Ablation study on five models of Table 2 on a web photograph. Top: back near the input view. Middle: a place a little further along. Bottom: the return to it.
Figure 6: Revisit against return interval over 400 s rollouts on the web photographs; error bars are the standard deviation over five rollout seeds. (a) Revisit per model, grouped by the interval between the two visits; n counts returns per seed. (b) Difference from 10 s top- K on the same returns. Shading marks the first 60 s, which the tables score.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Translations rescaled by
Pose distance D ( w=4 )
Frustum overlap O
0.1×
57.1%
39.5%
1× (as posed)
55.9%
47.4%
10×
46.9%
53.0%
Appendix
Table 3: Sensitivity to pose scale. Achievable gain when translations are rescaled with each cue’s constants held fixed.
Per-chunk retrieval K
1
2
3
4
6
8
12
16
24
Achievable gain (%) ↑
59.6
60.0
58.9
57.4
55.4
54.2
51.1
49.1
45.9
Bank capacity
4
6
10
16
24
40
64
∞
Achievable gain (%) ↑
46.5
46.9
49.1
49.7
50.4
51.4
55.0
55.6
Appendix
Table 4: Retrieval width and bank capacity. Top: each width K against its own ceiling, the K most similar candidates. Bottom: six frames read from a bank that admits frames by coverage, dropping near-duplicates, and is capped at the given size, together with the latents already completed in the chunk’s window, which supply the rest when the cap is below six; ∞ lifts the cap but not the de-duplication, so it differs slightly from K=6 above.
Group
Setting
Value
Model and data
Initialization
Wan2.2-TI2V-5B ( Wan Team, 2025 )
Camera conditioning
PRoPE ( Li et al., 2025a )
Video
864×480 at 16 fps
Training sample
Latents per chunk c
4
Scored chunks k , recent chunks r
10 , 1
Span m (chunks)
U[10,mmax] , mmax=100 for Memorizon ( 41 – 401 latents)
Appendix
Table 5: Settings. Everything the runs of Table 2 share; the ablation rows differ only as Sec. 4.3 describes.
Definition
Position
Heading
Also required
Used for
Corpus return ( App. B.2 )
≤τ=0.5
—
camera away >ε=1.5 in between
could one window hold both visits
Evaluation return ( App. C.1 )
≤0.5
≤15∘
≥8 s apart; away >1.5 or 90∘
Revisit and Gain
Fold-back ( App. C.1 )
an evaluation return whose path back retraces the path out within 0.5
camera-following check
Appendix
Table 6: Three definitions of a return. Distances in the units of the corpus (one unit is 4 m).
Figure 7: Revisits in the training corpus. (a) Shortest return interval for three definitions. (b) Share of training windows of each length that contain a return. (c) Per scene, sorted by median, at τ=0.5 , ε=1.5 . Dashed lines mark spans of 10 , 100 , 200 and 400 s.
Fidelity to GT
Memory
VBench
Training
PSNR ↑
LPIPS ↓
Revisit ↑
Gain ↑
DINO ↑
Subj. ↑
Backg. ↑
Imag. ↑
Seen scenes, new trajectories
10 s, no retrieval
11.48 ± .18
0.610 ± .008
0.333 ± .017
0.074 ± .021
0.633 ± .013
0.801 ± .002
0.882 ± .003
0.648 ± .005
10 s, 10 -chunk window
11.08 ± .14
0.641 ± .008
0.321 ± .019
0.081 ± .017
0.599 ± .008
0.774 ± .004
0.875 ± .002
0.635 ± .005
10 s, top- K , no bank
11.41 ± .28
0.609 ± .008
0.427 ± .026
0.186 ± .026
0.744 ± .011
0.799 ± .006
0.883 ± .004
0.654 ± .008
100 s, per-segment retrieval, bank
11.92 ± .10
0.609 ± .006
0.454 ± .025
0.226 ± .031
0.761 ± .020
0.797 ± .005
0.879 ± .003
0.658 ± .008
Appendix
Table 7: Ablation of Table 2 on seen scenes , held-out episodes of the training scenes along new trajectories. Same runs, protocol and columns; best in bold, second underlined.
Starting pose
Mid-path
VBench
Method
Revisit ↑
Gain ↑
DINO ↑
D. Gain ↑
Revisit ↑
Gain ↑
DINO ↑
D. Gain ↑
Subj. ↑
Backg. ↑
Imag. ↑
LingBot-World 2.0
0.143 ± .026
− 0.006 ± .032
0.526 ± .013
− 0.001 ± .029
0.147 ± .017
0.011 ± .015
0.546 ± .020
0.017 ± .019
0.765 ± .004
0.852 ± .004
0.741 ± .004
DreamX-World 1.0
0.146 ± .009
0.033 ± .021
0.547 ± .016
− 0.045 ± .013
0.109 ± .006
0.009 ± .007
0.591 ± .013
0.007 ± .013
0.781 ± .001
0.868 ± .001
0.684 ± .006
HY-WorldPlay 1.5 †
0.243 ± .014
0.040 ± .034
0.679 ± .022
0.025 ± .031
0.194 ± .011
0.037 ± .010
0.674 ± .004
0.027 ± .004
0.839 ± .002
0.908 ± .001
0.702 ± .008
Matrix-Game 3.0
0.254 ± .018
0.075 ± .017
0.717 ± .013
0.042 ± .019
0.283 ± .012
0.089 ± .020
0.760 ± .016
0.100 ± .016
0.813 ± .004
0.895 ± .002
0.761 ± .003
Infinite-World
0.109 ± .022
0.019 ± .031
0.610 ± .026
0.013 ± .022
0.168 ± .017
0.070 ± .012
0.646 ± .013
0.052 ± .007
0.789 ± .005
0.856 ± .003
0.746 ± .006
Appendix
Table 8: Table 1 with its spreads. Mean over five rollout seeds, with the standard deviation of the per-seed means in small type.
Interval between the visits
Type
Training
8 – 10 s
10 – 20 s
20 – 60 s
Start
Mid
All
Seen scenes
10 s, no retrieval
0.230
0.064
0.073
0.160
0.067
0.078
10 s, top- K
0.397
0.174
0.185
0.130
0.199
0.191
100 s, top- K , bank
0.337
0.216
0.253
0.213
0.252
0.248
Unseen scenes
Appendix
Table 9: Gain by return interval and type. Pixel Gain with the five rollout seeds pooled; seen and unseen scenes as in Table 2 , web photographs as in Table 1 . Start : the earlier frame lies in the first second; mid : every other return. Pooling pairs rather than clips makes all differ slightly from Table 2 . For the 100 s run the three interval bins hold 25 , 140 and 470 returns on seen scenes, 30 , 105 and 230 on unseen ones and 20 , 100 and 605 on web photographs.
Training
Sequence slots
Step time (s)
Wall clock (h)
Peak mem. (GB)
10 s, no retrieval
41 / 41
3.49
5.8
16.05
10 s, 10 -chunk window
41 / 41
4.02
6.7
16.05
10 s, top- K , no bank
41 / 41
3.68
6.1
16.05
100 s, per-segment retrieval, bank
51 / 51
4.17
6.9
17.99
100 s, top- K , ordered bank
65 / 105
8.18
13.6
28.45
100 s, top- K , bank
65 / 105
8.02
13.4
28.45
Appendix
Table 10: Training cost of the runs in Table 2 , from their logs. Slots: median and maximum over logged steps. Wall clock is the step time over the 6,000 steps every run trains for.
System
Correlation ↑
Direction ↑
Implied hfov †
Calibration
Rendered frames, engine poses
0.906
0.974
89.1∘
Open world models
LingBot-World 2.0
0.855
0.966
73.3∘
DreamX-World 1.0
0.825
0.955
132.5∘
HY-WorldPlay 1.5
0.697
0.909
123.6∘
Appendix
Table 11: Camera following : yaw only, web photographs, first 60 s, rollout seed 0 . † Field of view implied by image shift per degree of prescribed yaw; the paths assume 90∘ . Memorizon : mean over six runs of Table 2 , the three trained on 10 s clips and the 100 , 200 and 400 s runs.
Starting pose
Mid-path
Bank contents
Revisit ↑
Gain ↑
DINO ↑
D. Gain ↑
Revisit ↑
Gain ↑
DINO ↑
D. Gain ↑
Top- K (ours)
0.377
0.213
0.794
0.218
0.440
0.319
0.826
0.256
Empty
0.176
0.053
0.553
0.085
0.105
0.019
0.504
0.035
K random frames
0.153
0.055
0.583
0.044
0.207
0.109
0.634
0.087
K frames, another episode
0.071
0.012
0.445
− 0.020
0.073
0.016
0.459
− 0.014
Appendix
Table 12: Bank contents at inference. The 100 s run with only the bank’s frames replaced: its top- K (ours), nothing, K random frames from the same history, or K frames from another episode. Web photographs, first 60 s, returns split as in Table 1 , DINO columns as there; rollout seed 0 .
Figure 8: More returns from open world models , extending Figure 4 to three more web photographs. Each block follows one photograph’s walk. Blue tags mark returns to the starting pose; orange tags the two ends of a return to a pose met mid-path. Rollout seed 0 .
Imaging ↑
Subject ↑
Training
Generated
GT bank
GT context
Generated
GT bank
GT context
Seen scenes
10 s, no retrieval
0.648 ± .005
–
0.671 ± .001
0.801 ± .002
–
0.820 ± .000
10 s, top- K , no bank in training
0.654 ± .008
0.663 ± .004
0.673 ± .000
0.799 ± .006
0.811 ± .002
0.824 ± .000
100 s, top- K , bank
0.646 ± .015
0.666 ± .004
0.668 ± .000
0.804 ± .001
0.818 ± .002
0.820 ± .000
200 s, top- K , bank
0.637 ± .012
0.656 ± .003
0.666 ± .000
0.806 ± .006
0.820 ± .001
0.821 ± .000
Appendix
Table 13: Quality against what fills the context on seen and unseen scenes, first 60 s, five seeds. Generated : ordinary rollout. GT bank : the bank holds ground-truth latents. GT context : bank and recent window both ground truth. The ground-truth videos themselves score 0.678 and 0.803 on seen scenes and 0.719 and 0.774 on unseen ones. With a clean history each chunk generates only four latents, so the seed spread rounds to zero.
Figure 9: More returns from Memorizon on six web photographs. Blue: a pose close to the first frame’s. Orange and green: two returns, each shown at its first visit and on coming back; they meet the Revisit rule of Sec. 4.1 , or, where a path has only one such return, the same rule at twice its tolerance.
Institute of Foundation Models, Mohamed bin Zayed University of Artificial Intelligence · Mohamed bin Zayed University of Artificial Intelligence · Pinscreen