Long-horizon video world models require persistent memory to preserve scene consistency over extended rollouts. Softmax attention retains the full generation history through a growing KV cache, whereas recurrent linear attention compresses history into fixed-size states with substantially lower memory cost. However, we identify severe long-range forgetting in Gated DeltaNet (GDN), where information from distant but relevant scenes is progressively attenuated by subsequent state updates. To address this limitation, we propose HLA-WM, a training-free hybrid linear-attention framework that combines coarse-grained geometry-guided retrieval with fine-grained recurrent linear-state computation. HLA-WM exploits the affine structure of GDN to cache compact chunk-wise transition summaries, retrieve scene-relevant historical chunks using camera geometry, and recompose them into query-specific recurrent states. On the 60-second SANA-WM-Bench, HLA-WM improves all six aggregate revisit-consistency and camera-control metrics of the base autoregressive generator without additional training, including a 0.74 dB PSNR gain and a 28.5% reduction in rotation error. The improvements persist after downstream refinement and generalize to MBench-A, where HLA-WM consistently improves all three revisit-consistency metrics across all four subsets and all evaluated inference modes over 547 samples. At a 60-second context, HLA-WM reduces historical-state memory by 12× relative to full KV caching while incurring at most a 1.6% reduction in inference throughput. These results demonstrate that selectively addressable recurrent memory can improve long-range scene recall while preserving the efficiency advantages of GDN. Project page: https://caesarhhh.github.io/hla-wm/
Figures & tables
Figure 1 : Long-range forgetting in GDN-based world-model rollouts. Upper left: Revisit SSIM over increasing temporal gaps. Lower left: Camera displacement along the qualitative trajectory. Right: Qualitative comparison of the first visit, far excursion, and revisit at the similar camera pose. After the long excursion, the baseline SANA-WM shows noticeable changes in scene layout and object arrangement, whereas our HLA-WM better preserves the first-visit scene content.
Figure 2 : Long-range forgetting in GDN. The camera follows a long loop and eventually returns near the region observed at Chunk 3 around Chunk 35 . We visualize representative generated chunks along the trajectory together with the cumulative retention W3→j of Chunk 3 . Although the early scene becomes relevant again near the end of the rollout, its retained influence continuously decreases as intermediate chunks are generated, reaching only 0.0416 at Chunk 35 .
Figure 3 : Overview of HLA-WM. Historical video chunks are represented by independent GDN summaries. Given the current query, geometry-guided retrieval selects scene-relevant historical chunks together with persistent sink and recent context. The selected summaries are then composed in chronological order to construct a query-specific recurrent state.
Pipeline
Method
Revisit Consistency
Camera Control
Efficiency
PSNR ↑
SSIM ↑
LPIPS ↓
RotErr ( ∘ ) ↓
TransErr ↓
CamMC ↓
FPS ↑
Simple-Trajectory Split
Stage-1
SANA-WM
9.18
0.1729
0.6327
19.7040
2.2625
2.3974
22.403
HLA-WM
9.91
0.1999
0.6049
13.5864
2.0928
2.1798
22.053
Stage-1 + AR Refine
SANA-WM
14.44
0.2770
0.5738
20.9315
2.1587
2.3121
8.280
HLA-WM
14.82
0.2872
0.5682
20.2933
2.1910
2.3240
8.211
Table 1 : Main results on SANA-WM-Bench under different generation stages and refinement modes. We report revisit consistency, camera controllability, and inference throughput. Best results within each setting are shown in bold.
Figure 5
Figure 5 : Qualitative comparison on long-range revisit scenarios under causal AR refinement (top) and bidirectional refinement (bottom). Each example shows the first visit, intermediate views, and a later revisit along the same camera trajectory. Revisit metrics are computed between the first-visit and revisit frames within each method.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Variant
Revisit Consistency
Camera Control
PSNR ↑
SSIM ↑
LPIPS ↓
RotErr ( ∘ ) ↓
TransErr ↓
CamMC ↓
HLA-WM
10.11
0.1997
0.5848
15.0887
1.8388
1.9482
w/o Sink
9.71
0.1867
0.6019
15.9351
1.9326
2.0531
w/o Select A
9.61
0.1887
0.5901
15.2426
1.8552
2.0562
w/o Select B
9.64
0.1892
0.5921
15.2326
1.9830
2.0564
w/o Turn-aware
9.96
0.1917
0.5876
15.1784
2.1096
1.9899
Appendix
Table 3 : Component ablation on the Hard-Trajectory split using Stage 1 outputs. Best results are shown in bold.
Pipeline
Method
PSNR ↑
SSIM ↑
LPIPS ↓
RotErr ( ∘ ) ↓
TransErr ↓
CamMC ↓
Simple-Trajectory Split
Stage 1
Baseline
9.18
0.1729
0.6327
19.7040
2.2625
2.3974
GDN-only
9.45
0.1812
0.6219
17.9934
2.2832
2.4083
KV-only
9.50
0.1866
0.6149
14.4283
1.9569
2.0556
HLA-WM
9.91
0.1999
0.6049
13.5864
2.0928
2.1798
Stage 1 + AR Refine
Baseline
14.44
0.2770
0.5738
20.9315
2.1587
2.3121
Appendix
Table 4: Ablation of GDN-state recomposition and KV-cache alignment on SANA-WM-Bench under Top- 1 retrieval. Best results within each split and inference setting are shown in bold.
Pipeline
Method
PSNR ↑
SSIM ↑
LPIPS ↓
Stage 1
Baseline
10.17
0.2545
0.6251
GDN-only
10.61
0.2626
0.6052
KV-only
10.43
0.2606
0.6091
HLA-WM
11.01
0.2848
0.5823
Causal AR Refine
Baseline
12.36
0.2996
0.6554
GDN-only
12.57
0.3064
0.6492
Appendix
Table 5: Ablation of GDN-state recomposition and KV-cache alignment on MBench-A under Top- 1 retrieval. Best results within each inference setting are shown in bold.
K
Revisit Consistency
Camera Control
PSNR ↑
SSIM ↑
LPIPS ↓
RotErr ( ∘ ) ↓
TransErr ↓
CamMC ↓
1
10.11
0.1997
0.5848
15.0887
1.8388
1.9482
3
10.06
0.1973
0.5865
14.1537
1.7658
1.8661
4
10.00
0.1936
0.5885
12.2795
1.6985
1.7832
5
9.89
0.1886
0.5915
11.5435
1.7200
1.7970
Appendix
Table 6 : Ablation on the retrieval budget K on the Hard-Trajectory split using Stage 1 outputs. Best results are shown in bold.
Selection Policy
PSNR ↑
SSIM ↑
LPIPS ↓
RotErr ( ∘ ) ↓
TransErr ↓
CamMC ↓
Random
9.70
0.1864
0.6049
14.1883
1.7822
1.8924
Recent
9.29
0.1696
0.6179
27.6876
2.0142
2.2408
Uniform
9.91
0.1910
0.5957
12.3324
1.6881
1.7838
HLA-WM
10.00
0.1936
0.5885
12.2795
1.6985
1.7832
Appendix
Table 7 : Comparison of historical state selection strategies on Hard80 under Stage 1 inference with K=4 . Higher PSNR and SSIM are better, while lower LPIPS and camera errors are better.
Figure 6 : Performance improvement over SANA-WM across different revisit intervals on the Hard-Trajectory split. We report PSNR improvement (left), SSIM improvement (middle), and LPIPS reduction (right) for Stage 1, AR refinement, and bidirectional refinement.
Category
Method
PSNR ↑
SSIM ↑
LPIPS ↓
RotErr ( ∘ ) ↓
TransErr ↓
CamMC ↓
Simple-Trajectory Split
Game Style
SANA-WM
9.85
0.2070
0.6197
16.5064
1.9242
2.0482
HLA-WM
11.06
0.2639
0.5866
12.8640
2.0047
2.0939
Indoor
SANA-WM
9.38
0.2105
0.6576
38.6940
1.6993
2.0227
HLA-WM
10.05
0.2338
0.6306
29.0323
1.6196
1.8481
Outdoor City
SANA-WM
8.96
0.1323
0.6390
12.2424
2.0782
2.1372
Appendix
Table 8 : Category-level results on SANA-WM-Bench using Stage 1 outputs. Best results within each category are shown in bold.
Category
Method
PSNR ↑
SSIM ↑
LPIPS ↓
RotErr ( ∘ ) ↓
TransErr ↓
CamMC ↓
Simple-Trajectory Split
Game Style
SANA-WM
15.19
0.3225
0.5582
19.8894
2.1729
2.3106
HLA-WM
15.79
0.3505
0.5492
16.5280
2.0506
2.1496
Indoor
SANA-WM
13.64
0.3127
0.6013
35.3855
1.5308
1.8652
HLA-WM
14.08
0.3172
0.5956
34.3931
1.6263
1.9040
Outdoor City
SANA-WM
14.21
0.2353
0.5803
15.1111
1.9558
2.0376
Appendix
Table 9 : Category-level results on SANA-WM-Bench after causal AR refinement. Best results within each category are shown in bold.
Category
Method
PSNR ↑
SSIM ↑
LPIPS ↓
RotErr ( ∘ ) ↓
TransErr ↓
CamMC ↓
Simple-Trajectory Split
Game Style
SANA-WM
13.74
0.3351
0.5485
9.7648
1.7510
1.8012
HLA-WM
14.20
0.3553
0.5377
6.8672
1.8842
1.9156
Indoor
SANA-WM
12.67
0.3174
0.6045
21.6798
1.4978
1.6675
HLA-WM
13.03
0.3401
0.5980
16.1626
1.0450
1.1800
Outdoor City
SANA-WM
12.79
0.2479
0.5767
7.1726
1.7286
1.7495
Appendix
Table 10 : Category-level results on SANA-WM-Bench after bidirectional refinement. Best results within each category are shown in bold.
Subset
n
Method
PSNR ↑
SSIM ↑
LPIPS ↓
Stage 1
Causal
100
SANA-WM
10.4068
0.28982
0.65461
HLA-WM
11.0022
0.31830
0.61659
Human
120
SANA-WM
9.5147
0.26845
0.67977
HLA-WM
10.2161
0.30275
0.63759
Environment
227
SANA-WM
10.4140
0.24056
0.59436
Appendix
Table 11 : Detailed Top- 1 results on the four official MBench-A subsets under different inference modes. Best results within each subset and mode are shown in bold.
Figure 7 : Additional qualitative comparisons for Stage 1 outputs. Each example shows the first visit, two intermediate views, and a later revisit along the same camera trajectory. Revisit consistency is evaluated between the first-visit and revisit frames within each method.
Figure 8 : Additional qualitative comparisons under causal AR refinement. Each example shows the first visit, two intermediate views, and a later revisit along the same camera trajectory. Revisit consistency is evaluated between the first-visit and revisit frames within each method.
Figure 9 : Additional qualitative comparisons under bidirectional refinement. Each example shows the first visit, two intermediate views, and a later revisit along the same camera trajectory. Revisit consistency is evaluated between the first-visit and revisit frames within each method.
When a camera revisits a previously observed region, a video world model should reproduce what was there before. This requires both remembering past observations and retrieving the right one for the current viewpoint. Key-value caches preserve visual detail but grow with video length; recurrent memory is compact but compresses history into a fixed-size state, so individual past observations are no longer directly accessible. We introduce LOCI, a hybrid spatial-memory architecture that keeps both representations. In half of the transformer blocks, main attention keeps a key-value cache of past observations; in the other half, it is restricted to the current chunk and complemented by a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry, so viewpoint enters both memory addressing and stored content. Recurrent readouts flow into subsequent cache-backed blocks and supply their queries with accumulated scene context. On the public MIND memory benchmark and on held-out recorded trajectories, LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model; with full history, it lowers peak memory at equal length by about 30% relative to full softmax. With a bounded bank of retained observations, it streams long videos at constant memory and remains more faithful than full softmax under the same budget.
Ji Xia, Tingting Liao, Xuezhi Liang +2
Institute of Foundation Models, Mohamed bin Zayed University of Artificial Intelligence · Mohamed bin Zayed University of Artificial Intelligence · Pinscreen
Video world models hold promise for simulating interactive environments, yet maintaining consistent long-term memory across complex camera trajectories remains a critical challenge. Existing methods typically rely on computationally expensive context scaling or rigid heuristic retrieval mechanisms, which lacks generalization to varying camera trajectories and environments. In this paper, we propose Compression and Retrieval (CaR), an attention-driven implicit memory retrieval mechanism to overcome these limitations. By injecting viewpoint information via positional encoding, our method performs flexible memory retrieval through attention computation. To efficiently process extended contexts with minimal computational overhead, we further introduce a lightweight context compression network. Furthermore, we construct SceneFly, a large-scale synthetic dataset featuring realistic camera trajectories and frame-level annotations to train and evaluate long-horizon video world models. Extensive experiments demonstrate that our approach achieves state-of-the-art results on established benchmarks and exhibits strong generalization to open-domain scenes.
Zhan Peng, Jie Ma, Huiqiang Sun +6
Huazhong University of Science and Technology, China · HUJING Digital Media & Entertainment Group, China · Sun Yat-sen University, China
Video world models aim to simulate controllable visual environments, but long-horizon rollouts depend on what the model remembers after observations leave its native context window. Explicit memories retain frames or online 3D reconstructions, which can suffer from heuristic retrieval errors, redundant appearance storage, or reconstruction artifacts. Implicit memories compress history into a compact state, but existing designs are not explicitly constrained to encode cross-view scene geometry. We propose GIM-World, a geometry-aware implicit memory framework for video world models. A lightweight transformer encoder compresses variable-length history into fixed-size memory tokens, a camera-queryable geometry head distills 3D scene structure from a frozen foundation model into the memory during training, and an information-guided pruning rule keeps encoding cost bounded as history grows. The geometry teacher is discarded at inference, leaving a lightweight memory module. Experiments on MIND show that GIM-World better preserves long-horizon geometric and visual consistency than both explicit- and implicit-memory baselines.
Zhengxuan Wei, Xu Guo, Xinghui Li +8
School of Intelligence Science and Technology, Nanjing University · 2Kling Team, Kuaishou Technology · 3Tsinghua University