Long-horizon video world models require persistent memory to preserve scene consistency over extended rollouts. Softmax attention retains the full generation history through a growing KV cache, whereas recurrent linear attention compresses history into fixed-size states with substantially lower memory cost. However, we identify severe long-range forgetting in Gated DeltaNet (GDN), where information from distant but relevant scenes is progressively attenuated by subsequent state updates. To address this limitation, we propose HLA-WM, a training-free hybrid linear-attention framework that combines coarse-grained geometry-guided retrieval with fine-grained recurrent linear-state computation. HLA-WM exploits the affine structure of GDN to cache compact chunk-wise transition summaries, retrieve scene-relevant historical chunks using camera geometry, and recompose them into query-specific recurrent states. On the 60-second SANA-WM-Bench, HLA-WM improves all six aggregate revisit-consistency and camera-control metrics of the base autoregressive generator without additional training, including a 0.74 dB PSNR gain and a 28.5% reduction in rotation error. The improvements persist after downstream refinement and generalize to MBench-A, where HLA-WM consistently improves all three revisit-consistency metrics across all four subsets and all evaluated inference modes over 547 samples. At a 60-second context, HLA-WM reduces historical-state memory by 12× relative to full KV caching while incurring at most a 1.6% reduction in inference throughput. These results demonstrate that selectively addressable recurrent memory can improve long-range scene recall while preserving the efficiency advantages of GDN. Project page: https://caesarhhh.github.io/hla-wm/
Figures & tables
Figure 1 : Long-range forgetting in GDN-based world-model rollouts. Upper left: Revisit SSIM over increasing temporal gaps. Lower left: Camera displacement along the qualitative trajectory. Right: Qualitative comparison of the first visit, far excursion, and revisit at the similar camera pose. After the long excursion, the baseline SANA-WM shows noticeable changes in scene layout and object arrangement, whereas our HLA-WM better preserves the first-visit scene content.
Figure 2 : Long-range forgetting in GDN. The camera follows a long loop and eventually returns near the region observed at Chunk 3 around Chunk 35 . We visualize representative generated chunks along the trajectory together with the cumulative retention W3→j of Chunk 3 . Although the early scene becomes relevant again near the end of the rollout, its retained influence continuously decreases as intermediate chunks are generated, reaching only 0.0416 at Chunk 35 .
Figure 3 : Overview of HLA-WM. Historical video chunks are represented by independent GDN summaries. Given the current query, geometry-guided retrieval selects scene-relevant historical chunks together with persistent sink and recent context. The selected summaries are then composed in chronological order to construct a query-specific recurrent state.
Pipeline
Method
Revisit Consistency
Camera Control
Efficiency
PSNR ↑
SSIM ↑
LPIPS ↓
RotErr ( ∘ ) ↓
TransErr ↓
CamMC ↓
FPS ↑
Simple-Trajectory Split
Stage-1
SANA-WM
9.18
0.1729
0.6327
19.7040
2.2625
2.3974
22.403
HLA-WM
9.91
0.1999
0.6049
13.5864
2.0928
2.1798
22.053
Stage-1 + AR Refine
SANA-WM
14.44
0.2770
0.5738
20.9315
2.1587
2.3121
8.280
HLA-WM
14.82
0.2872
0.5682
20.2933
2.1910
2.3240
8.211
Table 1 : Main results on SANA-WM-Bench under different generation stages and refinement modes. We report revisit consistency, camera controllability, and inference throughput. Best results within each setting are shown in bold.
Figure 5
Figure 5 : Qualitative comparison on long-range revisit scenarios under causal AR refinement (top) and bidirectional refinement (bottom). Each example shows the first visit, intermediate views, and a later revisit along the same camera trajectory. Revisit metrics are computed between the first-visit and revisit frames within each method.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Variant
Revisit Consistency
Camera Control
PSNR ↑
SSIM ↑
LPIPS ↓
RotErr ( ∘ ) ↓
TransErr ↓
CamMC ↓
HLA-WM
10.11
0.1997
0.5848
15.0887
1.8388
1.9482
w/o Sink
9.71
0.1867
0.6019
15.9351
1.9326
2.0531
w/o Select A
9.61
0.1887
0.5901
15.2426
1.8552
2.0562
w/o Select B
9.64
0.1892
0.5921
15.2326
1.9830
2.0564
w/o Turn-aware
9.96
0.1917
0.5876
15.1784
2.1096
1.9899
Appendix
Table 3 : Component ablation on the Hard-Trajectory split using Stage 1 outputs. Best results are shown in bold.
Pipeline
Method
PSNR ↑
SSIM ↑
LPIPS ↓
RotErr ( ∘ ) ↓
TransErr ↓
CamMC ↓
Simple-Trajectory Split
Stage 1
Baseline
9.18
0.1729
0.6327
19.7040
2.2625
2.3974
GDN-only
9.45
0.1812
0.6219
17.9934
2.2832
2.4083
KV-only
9.50
0.1866
0.6149
14.4283
1.9569
2.0556
HLA-WM
9.91
0.1999
0.6049
13.5864
2.0928
2.1798
Stage 1 + AR Refine
Baseline
14.44
0.2770
0.5738
20.9315
2.1587
2.3121
Appendix
Table 4: Ablation of GDN-state recomposition and KV-cache alignment on SANA-WM-Bench under Top- 1 retrieval. Best results within each split and inference setting are shown in bold.
Pipeline
Method
PSNR ↑
SSIM ↑
LPIPS ↓
Stage 1
Baseline
10.17
0.2545
0.6251
GDN-only
10.61
0.2626
0.6052
KV-only
10.43
0.2606
0.6091
HLA-WM
11.01
0.2848
0.5823
Causal AR Refine
Baseline
12.36
0.2996
0.6554
GDN-only
12.57
0.3064
0.6492
Appendix
Table 5: Ablation of GDN-state recomposition and KV-cache alignment on MBench-A under Top- 1 retrieval. Best results within each inference setting are shown in bold.
K
Revisit Consistency
Camera Control
PSNR ↑
SSIM ↑
LPIPS ↓
RotErr ( ∘ ) ↓
TransErr ↓
CamMC ↓
1
10.11
0.1997
0.5848
15.0887
1.8388
1.9482
3
10.06
0.1973
0.5865
14.1537
1.7658
1.8661
4
10.00
0.1936
0.5885
12.2795
1.6985
1.7832
5
9.89
0.1886
0.5915
11.5435
1.7200
1.7970
Appendix
Table 6 : Ablation on the retrieval budget K on the Hard-Trajectory split using Stage 1 outputs. Best results are shown in bold.
Selection Policy
PSNR ↑
SSIM ↑
LPIPS ↓
RotErr ( ∘ ) ↓
TransErr ↓
CamMC ↓
Random
9.70
0.1864
0.6049
14.1883
1.7822
1.8924
Recent
9.29
0.1696
0.6179
27.6876
2.0142
2.2408
Uniform
9.91
0.1910
0.5957
12.3324
1.6881
1.7838
HLA-WM
10.00
0.1936
0.5885
12.2795
1.6985
1.7832
Appendix
Table 7 : Comparison of historical state selection strategies on Hard80 under Stage 1 inference with K=4 . Higher PSNR and SSIM are better, while lower LPIPS and camera errors are better.
Figure 6 : Performance improvement over SANA-WM across different revisit intervals on the Hard-Trajectory split. We report PSNR improvement (left), SSIM improvement (middle), and LPIPS reduction (right) for Stage 1, AR refinement, and bidirectional refinement.
Category
Method
PSNR ↑
SSIM ↑
LPIPS ↓
RotErr ( ∘ ) ↓
TransErr ↓
CamMC ↓
Simple-Trajectory Split
Game Style
SANA-WM
9.85
0.2070
0.6197
16.5064
1.9242
2.0482
HLA-WM
11.06
0.2639
0.5866
12.8640
2.0047
2.0939
Indoor
SANA-WM
9.38
0.2105
0.6576
38.6940
1.6993
2.0227
HLA-WM
10.05
0.2338
0.6306
29.0323
1.6196
1.8481
Outdoor City
SANA-WM
8.96
0.1323
0.6390
12.2424
2.0782
2.1372
Appendix
Table 8 : Category-level results on SANA-WM-Bench using Stage 1 outputs. Best results within each category are shown in bold.
Category
Method
PSNR ↑
SSIM ↑
LPIPS ↓
RotErr ( ∘ ) ↓
TransErr ↓
CamMC ↓
Simple-Trajectory Split
Game Style
SANA-WM
15.19
0.3225
0.5582
19.8894
2.1729
2.3106
HLA-WM
15.79
0.3505
0.5492
16.5280
2.0506
2.1496
Indoor
SANA-WM
13.64
0.3127
0.6013
35.3855
1.5308
1.8652
HLA-WM
14.08
0.3172
0.5956
34.3931
1.6263
1.9040
Outdoor City
SANA-WM
14.21
0.2353
0.5803
15.1111
1.9558
2.0376
Appendix
Table 9 : Category-level results on SANA-WM-Bench after causal AR refinement. Best results within each category are shown in bold.
Category
Method
PSNR ↑
SSIM ↑
LPIPS ↓
RotErr ( ∘ ) ↓
TransErr ↓
CamMC ↓
Simple-Trajectory Split
Game Style
SANA-WM
13.74
0.3351
0.5485
9.7648
1.7510
1.8012
HLA-WM
14.20
0.3553
0.5377
6.8672
1.8842
1.9156
Indoor
SANA-WM
12.67
0.3174
0.6045
21.6798
1.4978
1.6675
HLA-WM
13.03
0.3401
0.5980
16.1626
1.0450
1.1800
Outdoor City
SANA-WM
12.79
0.2479
0.5767
7.1726
1.7286
1.7495
Appendix
Table 10 : Category-level results on SANA-WM-Bench after bidirectional refinement. Best results within each category are shown in bold.
Subset
n
Method
PSNR ↑
SSIM ↑
LPIPS ↓
Stage 1
Causal
100
SANA-WM
10.4068
0.28982
0.65461
HLA-WM
11.0022
0.31830
0.61659
Human
120
SANA-WM
9.5147
0.26845
0.67977
HLA-WM
10.2161
0.30275
0.63759
Environment
227
SANA-WM
10.4140
0.24056
0.59436
Appendix
Table 11 : Detailed Top- 1 results on the four official MBench-A subsets under different inference modes. Best results within each subset and mode are shown in bold.
Figure 7 : Additional qualitative comparisons for Stage 1 outputs. Each example shows the first visit, two intermediate views, and a later revisit along the same camera trajectory. Revisit consistency is evaluated between the first-visit and revisit frames within each method.
Figure 8 : Additional qualitative comparisons under causal AR refinement. Each example shows the first visit, two intermediate views, and a later revisit along the same camera trajectory. Revisit consistency is evaluated between the first-visit and revisit frames within each method.
Figure 9 : Additional qualitative comparisons under bidirectional refinement. Each example shows the first visit, two intermediate views, and a later revisit along the same camera trajectory. Revisit consistency is evaluated between the first-visit and revisit frames within each method.
Institute of Foundation Models, Mohamed bin Zayed University of Artificial Intelligence · Mohamed bin Zayed University of Artificial Intelligence · Pinscreen