cs.CVOct 5, 2026

HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models

Authors: Zhuokun Chen, Feng Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang

Organizations: Monash University · The University of Adelaide · Zhejiang University

Abstract

Long-horizon video world models require persistent memory to preserve scene consistency over extended rollouts. Softmax attention retains the full generation history through a growing KV cache, whereas recurrent linear attention compresses history into fixed-size states with substantially lower memory cost. However, we identify severe long-range forgetting in Gated DeltaNet (GDN), where information from distant but relevant scenes is progressively attenuated by subsequent state updates. To address this limitation, we propose HLA-WM, a training-free hybrid linear-attention framework that combines coarse-grained geometry-guided retrieval with fine-grained recurrent linear-state computation. HLA-WM exploits the affine structure of GDN to cache compact chunk-wise transition summaries, retrieve scene-relevant historical chunks using camera geometry, and recompose them into query-specific recurrent states. On the 6060-second SANA-WM-Bench, HLA-WM improves all six aggregate revisit-consistency and camera-control metrics of the base autoregressive generator without additional training, including a 0.740.74 dB PSNR gain and a 28.5%28.5\% reduction in rotation error. The improvements persist after downstream refinement and generalize to MBench-A, where HLA-WM consistently improves all three revisit-consistency metrics across all four subsets and all evaluated inference modes over 547547 samples. At a 6060-second context, HLA-WM reduces historical-state memory by 12×12\times relative to full KV caching while incurring at most a 1.6%1.6\% reduction in inference throughput. These results demonstrate that selectively addressable recurrent memory can improve long-range scene recall while preserving the efficiency advantages of GDN. Project page: https://caesarhhh.github.io/hla-wm/

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LOCI: Spatial Linear Memory for Streaming World Models

    Sep 30, 2026Ji Xia, Tingting Liao, Xuezhi Liang +2Spatial MemoryRecurrent Model

  2. Compression and Retrieval: Implicit Memory Retrieval for Video World Models

    Jun 22, 2026Zhan Peng, Jie Ma, Huiqiang Sun +6Video World ModelsLong-Term Memory

  3. Geometry-Aware Implicit Memory for Video World Models

    Jun 1, 2026Zhengxuan Wei, Xu Guo, Xinghui Li +8Video World ModelsGeometry-Aware