LOCI: Spatial Linear Memory for Streaming World Models
Organizations: Institute of Foundation Models, Mohamed bin Zayed University of Artificial Intelligence · Mohamed bin Zayed University of Artificial Intelligence · Pinscreen
Abstract
When a camera revisits a previously observed region, a video world model should reproduce what was there before. This requires both remembering past observations and retrieving the right one for the current viewpoint. Key-value caches preserve visual detail but grow with video length; recurrent memory is compact but compresses history into a fixed-size state, so individual past observations are no longer directly accessible. We introduce LOCI, a hybrid spatial-memory architecture that keeps both representations. In half of the transformer blocks, main attention keeps a key-value cache of past observations; in the other half, it is restricted to the current chunk and complemented by a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry, so viewpoint enters both memory addressing and stored content. Recurrent readouts flow into subsequent cache-backed blocks and supply their queries with accumulated scene context. On the public MIND memory benchmark and on held-out recorded trajectories, LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model; with full history, it lowers peak memory at equal length by about 30% relative to full softmax. With a bounded bank of retained observations, it streams long videos at constant memory and remains more faithful than full softmax under the same budget.
Figures & tables
| Method | Params | Trained on MIND | MSE | PSNR | SSIM | LPIPS | |
| Reference (reported in the GIM-World paper) | |||||||
| SSM ( Po et al., 2025 ) | – | ✓ | – | 0.0796 | 11.96 | 0.395 | 0.744 |
| FramePack ( Zhang et al., 2025 ) | – | ✓ | – | 0.0764 | 12.04 | 0.376 | 0.706 |
| Context-as-Memory ( Yu et al., 2025 ) | – | ✓ | – | 0.0706 | 12.50 | 0.392 | 0.695 |
| GIM-World ( Wei et al., 2026 ) | 1.3B | ✓ | – | 0.0614 | 13.40 | 0.414 | 0.630 |
| Run by us (entire prediction segment) | |||||||
| PSNR | LPIPS | |||||||
|---|---|---|---|---|---|---|---|---|
| Data mixture | History | Softmax | Loci | [95% CI] | Softmax | Loci | [95% CI] | |
| Uniform | bounded | 50 | 12.17 | 12.85 | [ , ] | 0.700 | 0.680 | [ , ] |
| UE-weighted | bounded | 50 | 12.13 | 13.02 | [ , ] | 0.704 | 0.679 | [ , ] |
| UE-weighted | full | 32 | 13.90 | 14.25 | [ , ] | 0.669 | 0.641 | [ , ] |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Mode | Recurrent state | Main KV | Camera KV |
|---|---|---|---|
| Full softmax, dense | None | ||
| Loci , dense | 15 fixed layer states | ||
| Full softmax, sparse | None | ||
| Loci , sparse | 15 fixed layer states |
| Method | Rot. ( ∘ ) | Trans. ( ∘ ) | Mag. 1 |
|---|---|---|---|
| CaR ( Peng et al., 2026 ) | 2.14 | 8.0 | 0.74 |
| HY-WorldPlay ( Sun et al., 2026 ) | 2.75 | 9.6 | 1.38 |
| Matrix-Game 3.0 ( Wang et al., 2026 ) | 1.33 | 6.8 | 2.09 |
| LingBot-World ( Robbyant Team et al., 2026 ) | 0.72 | 4.9 | 1.06 |
| AlayaWorld ( AlayaWorld Team et al., 2026 ) | 7.49 | — | — |
| SANA-WM ( Zhu et al., 2026 ) | 1.71 | 4.5 | 3.02 |
| Comparison | Quantity | Value [95% CI] | Count |
|---|---|---|---|
| Loci vs. full softmax, full history | co-visible attention share | 11.1% vs. 10.45% | |
| enrichment | [ , ] | 8 / 8 clips | |
| Loci vs. full softmax, bounded | co-visible attention share | 9.66% vs. 8.91% | |
| enrichment | [ , ] | 8 / 8 clips | |
| Readout on off ( Loci ) | enrichment, downstream softmax layers | [ , ] | |
| enrichment, right after a hybrid block | [ , ] |
| Method | Params | Steps | Revisit (A+B) | Whole clip (A) | Identity | Cost | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| PSNR | LPIPS | PSNR | SSIM | LPIPS | DINO | s/v-s | GiB | |||
| CaR ( Peng et al., 2026 ) | 5B | 50 | 9.73 | 0.627 | 8.89 | 0.360 | 0.647 | 0.679 | 31.4 | 17.7 |
| HY-WorldPlay ( Sun et al., 2026 ) | 8B | 4 | 10.01 | 0.640 | 9.70 | 0.408 | 0.625 | 0.519 | 14.7 | 71.1 |
| Matrix-Game 3.0 ( Wang et al., 2026 ) | 5B | 3 | 8.40 | 0.687 | 7.85 | 0.313 | 0.678 | 0.626 | 6.2 | 35.0 |
| LingBot-World ( Robbyant Team et al., 2026 ) | 28B | 4 | 8.70 | 0.605 | 9.67 | 0.442 | 0.564 | 0.517 | 20.5 | 88.3 |
| AlayaWorld ( AlayaWorld Team et al., 2026 ) | 15B | 4 | 8.52 | 0.646 | 8.34 | 0.386 | 0.640 | – | 24.9 | 137.6 |
| Versus | PSNR | LPIPS | Ours lower | ||
|---|---|---|---|---|---|
| Mean | 95% CI | Mean | 95% CI | ||
| Full softmax (same recipe) | +0.62 | [+0.13, +1.12] | +0.018 | [ 0.006, +0.042] | 14 / 20 |
| HY-WorldPlay ( Sun et al., 2026 ) | +0.63 | [ 0.49, +1.71] | +0.079 | [+0.037, +0.122] | 14 / 20 |
| CaR ( Peng et al., 2026 ) | +0.91 | [ 0.10, +2.16] | +0.066 | [+0.007, +0.128] | 13 / 20 |
| SANA-WM ( Zhu et al., 2026 ) | +1.02 | [+0.03, +2.04] | +0.210 | [+0.146, +0.276] | 19 / 20 |
| LingBot-World ( Robbyant Team et al., 2026 ) | +1.95 | [+0.96, +2.99] | +0.044 | [ 0.027, +0.112] | 12 / 20 |
| Method | MSE | PSNR | SSIM | LPIPS | |
|---|---|---|---|---|---|
| HY-WorldPlay ( Sun et al., 2026 ) | 50 | 0.0672 | 12.84 | 0.411 | 0.695 |
| Matrix-Game 3.0 ( Wang et al., 2026 ) | 49 | 0.0737 | 12.51 | 0.387 | 0.635 |
| AlayaWorld ( AlayaWorld Team et al., 2026 ) | 50 | 0.0815 | 11.97 | 0.379 | 0.659 |
| Alaya-EVOKE ( Yin et al., 2026 ) | 50 | 0.0814 | 11.65 | 0.369 | 0.687 |
| LingBot-World ( Robbyant Team et al., 2026 ) | 50 | 0.0855 | 11.96 | 0.363 | 0.669 |
| Loci (ours) | 50 | 0.0478 | 14.12 | 0.454 | 0.635 |
| Method | Memory-segment feed | Control signal | Res. / fps | Steps / CFG |
|---|---|---|---|---|
| HY-WorldPlay ( Sun et al., 2026 ) | Retrieval-augmented KV-cache: memory GT frames VAE-encoded to clean latents and written into the latent tensor in place of the noise-initialized first chunk; later chunks retrieve by camera-FOV overlap over the full latent history (avg. 1141 frames, range 255–2406) | Continuous GT pose per predicted frame (no WASD discretization) | 832 480, 24 fps | 4 / none |
| Matrix-Game 3.0 ( Wang et al., 2026 ) | Long-range FOV-retrieval bank ( x_memory , selected from the official AR history tensors) plus a 16-frame local re-diffusion window; memory GT frames VAE-encoded to clean latents and prefilled into the same history tensors (avg. 1144 frames, range 255–2406) | Discrete WSAD + mouse actions, quantized from MIND’s continuous GT poses (12.35 cm/frame translation; yaw /frame) | 1280 704 native (17 fps) 1280 720 at 24 fps for scoring | 3 / none |
| LingBot-World ( Robbyant Team et al., 2026 ) | Rolling self-attention KV-cache (sink 6 + recent 12 latent frames); the full memory segment is causally VAE-encoded chunk by chunk and written with the same clean-latent ( ) forward pass the model uses for its own history, but only the sink ( 1.3 s) and tail ( 3 s) remain in the window | Continuous GT pose via Plücker camera embedding | 832 464, 16 fps | 4 / none |
| AlayaWorld ( AlayaWorld Team et al., 2026 ) | Sink (1 frame) + 4-latent history window + 9-frame RGB motion neighbors from the last 25 memory frames, plus the entire memory segment written into a ViGeo spatial point-map bank (top-10 frames retrieved per step by target-pose coverage) | Continuous GT camera pose (OpenCV c2w) | 544 960, 24 fps | 4 / 3.0 |
| Alaya-EVOKE ( Yin et al., 2026 ) | World-state-library prefill (v2v mode): the entire memory segment and its GT poses are fed to a persistent point-cloud world-state library (per-chunk monocular-depth back-projection); read = pose-indexed retrieval, top-8 covisible sources, z-buffer warp | Continuous GT camera pose | 384 640, 24 fps | 3 / none (guidance 1.0) |
| CaR ( Peng et al., 2026 ) | Context frames: the entire memory segment (arbitrary length) is VAE-encoded and temporally resampled onto the memory-token grid by the official memory encoder | Continuous GT camera pose | 480 832, 24 fps | 50 / 3.0 |
| Task / conditioning | Softmax access | Recurrent path | Training | Memory evaluation | |
|---|---|---|---|---|---|
| Video SSM ( Po et al., 2025 ) | actions (Memory Maze, Minecraft) | dense local frame attention | block-wise state-space scan (per the authors, trading spatial consistency for memory) | trained on game data | spatial retrieval on synthetic scenes |
| Hybrid Forcing ( Li et al., 2026b ) | text-to-video | local window | additive linear state of evicted KV | dense-to-hybrid distillation | – |
| ARL 2 ( Li et al., 2026a ) | text-to-video, no camera | current frame (hybrid layers); causal softmax (others) | gated delta rule with per-token gates | conversion by layer-output and velocity distillation | VBench |
| SANA-WM ( Zhu et al., 2026 ) | single image + camera path | attention sink + local window | Gated DeltaNet, one recurrent step per latent frame; UCPE camera encoding | multi-stage training | revisit vs. another generated frame |
| Loci (ours) | memory segment or first frame + camera path | all retained observations (full, or first + bank of 20 + 8 most recent) | KDA, token-wise writes, retention once per chunk, reads from the preceding state; PRoPE on | direct diffusion-forcing training from Wan2.2 | revisit vs. recorded ground truth |
| # | Model | Score | # | Model | Score |
|---|---|---|---|---|---|
| 1 | HY-World 1.5 ( Sun et al., 2026 ) | 84.92 | 12 | Wan 2.7 | 71.00 |
| 2 | Loci (ours) | 82.00 | 13 | LTX 2.3 | 70.17 |
| 3 | Matrix-Game 3.0 ( Wang et al., 2026 ) | 80.37 | 14 | LingBot-World ( Robbyant Team et al., 2026 ) | 67.14 |
| 4 | Genie 3 | 78.39 | 15 | InSpatio-World | 66.47 |
| 5 | Happy Oyster | 75.83 | 16 | LongCat-Video | 66.23 |
| 6 | Kling 3.0 | 75.14 | 17 | Matrix-Game 2.0 | 64.51 |
| Model | History | Reached | 60 s | 120 s | 180 s | Max |
|---|---|---|---|---|---|---|
| Full softmax | full | OOM at 157 s | 63.6 | 102.8 | — | 128.8 |
| Loci | full | OOM at 225 s | 45.5 | 69.0 | 100.2 | 114.9 |
| Full softmax | sparse | 300 s | 27.4 | 27.5 | 27.5 | 27.6 |
| Loci | sparse | 300 s | 23.4 | 23.4 | 23.5 | 23.6 |
| Revisit | Whole clip | Peak memory | |||
|---|---|---|---|---|---|
| Model, history access | PSNR | LPIPS | PSNR | LPIPS | (GiB, grows / constant) |
| Full softmax, full history | 10.16 | 0.576 | 10.18 | 0.587 | grows (OOM at 157 s) |
| Loci , full history | 10.67 | 0.544 | 10.36 | 0.577 | grows (OOM at 225 s) |
| Full softmax, bounded sparse | 10.15 | 0.570 | 10.07 | 0.578 | 27.6, constant |
| Loci , bounded sparse | 11.14 | 0.524 | 10.49 | 0.563 | 23.6, constant |
| Revisit gap | Segments | local error | Chunk lower |
|---|---|---|---|
| 8–20 s | 26 | [ , ] | 17 / 26 |
| 20–60 s | 44 | [ , ] | 28 / 44 |
| 60 s | 35 | [ , ] | 27 / 35 |
| All | 50 | [ , ] | 35 / 50 |