Long-horizon autoregressive video generation is limited by a finite context window. When an object or scene falls out of context, its fine-grained visual details may be lost and difficult to recover upon reappearance. To retain access to such visual details, we introduce MosaiChunk, a spatio-temporal memory mechanism that composes a mosaic of selected historical key-value (KV) entries across space and time. Our approach is motivated by the observation that a frozen video generator can directly consume such non-contiguous historical KV and recover the corresponding visual content. We therefore keep the generator fixed and learn only a lightweight router that determines which historical sections to include in the mosaic under a fixed active-memory budget. We further introduce RememBench, a benchmark of long-horizon revisits with prompt-driven text-to-video (T2V) and camera-driven image-to-video (I2V) splits. Our experiments show that MosaiChunk consistently improves revisit consistency over both sliding-window inference and whole-chunk retrieval under matched memory budgets, across both T2V and I2V settings.
Figures & tables
Figure 2: Can a frozen backbone reuse a non-contiguous subset of historical KV? To test whether a frozen video generator can consume a non-contiguous subset of KV entries, we select the KV at positions corresponding to the marked cookie region. Only the colored entries are supplied alongside the sliding window. The frozen video generator preserves the cookie’s appearance when the tin opens again.
Figure 3: Memory architecture. (a) Each chunk is partitioned into sections and each section is encoded into a descriptor. (b) Sections in the latest chunk query history outside the sliding window. (c) The N sections with the highest scores are extracted and composed into a MosaiChunk .
Figure 4: Training pipeline. The far memory consists of whole historical chunks for the teacher and a smaller MosaiChunk composed by our memory router for the student. The far memory is joined with the same sliding window and sent into the frozen DiT. Matching their velocity predictions updates the descriptor space of the encoder.
∥Fc∥=1
∥Fc∥=2
Method
CLIP ↑
LPIPS ↓
TempSSIM ↑
Drift ↓
CLIP ↑
LPIPS ↓
TempSSIM ↑
Drift ↓
Base
0.755
0.659
0.923
0.051
0.751
0.652
0.926
0.050
MoC
0.832
0.594
0.928
0.052
0.841
0.569
0.930
0.053
Ours
0.899
0.567
0.932
0.052
0.936
0.500
0.932
0.055
Table 1: T2V results on RememBench . Median CLIP and LPIPS on 100 marked departure–revisit pairs, with rollout-quality metrics alongside. Budgets are in chunk equivalents; Base uses the same total active KV cache budget. Best CLIP and LPIPS at each budget are bold.
∥Fc∥=1
∥Fc∥=2
Method
CLIP ↑
LPIPS ↓
TempSSIM ↑
Drift ↓
CLIP ↑
LPIPS ↓
TempSSIM ↑
Drift ↓
Rotation
Base
0.768
0.655
0.457
0.051
0.793
0.650
0.454
0.052
MoC
0.807
0.645
0.457
0.050
0.807
0.642
0.460
0.049
Ours
0.840
0.625
0.456
0.051
0.863
0.609
0.458
0.051
Rotation + translation
Table 2: I2V results on RememBench . Median revisit CLIP and LPIPS for 180∘ turns: 150 scenes with rotation and 100 with translation. Budget and bolding conventions follow Table 1 .
Figure 5: T2V revisits at a two-chunk far-memory budget. Rows (Base, MoC, and Ours) share the prompt and noise within each scene. The first and last columns show the departure and revisit frames selected for evaluation. Only MosaiChunk preserves the vegetables and their arrangement (top) and the cupboard’s contents (bottom).
WBench
WorldMark
Consistency
Video quality
Consistency
Video quality
∥Fc∥
Method
Spatial
Gated Spatial
Geom.
Subject
Aesth.
Imaging
Revisit Memory
Aesth.
Percept.
1
Base
78.28
75.47
85.98
88.13
61.75
67.31
80.14
61.43
85.49
Ours
79.63
76.69
86.74
88.36
61.86
67.42
84.89
62.48
86.45
2
Base
79.43
76.78
86.14
88.10
61.76
67.40
84.63
61.05
84.97
Ours
80.73
78.18
87.08
88.38
61.92
67.45
84.87
62.89
86.64
Table 3: I2V results on public benchmarks. WBench uses the navigation split; WorldMark uses the first-person splits. Budgets specify MosaiChunk ’s far memory in chunk equivalents; Base matches the total active KV cache budget. Higher is better for all metrics; bold marks the better consistency score at each budget.
Figure 6: I2V revisits at two-chunk (top) and one-chunk (bottom) far-memory budgets. Rows (Base, MoC, and Ours) share the conditioning frame and 180∘ camera trajectory within each scene. The first column shows the departure frame; the last shows each method’s revisit frame, selected from its Pi3X reconstruction. Intermediate columns sample the camera turn; camera insets schematically show the commanded rotation. MosaiChunk better recovers the garden entrance’s archway (top) and the playroom’s layout and wall decorations (bottom).
Router design
∥Fc∥=1
∥Fc∥=2
Learned descriptors
Global top- N
CLIP ↑
LPIPS ↓
CLIP ↑
LPIPS ↓
✘
✔
0.815
0.637
0.841
0.640
✔
✘
0.792
0.656
0.792
0.655
✔
✔
0.840
0.625
0.863
0.609
Table 4: Memory architecture ablation. Median revisit CLIP and LPIPS on 150 I2V rotation scenes. Disabling descriptor learning uses mean-pooled keys; disabling global selection uses per-query top- k . The last row is the full MosaiChunk router, with results from Table 2 .
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
T2V (H3-AR)
I2V (LingBot-World-Infinity)
Transformer layers
50
40
Attention heads × head width
56×128
40×128
Key width across heads
7168
5120
Latent frames per chunk
5
4
Spatial patch grid
24×43
30×52
Visual tokens per chunk
5160
6240
Appendix
Table 5: Memory router settings for Stage 1. Token counts refer to visual tokens only; layer indices are zero-based. Keys from the clustering layer are used to partition each chunk into equal-size sections. Pooled keys from the descriptor layers are used to compute each section’s descriptor.
T2V (H3-AR)
I2V (LingBot-World-Infinity)
Hardware
32 NVIDIA H200
16 NVIDIA H200
Parallelism
FSDP, sequence parallel 4
DDP
Sequences per step
8
16
Trainable parameters
13.65 M
11.55 M
Optimizer
AdamW
AdamW
Learning rate
10−4
2×10−4
Appendix
Table 6: Self-distillation configuration. Only the router is optimized. Scheduled steps describe the configured training schedule; reported checkpoints identify the weights used for evaluation.
Figure 7: T2V prompt-compliance failures. Each row shows a discarded scene under sliding-window inference. The upper band marks the prompt transitions and sampled instants. Top: the container closes but never reopens. Bottom: the container never closes.
Purpose
Prompt
Scene class
Is this an indoor scene or an outdoor scene? Answer with one word: indoor or outdoor.
Eye-level view
Is this photo taken from roughly a standing person’s eye level, looking horizontally – not from high above the scene, not tilted down at the ground, not tilted up at the sky? Answer with one word: yes or no.
Scene prompt
Describe this scene in two or three sentences, as a caption for the image. Name the place, the main structures and surfaces, the materials, the lighting and the weather. Do NOT describe any camera motion, and do not say ’the camera’ or ’the video’ – describe only what is visible.
Appendix
Table 7: Prompts for I2V scene preparation. These prompts cover scene classification, eye-level screening, and captioning for both indoor and outdoor frames.
Figure 8: Example frames from the two splits. Top: model-generated frames from the T2V split when the prompt first reveals the object. Bottom: conditioning frames from the I2V split, taken from the first frames of DL3DV videos. Sixteen scenes are sampled at random from each split.
T2V
I2V
Scenes
100
50 indoor + 100 outdoor
Backbone
H3-AR
LingBot-World-Infinity
Resolution
1376×768
832×480
Frames / frame rate
379 / 24 fps
253 / 16 fps
Revisit control
Four-segment prompt
Camera trajectory
Rotation
—
90∘ , 180∘ , 360∘ ; 150 scenes
Appendix
Table 8: RememBench evaluation settings. The same inputs are used for all methods. Outdoor I2V scenes support both trajectory types.
Figure 9: T2V frame annotation. During revisit annotation, the timeline controls only the selected rollout (red outline); the other videos remain paused. Q marks the shared departure frame, and E marks the selected rollout’s revisit frame. Saved departure–revisit pairs appear below the timeline.
Metric
Definition
CLIP ↑
Cosine similarity of L2-normalized image embeddings from CLIP ViT-H/14, using the LAION-2B checkpoint and its standard image processor.
LPIPS ↓
LPIPS with AlexNet (version 0.1), evaluated at each backbone’s native image resolution after scaling pixels to [−1,1] .
TempSSIM ↑
Mean SSIM over all consecutive decoded frame pairs in grayscale, with an 11×11 Gaussian window and σ=1.5 .
Drift ↓
Mean cosine distance between adjacent chunks. Each chunk is represented by the normalized average of CLIP embeddings from four evenly spaced frames within that chunk.
Appendix
Table 9: Metrics on RememBench . CLIP and LPIPS assess revisit consistency; TempSSIM and Drift provide complementary measures of rollout quality. Arrows indicate the preferred direction.
Figure 10: Additional T2V revisits at a two-chunk far-memory budget. The first column shows the departure frame before the target content leaves view; the last shows each method’s marked revisit frame after it returns. Intermediate columns sample the rollout. Ours denotes MosaiChunk .
Figure 11: Additional I2V revisits at a one-chunk far-memory budget. All methods start from the same conditioning frame and follow the same 180∘ trajectory. The first column shows the conditioning frame, the intermediate columns sample the camera turn, and the last shows each method’s selected revisit frame. Camera insets schematically show the commanded rotation. Ours denotes MosaiChunk .
1 chunk
2 chunks
Metric
Base
MosaiChunk
Base
MosaiChunk
Consistency
Background
91.33
91.57
91.44
91.52
Spatial
78.28
79.63
79.43
80.73
Gated Spatial
75.47
76.69
76.78
78.18
Segment
97.47
97.47
97.47
94.94
Appendix
Table 10: Full WBench results. Consistency and Video Quality metrics on the navigation split: Spatial and Gated Spatial use 60 round-trip cases, Subject uses 103 cases, and the remaining metrics use all 158 cases. Budgets specify MosaiChunk ’s far memory in chunk equivalents; Base matches the total active KV cache budget by retaining additional recent chunks. Higher scores are better. The higher consistency score at each budget is bold, except ties; video-quality scores are not bolded.
1 chunk
2 chunks
Metric
Base
MosaiChunk
Base
MosaiChunk
Consistency
Local Memory
95.55
95.29
93.70
93.46
Global Memory
62.15
62.22
61.75
63.00
Revisit Memory
80.14
84.89
84.63
84.87
Video Quality
Appendix
Table 11: Full WorldMark results. All World Memory and Visual Quality metrics, macro-averaged over the Real and Stylized first-person splits, with 125 videos per method and budget in each split. Budgets specify MosaiChunk ’s far memory in chunk equivalents; Base matches the total active KV cache budget by retaining additional recent chunks. Higher scores are better. The higher consistency score at each budget is bold; video-quality scores are not bolded.
CLIP ↑
Trajectory
Budget
Base
MoC
MosaiChunk
WorldKV
MosaiChunk-Pose
90∘
1 chunk
0.910
0.910
0.913
0.915
0.912
2 chunks
0.907
0.908
0.906
0.915
0.910
90∘ + translation
1 chunk
0.882
0.884
0.889
0.900
0.896
2 chunks
0.881
0.875
0.887
0.895
0.889
360∘
1 chunk
0.716
0.722
0.717
0.728
0.704
Appendix
Table 12: I2V results at 90∘ and 360∘ . Median revisit CLIP and LPIPS on 150 rotation scenes and 100 scenes with rotation and translation. Retrieval methods use one- or two-chunk far-memory budgets with a fixed sliding window; Base matches the total active KV cache budget by retaining additional recent chunks. WorldKV and MosaiChunk-Pose use camera poses; MoC scores pooled keys, while MosaiChunk uses learned section descriptors.
Method
CLIP ↑
LPIPS ↓
TempSSIM ↑
Drift ↓
Far memory: 1 chunk
Rotation
MosaiChunk
0.840
0.625
0.456
0.051
WorldKV
0.853
0.610
0.458
0.051
MosaiChunk-Pose
0.843
0.624
0.459
0.051
Rotation + translation
Appendix
Table 13: Effect of pose-guided retrieval. Revisit CLIP and LPIPS (medians) and rollout-quality metrics on 150 rotation scenes and 100 scenes with rotation and translation, all with 180∘ turns. WorldKV and MosaiChunk-Pose use camera poses; MosaiChunk uses learned section descriptors without pose guidance. Far-memory budgets are measured in original video-chunk equivalents, with the sliding window fixed. Better CLIP and LPIPS among the two pose-guided methods are bold.
Figure 12: Pose-guided retrieval at a one-chunk far-memory budget. The left panel is the conditioning frame. Compared with WorldKV, MosaiChunk-Pose better preserves the background building’s staircase.
Figure 13: Revisit consistency across far-memory budgets. Median CLIP on 100 T2V scenes and 150 I2V 180∘ rotation scenes. MosaiChunk and MosaiChunk-Pose use half-chunk increments; other methods use one- and two-chunk budgets. Base matches the total active KV cache budget by retaining additional recent chunks.
Figure 14: Transplanting content between rollouts. Rows specify the tin and columns the cookie. Diagonal images are source frames; off-diagonal images show the swapped results.
Autoregressive (AR) video generation extends videos by producing latent chunks sequentially, but scaling to long videos requires repeated access to a growing historical KV cache. Existing methods reduce this cost by truncating the KV cache or compressing it into implicit memory, but both lose explicit access to query-relevant historical details. We propose OmniMem, an explicit full-range memory retrieval framework that performs sparse KV retrieval over the historical cache. To make this practical for chunk-based AR video generation, OmniMem addresses two issues: (i) local bias in sparse KV selection and (ii) Union Explosion in memory access. Adaptive Window Exclusion removes local-window blocks from the selection candidates when sufficient long-range history is available, preserving the sparse budget for informative long-range retrieval. Query-Shared KV Selection reduces cross-query diversity, while Per-Head Scattered KV Access avoids expanding head-specific selections into a large selected KV buffer. This allows each attention head to retrieve non-contiguous KV blocks according to its own selection pattern. Experiments on long-video generation show that OmniMem improves Dynamic Degree by 52.3% and preserves strong consistency over strong baselines, while maintaining comparable memory usage.
Autoregressive video generators synthesize long videos by generating successive temporal segments, but their historical KV cache grows with video length. Existing bounded-cache methods reduce this cost with local windows, sink tokens, or compressed memory states, yet they usually assign fixed roles to different parts of the history. We propose FadeMem, a distance-aware KV memory consolidation mechanism that organizes historical KV blocks into a temporal hierarchy under a fixed cache budget. This design is motivated by frequency-dependent temporal decay: fine details decorrelate quickly, while coarse scene structure and identity remain useful over longer horizons. During generation, new history is inserted as fine-grained entries, while older adjacent entries are progressively merged under a power-law temporal allocation schedule, yielding a dense-near, sparse-far memory within one cache. Without architectural changes, FadeMem preserves recent context for short-term dynamics and compact long-range anchors for identity and scene coherence. Experiments show improved subject consistency, background stability, and temporal coherence over existing bounded-cache strategies.
Yu Lu, Junjie Yang, Piotr Koniusz +2
Zhejiang University · University of New South Wales (UNSW) · Data61/CSIRO +1
Autoregressive (AR) video generation degrades over long horizons due to an overlooked train-inference discrepancy we term KV eviction mismatch: models train on short clips where all context frames reside in the KV cache, but at inference, memory constraints force distant frames to be evicted from the KV cache - removing context the model was conditioned on. Rather than simulating eviction via context truncation - which discards temporal information the model still needs and degrades motion coherence - we keep the context but while progressively reducing the influence of distant frames, making their eventual eviction negligible. To guide this design, we introduce the positional response R(Δ,tdenoise), a perturbation-based sensitivity measure revealing that context influence decays steeply with temporal distance and varies systematically across denoising steps. Motivated by this analysis, we propose Recency Forcing, which applies a non-positive, timestep-dependent bias, termed Temporal Response Bias (TRB), on pre-softmax attention logits derived directly from R, closing the train-inference gap without modifying context length or training objectives. We further introduce Biased Attention Reparameterization (BAR), an exact reformulation that moves the bias outside the softmax, making TRB a standard FlashAttention call at zero overhead. Recency Forcing operates in both training-free mode and training-based mode. Experiments on VBench and VBench-Long demonstrate state-of-the-art long-horizon generation quality at no additional inference cost.