FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation
Organizations: National University of Singapore · Harbin Institute of Technology (Shenzhen)
Abstract
Long-horizon video generation requires models to effectively leverage an increasingly long generation history. As the generated history grows, retaining all previous content becomes increasingly expensive and redundant, making effective historical selection essential. Existing approaches often determine historical relevance based on the current content. However, information relevant to the present is not necessarily useful for future generation, while seemingly less relevant history may become important later. Our key insight is that historical information should be selected according to its relevance to future information needs. Capturing these needs does not require generating the full future; instead, a compact representation of what becomes important next is sufficient to guide historical selection. Building on this insight, we propose FrameMorrow, a prospective frame selector that predicts a small set of prospective tokens representing future information needs and uses them to identify relevant information from history. FrameMorrow selects explicit historical frames rather than model-specific internal states, enabling plug-and-play integration across diverse generators, including closed-source models, with little additional inference cost. We evaluate FrameMorrow across five benchmarks and 11 generative models spanning long-video generation, interactive generation, and action-conditioned world models. Extensive experiments demonstrate consistent improvements in long-range consistency, visual quality, and action alignment across diverse generation settings.
Figures & tables
| Method | Subject | Background | Motion | Dynamic | Aesthetic | Imaging | Avg. Rank |
|---|---|---|---|---|---|---|---|
| Self-Forcing | 95.84 | 95.27 | 98.20 | 51.72 | 56.05 | 62.22 | 4.33 |
| -RoPE | 97.24 | 96.24 | 98.58 | 46.64 | 56.09 | 63.28 | 3.17 |
| +Deep Forcing | 96.08 | 95.38 | 98.24 | 41.44 | 56.68 | 60.81 | 4.17 |
| +LongLive-RAG | 97.60 | 96.51 | 98.70 | 44.69 | 57.19 | 64.97 | 2.00 |
| + FrameMorrow (Ours) | 97.71 | 96.45 | 98.60 | 64.56 | 57.27 | 68.47 | 1.33 |
| LongLive 1.0 | 97.13 | 95.89 | 98.61 | 44.56 | 58.17 | 67.56 | 3.83 |
| Overall | Segment-wise CLIP Score | |||||||||
| Model | Quality | Consistency | Aesthetic | 0–10s | 10–20s | 20–30s | 30–40s | 40–50s | 50–60s | Avg. |
| Single-shot | ||||||||||
| LongLive 1.0 | 83.49 | 92.62 | 64.21 | 30.71 | 29.41 | 27.35 | 28.45 | 27.63 | 28.95 | 28.75 |
| + FrameMorrow (Ours) | 84.06 | 93.21 | 64.56 | 30.72 | 29.65 | 28.11 | 28.78 | 28.03 | 29.18 | 29.08 |
| Self-Forcing | 78.71 | 84.95 | 58.16 | 30.40 | 29.52 | 26.93 | 25.10 | 23.37 | 23.02 | 26.39 |
| + FrameMorrow (Ours) | 81.74 | 89.30 | 60.62 | 30.56 | 29.67 | 28.34 | 27.12 | 26.60 | 26.01 | 28.05 |
| Model | Quality | Consistency | Aesthetic |
|---|---|---|---|
| Seedance 2.0 | 86.80 | 94.02 | 65.61 |
| + Uniform | 86.47 | 93.54 | 65.28 |
| + FrameMorrow | 87.46 | 95.31 | 65.92 |
| Kling O3 | 83.47 | 91.28 | 63.41 |
| + Uniform | 83.21 | 90.84 | 63.18 |
| + FrameMorrow | 84.06 | 92.61 | 63.77 |
| Visual Quality | Temporal Quality | Interaction | ||||
|---|---|---|---|---|---|---|
| Model | Subject | Background | Imaging | Anti-flicker | Motion | Action Alignment |
| Matrix-Game 3.0 | 0.801 | 0.894 | 68.20 | 0.936 | 0.950 | 0.842 |
| + FrameMorrow (Ours) | 0.816 | 0.910 | 68.35 | 0.939 | 0.949 | 0.861 |
| WorldMem | 0.782 | 0.924 | 70.23 | 0.910 | 0.916 | 0.866 |
| + FrameMorrow (Ours) | 0.793 | 0.932 | 70.47 | 0.913 | 0.918 | 0.889 |
| YuMe 1.5 | 0.765 | 0.872 | 50.98 | 0.944 | 0.965 | 0.883 |
| Strategy | Single-shot | Multi-shot |
|---|---|---|
| Recent- | 85.70 | 92.80 |
| Uniform | 86.60 | 92.21 |
| Context matching | 87.85 | 92.68 |
| Prompt matching | 88.35 | 93.17 |
| Direct scoring | 88.22 | 92.53 |
| FrameMorrow | 89.30 | 93.49 |
| Method | Sel. (ms) | Total (s) | (%) |
|---|---|---|---|
| Self-Forcing | – | 82.6 | – |
| + FrameMorrow | 112 | 84.0 | 1.7 |
| LongLive 1.0 | – | 67.8 | – |
| + FrameMorrow | 109 | 69.2 | 2.1 |
| Causal Forcing | – | 82.6 | – |
| + FrameMorrow | 115 | 84.0 | 1.7 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Configuration | Value |
|---|---|
| Visual encoder / teacher | DINOv2 ViT-B/14 |
| Visual feature dimension | 768 |
| Transformer layers / hidden dimension | 4 / 512 |
| Attention heads / FFN dimension | 8 / 2,048 |
| Normalization / activation | Pre-LayerNorm / GELU |
| Dropout | 0.0 |
| Symbol | Role | Value |
|---|---|---|
| Query-score aggregation | 0.10 | |
| Future-similarity aggregation | 0.05 | |
| Teacher ranking temperature | 0.10 | |
| Student ranking temperature | 1.00 | |
| Teacher margin for pair inclusion | 0.05 | |
| Pairwise loss weight | 0.50 |
| Configuration | Value |
|---|---|
| Optimizer | AdamW |
| Learning rate / weight decay | / 0.01 |
| Adam coefficients | |
| Schedule / warmup | Cosine / 5% of updates |
| Global batch size / epochs | 64 / 5 |
| Precision / gradient clipping | BF16 / global norm 1.0 |
| Backbone | Selector | Historical conditioning route |
|---|---|---|
| Self-Forcing | Text | Frame encoding followed by historical KV conditioning |
| LongLive 1.0 | Text | Historical frame encoding into the memory cache |
| Causal Forcing | Text | Frame encoding into causal context KV states |
| CausVid | Text | Historical latent encoding into context KV states |
| LongLive 2.0 | Text | Historical references in the streaming memory interface |
| ShotStream | Text | Selected images as cross-shot reference conditioning |
| Score | Criterion |
|---|---|
| 1.00 | All commanded motion directions and their temporal order are clearly followed. |
| 0.75 | The dominant commands are followed, with minor delay or one brief inconsistency. |
| 0.50 | Some commands are followed, but a substantial part is missing or ambiguous. |
| 0.25 | Only weak evidence supports the commands, with mostly inconsistent motion. |
| 0.00 | The commanded motion is absent, reversed, or unsupported by the visible rollout. |
| Visual Teacher | Consistency | |
|---|---|---|
| CLIP | 88.42 | +3.47 |
| SigLIP | 88.76 | +3.81 |
| DINOv2 | 89.30 | +4.35 |
| Objective | Consistency | |
|---|---|---|
| Pairwise only | 88.21 | +3.26 |
| Listwise only | 88.67 | +3.72 |
| Listwise + Pairwise | 89.30 | +4.35 |
| Consistency | ||
|---|---|---|
| 1 | 87.86 | +2.91 |
| 2 | 88.71 | +3.76 |
| 4 | 89.14 | +4.19 |
| 8 | 89.30 | +4.35 |
| Aggregation | Consistency | |
|---|---|---|
| Mean | 88.43 | +3.48 |
| Hard maximum | 88.91 | +3.96 |
| Smooth maximum | 89.30 | +4.35 |
| Ranking signal | Realized future | 95% CI | |
|---|---|---|---|
| Context similarity | No | 0.10 | [0.04, 0.16] |
| FrameMorrow prediction | No | 0.42 | [0.35, 0.49] |
| Future-grounded teacher | Yes | 0.51 | [0.44, 0.57] |