Standard video generators do not natively compact historical context into reusable memory tokens. As generation continues, the growing history makes it increasingly difficult to retain information from earlier frames due to long-context degradation. Key-frame-based approaches address this challenge by retaining selected past frames, but can discard information needed for future generation. Rather than relying on frame selection alone, we study whether a frozen video generator can supply the supervision needed to learn a compact representation of the history. We propose Prediction-Aligned Context Compaction (PACC), which uses a learned compressor to aggregate information across past frames into compact memory tokens. We train the compressor through on-policy distillation, using the same frozen generator both as a student when conditioned on compressed memory and as a teacher when conditioned on the full history. The student generates continuations, while the teacher provides targets for the same noisy inputs at each denoising step. Only the compressor is updated to align the student's predictions with these targets. We evaluate PACC on MBench, which jointly measures memory-event coverage and consistency. PACC outperforms the strongest baseline by 6.63 points on Causal-rCM and 3.19 points on Causal Forcing. Evaluation on VBench-Long using MovieGen prompts further shows that PACC produces minute-long videos with generation quality competitive with baselines. Together, these results show that learning to compact historical context can improve long-video memory without modifying the underlying generator.
Figures & tables
Figure 1: Frame-based context versus learned memory compaction. (a) Frame-based approaches select existing historical frames or tokens to condition future generation. (b) Prediction-Aligned Context Compaction (PACC) compacts each completed video block into learned memory tokens, stores them in an archive, and selectively retrieves compressed blocks within a bounded active context. Highlighted memory groups are retrieved; faded groups remain archived. Compaction changes the representation available for retrieval rather than eliminating selection.
Figure 2: PACC training through on-policy distillation. The compressor Cϕ maps a model-generated prefix P to compact memory M=Cϕ(P,c) . A short student rollout supplies noisy latents z and previously generated chunks y<j . The full-memory teacher and compressed-memory student share the frozen generator Gθ , text condition c , and rollout inputs, but condition on P and M , respectively. Matching their velocity predictions vt and vs updates only Cϕ through the student branch. The figure abbreviates zj,s , vj,st , and vj,ss as z , vt , and vs , where j indexes continuation chunks and s indexes denoising steps. The shared noise level ts and the compressor’s text-conditioning connection are omitted. The loss box shows one squared-error term of Equation 4 .
Method
MBench ↑
VBench-Long ↑
Human
Object
Causal
M-score
Avg.
Causal-rCM c3-3
Infinity-RoPE ( Yesiltepe et al., 2026 )
31.83
23.83
60.92
38.86
80.09
Relax Forcing ( Zhao et al., 2026 )
27.10
36.21
65.74
43.01
81.00
Rolling Sink ( Li et al., 2026a )
15.96
4.76
42.25
20.99
77.43
Deep Forcing ( Yi et al., 2025 )
31.57
34.56
49.20
38.44
79.94
Table 1: Long-horizon memory and generation quality. MBench evaluates 26-second videos; VBench-Long evaluates minute-long videos on MovieGenBench. Human, Object, and Causal average the paired M-scores for identity/appearance, geometry/texture, and state/correctness, respectively. M-score averages all six memory dimensions; VBench-Long Avg. averages six quality dimensions and is not the official VBench total. Scores are on a 0–100 scale; higher is better. Bold and underlining mark the best and second-best values in each column within each backbone.
Figure 3: Qualitative memory comparison on Causal Forcing. Columns follow five consecutive prompt segments, with excerpts shown below. The subject is asked to leave the frame and later return with her appearance unchanged. PACC depicts the empty-scene stage and subsequent return while retaining her hairstyle and black high-necked blouse. In contrast, the subject remains visible in both baselines at the displayed empty-scene stage.
Figure 4: Training and compression affect memory and imaging quality differently. (a,b) Training-duration sweep at ρ−1=10 ; stars mark the 5,000-update main checkpoint. (c,d) Compression-factor sweep at 2,000 updates. M-score uses the MBench subset; imaging quality uses 64 minute-long MovieGen videos per setting. The complete sweep and complementary quality dimensions are reported in Appendix C . Lines are visual guides.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Identity
Appear.
Geom.
Texture
State
Correct.
M-Score
Causal-rCM c3-3
Infinity-RoPE ( Yesiltepe et al., 2026 )
28.22
35.44
14.01
33.65
65.38
56.46
38.86
Relax Forcing ( Zhao et al., 2026 )
25.51
28.68
24.28
48.14
64.42
67.06
43.01
Rolling Sink ( Li et al., 2026a )
13.00
18.92
3.70
5.81
45.89
38.60
20.99
Deep Forcing ( Yi et al., 2025 )
27.40
35.74
21.23
47.89
52.05
46.35
38.44
MemRoPE ( Kim et al., 2026 )
28.62
35.31
11.88
50.14
64.20
56.34
41.08
Appendix
Table 2: MBench memory preservation at 26.06 seconds. Results use Seed 2.0 triggers and Qwen3-VL-Plus causal-reliability judges, with one seed-0 video per case and method. Scores are on a 0–100 scale; M-score is the unweighted mean of the six reported dimensions. Bold denotes the best value within each backbone.
Method
Subject
Backgr.
Motion
Aesthetic
Imaging
Dynamic
Avg.
Causal-rCM c3-3
Infinity-RoPE ( Yesiltepe et al., 2026 )
97.45
96.45
98.47
61.02
70.01
57.14
80.09
Relax Forcing ( Zhao et al., 2026 )
97.13
96.21
98.35
60.84
68.83
64.64
81.00
Rolling Sink ( Li et al., 2026a )
98.11
96.95
98.91
61.41
69.71
39.46
77.43
Deep Forcing ( Yi et al., 2025 )
97.12
96.32
98.48
59.64
67.89
60.16
79.94
MemRoPE ( Kim et al., 2026 )
97.47
96.50
98.64
60.39
69.27
55.56
79.64
Appendix
Table 3: Minute-long video quality on MovieGenBench. VBench-Long evaluation on the first 128 prompts with five seeds each (640 videos per method; 957 frames at 16 fps). Baselines are inference-policy adaptations on the same frozen backbone. Scores are on a 0–100 scale; Avg. is the unweighted mean of the six dimensions, not the official VBench total. Bold marks the best value per column within each backbone; underlining marks the second-best average.
Figure 5: Complementary quality dimensions across training. VBench-Long aesthetic quality and dynamic degree at ρ−1=10 , on the same 64-video subset used for imaging quality. Stars mark the 5,000-update main checkpoint. Relative to 2,000 updates, it has higher aesthetic quality but lower dynamic degree. These dimensions are reported separately, not combined into a new score.
Updates
Identity
Appear.
Geom.
Texture
State
Correct.
M-score
1k
30.55
32.99
30.24
62.06
63.27
64.46
47.26
2k
45.12
50.76
34.99
66.72
69.06
71.31
56.33
3k
41.05
48.44
31.86
57.26
69.32
70.53
53.08
4k
43.83
49.75
34.44
54.38
64.79
64.95
52.02
5k †
35.24
43.30
26.01
54.13
61.09
66.92
47.78
7k
40.99
38.49
36.26
59.03
60.54
64.79
50.02
Appendix
Table 4: MBench memory in the training-duration sweep at ρ−1=10 . Dimension-level M-scores on the MBench evaluation subset. Scores are on a 0–100 scale. M-score is the mean of the six reported memory dimensions. Bold and underlining mark the best and second-best values per column; tied values share their rank. † marks the main checkpoint.
Updates
Subject
Backgr.
Motion
Aesthetic
Imaging
Dynamic
Avg.
1k
97.62
96.40
98.53
61.12
67.64
57.50
79.80
2k
96.99
96.04
98.34
60.09
65.21
69.38
81.01
3k
97.32
96.20
98.25
60.89
67.18
67.66
81.25
4k
97.28
96.20
98.36
60.69
67.12
65.73
80.89
5k †
97.41
96.28
98.35
61.14
67.21
64.06
80.74
7k
97.49
96.37
98.38
61.40
68.18
62.66
80.74
Appendix
Table 5: VBench-Long quality in the training-duration sweep at ρ−1=10 . 64 approximately one-minute MovieGen videos per setting. Scores are on a 0–100 scale. Avg. is the unweighted mean of the six quality dimensions, not the official VBench total. Bold and underlining mark the best and second-best values per column; tied values share their rank. † marks the main checkpoint.
Factor ρ−1
Identity
Appear.
Geom.
Texture
State
Correct.
M-score
4
32.84
36.46
32.21
61.99
64.38
69.77
49.61
6
29.23
31.60
29.28
62.52
62.38
63.39
46.40
8
38.57
39.63
27.32
55.18
61.59
62.74
47.51
10
45.12
50.76
34.99
66.72
69.06
71.31
56.33
Appendix
Table 6: MBench memory in the compression-factor sweep at 2,000 updates. Dimension-level M-scores on the MBench evaluation subset. Scores are on a 0–100 scale. M-score is the mean of the six reported memory dimensions. Bold and underlining mark the best and second-best values per column.
Factor ρ−1
Subject
Backgr.
Motion
Aesthetic
Imaging
Dynamic
Avg.
4
97.51
96.46
98.43
61.64
68.29
62.40
80.79
6
97.41
96.28
98.35
61.34
68.43
70.52
82.06
8
97.29
96.23
98.42
60.63
65.64
66.61
80.80
10
96.99
96.04
98.34
60.09
65.21
69.38
81.01
Appendix
Table 7: VBench-Long quality in the compression-factor sweep at 2,000 updates. 64 approximately one-minute MovieGen videos per setting. Scores are on a 0–100 scale. Avg. is the unweighted mean of the six quality dimensions, not the official VBench total. Bold and underlining mark the best and second-best values per column.
Figure 6: Qualitative memory comparison on Causal-rCM c3-3. Columns follow five consecutive prompt segments, with excerpts shown below. A bird passes in front of a skier, obscuring the scene before leaving to reveal the skier on the snow. PACC depicts this occlusion-and-reveal sequence, whereas Full AR and Relax Forcing do not show the requested final reveal in the displayed frames.
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China · University of Chinese Academy of Sciences, Beijing, China · China University of Mining & Technology, Beijing +3