Long-horizon camera-controlled video generation relies on persistent memory to maintain scene consistency. Existing systems follow two strategies to achieve this consistency. Full-history approaches retain all generated observations, causing unbounded storage and retrieval costs. Selective-construction approaches reduce redundancy, but make one-time retention decisions that are never revisited, even as an observation's value changes with the evolving memory bank. Both strategies leave a shared question unresolved: as the generated history evolves, which stored observations should still remain in memory? Our key insight is that the value of a stored observation is not fixed, but relational: it depends on the alternatives currently available in the memory bank. A view supported by many geometrically and visually similar substitutes can be relinquished with little loss of coverage, whereas an observation with few viable alternatives should remain regardless of age. We introduce Keepsake, an online, training-free controller for fixed-capacity spatial memory. At each update, Keepsake constructs a pose-appearance graph over retained and newly generated observations, combining camera-pose proximity with visual similarity. A retention priority jointly captures the number of strong substitutes and the similarity of the closest alternative, allowing Keepsake to continually reassess memory value, preserve observations with little alternative support, and evict highly replaceable ones under a fixed budget. The controller modifies only the persistent-memory update; the host generator, denoising schedule, and retrieval rule remain unchanged. Across MemCam and WorldMem, Keepsake improves FVD and LPIPS under a fixed memory budget. On 180-second MemCam trajectories, it retains only 32 of 5,397 frames while reducing FVD by 35.1%.
Figures & tables
Figure 1: What should remain in memory? Unbounded memory versus our budgeted memory in MemCam. (a) Retained-frame count and CPU lookup latency per query. (b) DINO feature cosine distance from selected historical frames to reference frames over generation time; lower is better. (c) A camera-revisit example, with boxes highlighting corresponding scene regions.
Figure 2: Overview of Keepsake . Top: Keepsake is inserted into the long-horizon generation loop as a memory-update module. After each generated chunk, it updates the persistent memory while leaving the generator and retrieval rule unchanged. Bottom: Existing memory and newly generated observations are pooled into a candidate bank and organized as a pose-appearance graph, where edge weights combine camera-pose proximity and visual similarity. Node statistics capture both the number of strong substitutes and the similarity to the closest substitute, which jointly determine retention priority. Low-priority observations are evicted until the fixed budget B is met. Thumbnails and displayed scores are illustrative.
Consistency
VBench ↑
Efficiency
Models
FVD ↓
LPIPS ↓
Subject
Background
Aesthetic
Average
Stored items ↓
Retrieval ms/query ↓
MemCam
784.9
0.59
80.53
90.16
45.42
73.98
1825
1344.61
MemCam + FIFO
832.2
0.62
77.01
89.50
45.35
73.34
32
39.93
MemCam + MCE
773.9
0.60
79.78
89.97
46.05
73.96
32
54.20
MemCam + K-center
712.0
0.59
80.85
90.00
45.52
74.11
32
43.59
MemCam + KEEPSAKE (Ours)
690.5
0.58
81.78
90.47
46.14
74.65
32
43.77
Table 1: Main quality results on 15 matched 60-second videos within each system. All bounded methods use B=32 . For consistency, we report FVD and LPIPS. For efficiency, we report number of stored items in memory and retrieval time needed for the model for each generation chunk. For additional evaluation we also report Vbench scores across subject and background consistency, as well as aesthetics and average scores.
Figure 3: Generation quality and memory diagnostics. (a) LPIPS and FVD on fifteen matched 180-second MemCam rollouts. (b) View mismatch and memory corruption on fifteen matched 60-second rollouts, measured by DINOv2 cosine distance; whiskers denote 95% trajectory-bootstrap confidence intervals. FIFO and Keepsake use B=32 in (a,b). (c) Retention gap measures useful evidence lost through eviction, while selection gap measures the retriever’s failure to select the best evidence that remains in memory. Results are shown on fifteen 180-second trajectories across memory budgets B . Lower is better for all metrics.
90∘ , 153 frames
360∘ , 609 frames
Method
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
DFoT
16.76
0.47
0.39
8.94
0.25
0.61
GeometryForcing
16.57
0.49
0.35
10.07
0.40
0.57
MemCam
17.83
0.51
0.357
14.81
0.42
0.50
Keepsake
17.66
0.52
0.34
16.22
0.46
0.43
Table 2: Round-trip and beyond-context evaluation. Keepsake improves most consistency metrics, with larger gains on the more challenging 360∘ round-trip setting.
Figure 4: Qualitative comparison of camera revisits. Examples from 180-second MemCam rollouts (left) and 60-second WorldMem rollouts (right). Rows compare unbounded retention, FIFO, and Keepsake with B=32 ; columns show the first visit, an intervening view, and the revisit. Highlighted regions track scene content across visits. Keepsake better preserves the appearance and structure of previously observed regions when the camera returns.
Budget B
LPIPS ↓
FVD ↓
VBench ↑
Query time ↓
MemCam, 180 seconds
16
0.5876
446.1
74.90
24.80 ms
32
0.5876
476.6
74.65
43.77 ms
64
0.5865
493.9
73.99
105.15 ms
128
0.5913
515.5
73.82
209.78 ms
WorldMem, 60 seconds
Table 3: Memory-budget ablation study.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
C(S)=∑q∈Qwq(1−∏m∈S(1−K(q,m))).
Appendix
Algorithm 3 Marginal Coverage Eviction (MCE)
Parameter
Role
Value
B
Retained-observation budget
32
α
Pose–appearance weighting
0.65
T
Strong-substitute threshold
0.65
τ
Degree saturation threshold
3
β
Inverse-degree weight
0.5
λ
Closest-substitute weight
0.25
Appendix
Table 4: Default retention settings. The same values are used for MemCam and WorldMem.
Table 6: End-to-end generation time per completed video. Comparisons are made within each system.
Policy
View mismatch
Memory corruption
Complete retention
0.1256 [0.1041, 0.1479]
0.4871 [0.4012, 0.5716]
FIFO
0.1557 [0.1336, 0.1777]
0.5022 [0.4195, 0.5835]
Keepsake
0.3254 [0.2579, 0.3964]
0.4079 [0.3159, 0.4989]
Appendix
Table 7: Common-source memory diagnostics. Fifteen 60-second trajectories; bounded policies use B=32 . Entries are means with 95% trajectory-bootstrap confidence intervals. Lower distances are better.
FIFO
Keepsake
Budget B
Retention gap
Selection gap
Retention gap
Selection gap
16
0.1872
0.0448
0.0669
0.1037
32
0.1668
0.0646
0.0426
0.1470
64
0.1429
0.0867
0.0366
0.1596
128
0.1203
0.1084
0.0298
0.1783
Appendix
Table 8: Retention-selection diagnostic at 180 seconds. Rounded coordinates from Fig. 3(c) . Complete retention has retention gap 0 and selection gap 0.2188.
Comparison subset
Δ PSNR (dB)
Δ SSIM
Aggregate comparison
+4.629
+0.1512
Exclude initial-image selection disagreements
+1.060
+0.0449
Additionally require FOV overlap ≥0.80 for both selections
+0.103
+0.0098
Appendix
Table 9: Sensitivity of historical-image fidelity. Differences are Keepsake minus complete retention on common-source images, evaluated against ground truth at their historical indices. The restrictions are cumulative.
Figure 5: Additional camera-revisit examples.
Figure 6: Selected common-source retrieval examples. For identical queries, unbounded retention and Keepsake can retrieve substantially different historical observations.