Online novel view synthesis from multi-view streaming videos faces a fundamental trade-off: maintaining a persistent, long-horizon memory to reconstruct temporarily occluded regions while operating under strict real-time constraints. While Test-Time Training (TTT) offers a powerful memory mechanism, standard models mandate gradient-based memory updates at every frame to adapt to the changing motion in dynamic scenes. The computational cost of heavy memory updates precludes real-time application and can lead to instability over long contexts. Given that memory updates are more demanding than memory application and video content is largely redundant, we propose to decouple the frequencies of these two processes. Our approach performs periodic memory updates while applying the memory on a per-frame basis, using cross-view attention to manage deformations between the prior memory state and the current frame. To lock in the historical context, we introduce two critical mechanisms: an auxiliary Memory Loss that forces persistent internalization of the scene, and a Memory Caching strategy that regularizes active weights against catastrophic drift. Our method demonstrates state-of-the-art minute-scale memory persistence in online dynamic human scenes at amortized real-time speed.
Figures & tables
Figure 2 : NSTM Training Scheme. Left: Our alternating training scheme explicitly enforces memorization. In the Memory Supervision step, isolated memory tokens (target camera rays) perform strict self-attention — sharing weights with the cross-view attention layer — without attending to input views. Once the memory is updated by the inputs, these isolated tokens query the memory to reconstruct the target view. This forces a pure memory readout, ensuring context internalization. During Synthesis Supervision , we perform self-attention across input and target tokens to capture ongoing motion, and then query the previously updated memory state. Right: Visualizations of the inputs, targets, and memory in both steps. Crucially, the memory state predates the current observation during synthesis: here the memory holds a standing pose (arms down, bottom-right) while the current inputs show a throwing pose (arms raised, bottom-left). Our cross-view attention resolves this misalignment, transferring the memorized shirt-back pattern onto the current pose.
Figure 3 : Memory Stress Test at T=60 . Despite only observing the back at T=0 , NSTM accurately recalls the hoodie patterns and hairstyles while fusing with current motion. LVSM makes its best stateless guess in unseen regions, LaCT-NVS suffers from out-of-distribution memory drift, and Token-Mem fails to recall accurate back views. See our supplementary website for the video results.
Figure 4 : Memory Stress Test Over Time. We show the back view at T=0. For subsequent timestamps, we show only frontal views and task the model with synthesizing the occluded back. NSTM maintains high-fidelity recall over time while baseline models suffer from drift or collapse to stateless prior over time.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
mPSNR ↑
mSSIM ↑
LVSM
28.99
0.9662
0.0257
0.1034
19.34
0.5670
LaCT-NVS
25.85
0.9500
0.0360
0.1236
16.25
0.5158
Token-Mem
28.97
0.9661
0.0253
0.1019
19.85
0.6220
Ours
29.99
0.9701
0.0225
0.0919
20.71
0.6286
Table 1 : Long-horizon Memory Stress Test ( Sec. 4.2 ) , measured on last frame T=60 . NSTM ranks best on every metric, leading LaCT-NVS by 4.46 dB and Token-Mem by 0.86 dB. As a stateless model, LVSM fails to reconstruct details in the back view when only observing frontal views as seen in Fig. 3 and Fig. 4 . The best , second-best , and third-best results are highlighted.
Figure 5 : Qualitative ablation study. At T=60 , each model synthesizes the back view that has been occluded after T=0 . LaCT-NVS demonstrates the instability of the inner dot-product loss. In contrast, LaCT-NVS w/ L2 greatly improves stability through our L2 inner loss choice, though it still suffers from some drift. LaCT-NVS w/ L2 w/ Caching freezes into an averaged past pose, demonstrating that our memorization-synthesis decoupling with cross-view attention is essential for effective memory caching. Dropping Lmem from our method collapses recall, validating our disentangled memory supervision design ( Sec. 3.2 ). Furthermore, dropping memory caching reduces the recall horizon. Only the full model ( Ours ) provides minute-scale recall of the historical back view while capturing current motion. See our supplementary website for the video results.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
mPSNR ↑
mSSIM ↑
LVSM
28.81
0.9656
0.0296
0.1298
19.69
0.5272
LaCT-NVS
27.01
0.9551
0.0340
0.1414
17.93
0.4982
Token-Mem
28.27
0.9626
0.0293
0.1250
19.45
0.5413
Ours
29.68
0.9691
0.0250
0.1107
20.88
0.5997
Table 2 : Long-horizon Memory from Natural Rotations ( Tab. 2 ) , evaluated at T=60. NSTM ranks best on every metric, leading LaCT-NVS by 2.95 dB and Token-Mem by 1.43 dB in mPSNR. See Sec. B.1 for further qualitative and quantitative results and our supplementary website for video results.
Ablation Study (over 60 timesteps)
Method
PSNR ↑
SSIM ↑
LPIPS ↓
mPSNR ↑
mSSIM ↑
LaCT-NVS (dot-product) baseline
22.32
0.9275
0.0616
13.49
0.4942
LaCT-NVS w/ L2
27.77
0.9594
0.0275
19.08
0.6189
LaCT-NVS w/ L2 w/ Mem Caching
20.25
0.9135
0.0688
12.03
0.4707
Ours wo/ Lmem
28.23
0.9639
0.0251
19.36
0.6198
Ours wo/ Mem Caching
29.24
0.9673
0.0232
19.78
0.6321
Table 3 : Quantitative Ablation Study . We ablate all variants under stage 1 pretraining before finetuning. Metrics are averaged over 60 timesteps over 390 MVHumanNet++ [ 37 ] eval scenes.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Step
Latency (ms) ↓
FPS ↑
Memorization
58.14
17.19
Synthesis
27.01
37.02
Appendix
Table 4 : Computation Asymmetry. The memorization step (including both memory update and apply) is significantly slower than the synthesis step (apply only). Inference performance breakdown measured on a single H100 GPU with a resolution of 256x256.
Average over T=60
Last frame
Method
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
mPSNR ↑
mSSIM ↑
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
mPSNR ↑
mSSIM ↑
LVSM
28.81
0.9654
0.0286
0.1216
19.21
0.5069
28.81
0.9656
0.0296
0.1298
19.69
0.5272
LaCT-NVS
29.13
0.9667
0.0253
0.1045
19.98
0.5850
27.01
0.9551
0.0340
0.1414
17.93
0.4982
Token-Mem
27.76
0.9607
0.0299
0.1198
18.69
0.5254
28.27
0.9626
0.0293
0.1250
19.45
0.5413
Ours
29.53
0.9692
0.0247
0.1056
20.39
0.5928
29.68
0.9691
0.0250
0.1107
20.88
0.5997
Appendix
Table 5 : Memory from Natural Rotations Test. Best , second best , and third best are highlighted. Memory is critical for this evaluation protocol: the sequences feature natural rotations that reveal the back of the subject, which the models must accurately synthesize. Because the memory of stateful baselines degrades over time, NSTM significantly outperforms Token-Mem and LaCT-NVS by 1.43 dB and 2.95 dB in mPSNR, respectively, when measured on the final frame. Furthermore, by also surpassing the stateless LVSM, NSTM clearly demonstrates the advantage of our robust memory persistence. See our supplementary website for video results.
Novel View Synthesis Over 60 Timesteps
Method
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
mPSNR ↑
mSSIM ↑
LVSM
29.48
0.9699
0.0225
0.0917
20.10
0.6509
LaCT-NVS
26.84
0.9567
0.0281
0.0997
17.62
0.5869
Token-Mem
25.33
0.9470
0.0328
0.1056
16.50
0.5692
Ours
29.11
0.9696
0.0223
0.0894
19.93
0.6703
Appendix
Table 6 : General Novel View Synthesis Comparison. best , second best , and third best are highlighted. Given stereo inputs from fixed frontal cameras, all models synthesize arbitrary unseen viewpoints over time. The metrics are averaged over 60 timesteps across 390 evaluation scenes. NSTM significantly outperforms prior stateful baselines (LaCT-NVS and Token-Mem) by successfully mitigating long-term memory degradation. Because the sparsity of disocclusions in this protocol limits the benefits of historical context, NSTM performs on par with the stateless LVSM baseline, confirming its strong general NVS capabilities. See videos in our supplementary website for qualitative visual results.
Figure 6 : Update Interval Analysis ( T=60 sequence average over 17 Natural Rotation scenes). Our method is near stride-invariant, losing only 0.73 dB mPSNR ( 20.43→19.70 ) and 0.003 LPIPS ( .024→.027 ), whereas LaCT-NVS degrades by 11.73 dB mPSNR ( 19.98→8.25 ) and 4.2× in LPIPS ( .025→.107 ). Token-Mem is stable but capacity-limited, trailing ours at every interval.
Figure 7 : Memory from Natural Rotations Qualitative Results . We evaluate our method against the baselines on sequences featuring natural rotations that reveal the back of the subject, which the models must synthesize. Left column shows at which timestep the back is revealed. NSTM recalls the clothing graphics. The stateless LVSM cannot synthesize the unseen graphics, while LaCT-NVS and Token-Mem show a degraded memory. See supplementary website for more video results.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
LVSM
29.15
0.9589
0.0395
0.1282
LaCT-NVS
23.65
0.9253
0.1305
0.2957
Token-Mem
30.83
0.9673
0.0317
0.1109
Ours
34.18
0.9786
0.0196
0.0739
Appendix
Table 7 : Category-Agnostic Object-Centric NVS Test . (Evaluated on scenes of dynamic objects; metrics averaged over all 30s of 30 FPS frames). Because LaCT-NVS and Token-Mem update every frame, they are prone to memory drift when evaluated at 30FPS, with LaCT-NVS degrading severely. In contrast, NSTM ranks best across all metrics, leading Token-Mem by 3.35 dB, LVSM by 5.03 dB, and LaCT-NVS by 10.53 dB in PSNR. The best , second-best , and third-best results are highlighted.
Figure 8 : Category-Agnostic Object-Centric NVS Qualitative Results . We demonstrate our method’s prior-free nature by training and evaluating on a category-agnostic dynamic object-centric dataset. Top: The race car stripe is memorized at T=31 and recovered by NSTM at T=695 after 22 updates, while LVSM synthesizes generic color. Token-Mem’s memory degrades and LaCT-NVS drifts out of distribution. Note that NSTM is successful at recovering this pattern even with the input poses far-away from the camera while the synthesized target view is close-up. Middle: NSTM recovers the occluded stripe inside the helmet on frame T=663 that was memorized at T=61 , while LaCT-NVS again drifts out of distribution and Token-Mem produces blurry results due to capacity constraints. Bottom: Only NSTM can recall the textured pattern at T=680 (memorized at T=211 ).
NSTM
LaCT-NVS / Token-Mem
Dataset
Protocol
Frame rate
Timesteps
Update interval
# updates
Update interval
# updates
Synth. : update
MVHumanNet++
Training, Stage 1
1.2 FPS
T=4
2 frames
2
1 frame
4
1:1
Training, Stage 2 / 3
1.2 FPS
T=24
2 frames
12
1 frame
24
1:1
Evaluation (Stress, Natural, NVS, Ablation)
1.2 FPS
T=60
2 frames
30
1 frame
60
1:1
Update-Interval Analysis ( Sec. B.4 ), τ∈{1,2,4,8,16,32}
1.2 FPS
T=60
τ frames
60/τ
τ frames
60/τ
(τ−1) :1
Inference / rendering (supp. 30 FPS videos)
30 FPS
T=1800 (60 s)
30 frames (1 s)
60
1 frame (33 ms)
1800
29:1
Appendix
Table 8 : Memorization and synthesis schedules across all protocols. We summarize the schedules of all protocols described in our paper. Constrained by the annotation provided by MVHumanNet++ [ 37 ] , we train and evaluate with 1.2 FPS frames. This protocol fairly demonstrates our long-horizon quality, as our synthesis steps operate independently in between memorization steps. However, it provides an advantage to baselines without a decoupled strategy, as the lower frame rate means their memory will not degrade as much as it would at 30 FPS. Therefore, we further show video comparison results on real minute-long 30 FPS videos on our supplementary website. For the Category-Agnostic Object-Centric NVS Test, training clips are sampled from the 6 FPS training sequences with a temporal stride s drawn at random per clip ( Sec. C.3 ), while the evaluation is done on 30 FPS videos.
Figure 9 : Online Inference Pipeline and NSTM Layer Design. Left : During inference, we run the memorization step periodically while performing the synthesis step per-frame, which enables us to achieve amortized real-time speed. Right : We introduce the NSTM Layer consisting of cross-view attention, a fast-weight block and a feed-forward layer. We also incorporate a learned scale that allows the network to balance the contribution of memory and attention to each layer’s output.
Figure 10 : Limitations . Left: Update Frequency and Missed Events. Our periodic updates may miss events between memorization steps. The pattern shown at t=1 was presented only at this non-update timestep and is therefore not recoverable at t=60 . Right: Memory Primacy Bias Our method favors memorization during early steps. Here the back is revealed after 21 updates (at t=42 ) and our recovery at t=60 is only partial.
Long-horizon camera-controlled video generation relies on persistent memory to maintain scene consistency. Existing systems follow two strategies to achieve this consistency. Full-history approaches retain all generated observations, causing unbounded storage and retrieval costs. Selective-construction approaches reduce redundancy, but make one-time retention decisions that are never revisited, even as an observation's value changes with the evolving memory bank. Both strategies leave a shared question unresolved: as the generated history evolves, which stored observations should still remain in memory? Our key insight is that the value of a stored observation is not fixed, but relational: it depends on the alternatives currently available in the memory bank. A view supported by many geometrically and visually similar substitutes can be relinquished with little loss of coverage, whereas an observation with few viable alternatives should remain regardless of age. We introduce Keepsake, an online, training-free controller for fixed-capacity spatial memory. At each update, Keepsake constructs a pose-appearance graph over retained and newly generated observations, combining camera-pose proximity with visual similarity. A retention priority jointly captures the number of strong substitutes and the similarity of the closest alternative, allowing Keepsake to continually reassess memory value, preserve observations with little alternative support, and evict highly replaceable ones under a fixed budget. The controller modifies only the persistent-memory update; the host generator, denoising schedule, and retrieval rule remain unchanged. Across MemCam and WorldMem, Keepsake improves FVD and LPIPS under a fixed memory budget. On 180-second MemCam trajectories, it retains only 32 of 5,397 frames while reducing FVD by 35.1%.
Abdul Mohaimen Al Radi, Kunyang Li, Yuzhang Shang +2
Institute of Artificial Intelligence, University of Central Florida
Autoregressive video generators synthesize long videos by generating successive temporal segments, but their historical KV cache grows with video length. Existing bounded-cache methods reduce this cost with local windows, sink tokens, or compressed memory states, yet they usually assign fixed roles to different parts of the history. We propose FadeMem, a distance-aware KV memory consolidation mechanism that organizes historical KV blocks into a temporal hierarchy under a fixed cache budget. This design is motivated by frequency-dependent temporal decay: fine details decorrelate quickly, while coarse scene structure and identity remain useful over longer horizons. During generation, new history is inserted as fine-grained entries, while older adjacent entries are progressively merged under a power-law temporal allocation schedule, yielding a dense-near, sparse-far memory within one cache. Without architectural changes, FadeMem preserves recent context for short-term dynamics and compact long-range anchors for identity and scene coherence. Experiments show improved subject consistency, background stability, and temporal coherence over existing bounded-cache strategies.
Yu Lu, Junjie Yang, Piotr Koniusz +2
Zhejiang University · University of New South Wales (UNSW) · Data61/CSIRO +1
Online streaming video understanding requires models to process continuous visual inputs and respond to user queries in real time, where the unbounded stream and unpredictable query timing turn memory management into a central challenge. Existing methods typically compress visual tokens via visual similarity heuristics, or augment compression with KV-cache-level retrieval. However, compression decisions rarely incorporate semantic signals, and retrieval is often added after compression is finalized, making the two stages hard to coordinate. We present SAVEMem, a training-free dual-stage framework that brings semantic awareness into memory generation and lets the retrieval scope adapt per query. In Stage1, SAVEMem builds a three-tier streaming memory online under a constant memory budget. A fixed pseudo-question bank provides a lightweight semantic prior, so that long-term retention is shaped by semantic salience rather than visual similarity alone. In Stage2, SAVEMem performs query-aware retrieval over this memory. An anchor-conditioned recency gate adapts the retrieval scope from short-term to mid- and long-term memory based on whether the query targets the present or the distant past. Within this scope, late interaction between query and memory tokens selects candidate frames for answering. Applied to Qwen2.5-VL without training, SAVEMem improves the OVO-Bench overall score from 52.27 to 62.69 and yields consistent gains on StreamingBench and ODV-Bench, while reducing peak GPU memory by 48% at 128 frames over the backbone.
Hang Wu, Sherin Mary Mathews, Yujun Cai +2
University of California, Merced · US Bank · University of Queensland