Streaming vision-language models must process continuously growing video streams under a bounded compute budget, creating a persistent tension between real-time perception and long-term memory. Retrieving historical information provides a natural remedy, yet historical recall is not uniformly beneficial: unnecessary history may introduce irrelevant context into current reasoning and interfere with native real-time perception. Effective streaming memory should therefore address not only what to remember, but also when and how to access it. To this end, we introduce FlashBack, a training-free framework for selective, multi-level memory in streaming vision-language models. Before retrieving history, FlashBack draws on the semantic understanding of the frozen streaming VLM to infer whether a query calls for historical evidence. This assessment determines whether inference remains on the Native trajectory or invokes an isolated Recall trajectory. The Recall trajectory combines recent context with retrieved long-term memory through a query-local Side-KV pathway, preserving local temporal continuity without modifying the persistent Native state. We instantiate FlashBack on StreamingVLM and Mage-VL-4B and evaluate it on OVO-Bench and StreamingBench. The results show improvements on several long-horizon and memory-dependent tasks while largely preserving real-time perception, with performance competitive with strong training-based streaming methods despite requiring no additional training. Our code will be announced later.
Figures & tables
Figure 1: Motivation for FlashBack.
Figure 2: Overview of FlashBack.
Method
MA
CRR
Native
69.60
41.67
+ Recall
70.40
60.83
+ Direct Routing
69.20
41.67
+ Structured Routing
72.80
60.83
Table 1: Effect of Recall and Routing.
Memory
MA
CRR
Recent only ( K=0 )
73.60
55.42
K=1
72.80
60.83
K=3
72.80
60.83
K=5
72.40
60.83
Table 2: Effect of Memory Hierarchy.
Integration
MA
CRR
Random KV
33.20
39.58
Zero KV
26.80
39.17
Historical KV
64.40
45.83
Side-KV
70.40
60.83
Table 3: Effect of Side-KV.
Method
RT Avg.
Bwd. Avg.
Fwd. Avg.
Training-based methods
VideoLLM-online
20.79
17.73
–
EventMemAgent-8B
68.29
58.03
55.92
SelectStream-Qwen3-8B
82.76
62.20
56.13
StreamReady-7B
73.60
72.20
58.80
Training-free methods
Table 4: OVO-Bench group-level results.
Method
EPM
ASI
REC
CRR
HLD
StreamingVLM
54.88
56.76
19.34
41.67
27.96
StreamingVLM + FlashBack
54.88
56.76
20.06
60.83
23.12
Mage-VL
52.53
58.78
19.48
44.17
32.26
Mage-VL + FlashBack
53.20
64.86
29.80
54.17
31.72
Table 5: Memory-related tasks on OVO-Bench.
Method
RT Avg.
Omni Avg.
Context Avg.
Training-based methods
VideoLLM-online
35.99
28.45
–
EventMemAgent-8B
77.00
–
–
SelectStream-Qwen3-8B
82.67
–
–
Training-free methods
StreamRAG-ViSpeak-7B
78.12
–
–
Table 6: StreamingBench group-level results.
Method
CR
CS
EU
CT
ACU
StreamingVLM
79.69
87.70
80.38
27.66
52.40
StreamingVLM + FlashBack
81.25
88.33
80.38
30.32
53.60
Mage-VL
67.97
90.22
76.58
40.43
67.20
Mage-VL + FlashBack
71.88
90.22
81.01
46.28
67.20
Table 7: Memory-related tasks on StreamingBench.
Figure 3: Routing behavior and historical reach of FlashBack.
Figure 4: Efficiency of FlashBack.
Figure 5: Representative FlashBack visualization cases on OVO-Bench.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Selection policy or diagnostic
MA
CRR
Downstream accuracy (%)
Native
69.60
41.67
Always Recall
70.40
60.83
Direct Routing
69.20
41.67
Structured Routing
72.80
60.83
Trajectory Oracle
78.80
65.42
Appendix
Table 8: Outcome-based trajectory-selection analysis. The correctness partition is derived from the same fixed Native and Recall outputs used by all selection policies.
Humans effortlessly perceive the present while remembering the past, yet streaming VLMs often trade off real-time perception against long-term memory. Prior work shows that shortening the context can sharpen current-scene perception at the expense of long-range recall. To reconcile these abilities, we introduce StreamTTT, which writes long-range history into online-updated fast weights outside the attention context. This leaves a short sliding key-value cache dedicated to recent evidence, mitigating attention dilution. We train StreamTTT jointly on offline long-video QA and a newly constructed real-time QA corpus. On OVO-Bench, under each model's reported input protocol, StreamTTT-4B outperforms the same-scale SimpleStream-4B by 0.6 points in real-time perception and 5.3 points in backward tracing. It also surpasses the larger SimpleStream-8B by 0.73 points on StreamingBench's Real-Time Visual Understanding (RTVU) subset. Our code is publicly available at https://github.com/zeyun-zhong/StreamTTT.
Joya Chen, Zeyun Zhong, Mike Zheng Shou
National University of Singapore · Karlsruhe Institute of Technology
Streaming video understanding requires multimodal large language models (MLLMs) to preserve relevant evidence from continuously evolving streams under strict causality and bounded memory. Yet existing paradigms remain limited: model-based methods require intrusive backbone updates, while memory-based methods expend substantial visual-encoding computation on temporally redundant content and rely on rigid access to visual history. To address these limitations, we introduce StreamFlow, an efficient visual memory framework that enables dynamic, on-demand access to historical visual information. StreamFlow combines a lightweight, dynamics-aware mid-term memory that filters temporal redundancy before visual encoding with a latent long-term memory that consolidates historical video content into visual latents accessible to subsequent reasoning. During generation, an attention-guided retrieval mechanism injects relevant visual latents when the model's reliance on visual evidence weakens. StreamFlow achieves state-of-the-art streaming video understanding performance, reaching 67.73% overall accuracy on StreamingBench, while also delivering strong performance on offline long-video benchmarks. Relative to the vanilla setting, it improves the visual attention score (VAS) by 59.1% while reducing end-to-end latency and peak memory by 50.4% and 21.1%, respectively, enabling more visually grounded and efficient reasoning.
Muxin Fu, Yifan Zhang, Wentao Zhang +5
Tongji University · Nanyang Technological University · University of Michigan +2
Streaming video understanding models must answer queries at any moment during an ongoing stream, using only what they have observed so far and under fixed memory and computation budgets. Existing methods address this by adding memory banks, retrieval modules, or visual token compression to preserve long-range history. However, strong recent-window baselines show that indiscriminate history injection can dilute current-scene perception, suggesting that the key challenge is not whether to use memory, but how to allocate it selectively. We formulate this as budgeted online latent evidence allocation and propose \textbf{SelectStream}, a selective latent-memory framework that keeps the current observation directly visible to a frozen VLM while exposing historical information only through a compact, query-conditioned evidence budget. Three coordinated mechanisms govern when to write, what to preserve, and how to retrieve: surprise-driven adaptive windowing, priority-preserving consolidation, and query-conditioned graph reasoning over a fixed-capacity latent memory graph. Retrieved evidence is calibrated and injected as latent tokens for answer generation, without replaying frames or growing the context with stream length. Experimental results show that SelectStream achieves strong online streaming performance and preserves general video understanding, reaching 82.67% on StreamingBench, 67.03% on OVO-Bench, and 74.4% average accuracy on offline video benchmarks, while outperforming strong recent-window baselines and prior streaming memory methods.
Haonan Ge, Yiwei Wang, Hang Wu +1
University of California, Santa Barbara · University of California, Merced · The University of Queensland