Streaming vision-language models must process continuously growing video streams under a bounded compute budget, creating a persistent tension between real-time perception and long-term memory. Retrieving historical information provides a natural remedy, yet historical recall is not uniformly beneficial: unnecessary history may introduce irrelevant context into current reasoning and interfere with native real-time perception. Effective streaming memory should therefore address not only what to remember, but also when and how to access it. To this end, we introduce FlashBack, a training-free framework for selective, multi-level memory in streaming vision-language models. Before retrieving history, FlashBack draws on the semantic understanding of the frozen streaming VLM to infer whether a query calls for historical evidence. This assessment determines whether inference remains on the Native trajectory or invokes an isolated Recall trajectory. The Recall trajectory combines recent context with retrieved long-term memory through a query-local Side-KV pathway, preserving local temporal continuity without modifying the persistent Native state. We instantiate FlashBack on StreamingVLM and Mage-VL-4B and evaluate it on OVO-Bench and StreamingBench. The results show improvements on several long-horizon and memory-dependent tasks while largely preserving real-time perception, with performance competitive with strong training-based streaming methods despite requiring no additional training. Our code will be announced later.
Figures & tables
Figure 1: Motivation for FlashBack.
Figure 2: Overview of FlashBack.
Method
MA
CRR
Native
69.60
41.67
+ Recall
70.40
60.83
+ Direct Routing
69.20
41.67
+ Structured Routing
72.80
60.83
Table 1: Effect of Recall and Routing.
Memory
MA
CRR
Recent only ( K=0 )
73.60
55.42
K=1
72.80
60.83
K=3
72.80
60.83
K=5
72.40
60.83
Table 2: Effect of Memory Hierarchy.
Integration
MA
CRR
Random KV
33.20
39.58
Zero KV
26.80
39.17
Historical KV
64.40
45.83
Side-KV
70.40
60.83
Table 3: Effect of Side-KV.
Method
RT Avg.
Bwd. Avg.
Fwd. Avg.
Training-based methods
VideoLLM-online
20.79
17.73
–
EventMemAgent-8B
68.29
58.03
55.92
SelectStream-Qwen3-8B
82.76
62.20
56.13
StreamReady-7B
73.60
72.20
58.80
Training-free methods
Table 4: OVO-Bench group-level results.
Method
EPM
ASI
REC
CRR
HLD
StreamingVLM
54.88
56.76
19.34
41.67
27.96
StreamingVLM + FlashBack
54.88
56.76
20.06
60.83
23.12
Mage-VL
52.53
58.78
19.48
44.17
32.26
Mage-VL + FlashBack
53.20
64.86
29.80
54.17
31.72
Table 5: Memory-related tasks on OVO-Bench.
Method
RT Avg.
Omni Avg.
Context Avg.
Training-based methods
VideoLLM-online
35.99
28.45
–
EventMemAgent-8B
77.00
–
–
SelectStream-Qwen3-8B
82.67
–
–
Training-free methods
StreamRAG-ViSpeak-7B
78.12
–
–
Table 6: StreamingBench group-level results.
Method
CR
CS
EU
CT
ACU
StreamingVLM
79.69
87.70
80.38
27.66
52.40
StreamingVLM + FlashBack
81.25
88.33
80.38
30.32
53.60
Mage-VL
67.97
90.22
76.58
40.43
67.20
Mage-VL + FlashBack
71.88
90.22
81.01
46.28
67.20
Table 7: Memory-related tasks on StreamingBench.
Figure 3: Routing behavior and historical reach of FlashBack.
Figure 4: Efficiency of FlashBack.
Figure 5: Representative FlashBack visualization cases on OVO-Bench.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Selection policy or diagnostic
MA
CRR
Downstream accuracy (%)
Native
69.60
41.67
Always Recall
70.40
60.83
Direct Routing
69.20
41.67
Structured Routing
72.80
60.83
Trajectory Oracle
78.80
65.42
Appendix
Table 8: Outcome-based trajectory-selection analysis. The correctness partition is derived from the same fixed Native and Recall outputs used by all selection policies.