Streaming video understanding requires models to process unbounded visual streams while preserving rich visual semantics across vast temporal horizons, posing a fundamental challenge for memory modeling. Existing approaches primarily focus on increasing memory capacity, either by compressing historical information into fixed-size representations or by extending storage beyond GPU memory. However, these methods largely rely on global or coarse-grained representations, inevitably losing fine-grained visual information. In this work, we argue that streaming video memory should explicitly encode structured and semantically meaningful representations, particularly at the entity level. To this end, we propose MEMO, a novel framework that models streaming video through multi-level, entity-aware structured memory. MEMO performs multi-level perception to jointly capture global semantics, entity dynamics, and spatial structures, partitioning streaming video into semantically coherent chunks. Each chunk is organized into a structured memory, where lightweight global and entity-level representations serve as retrieval indices, while the corresponding high-resolution visual content is retained separately for on-demand access. At inference time, MEMO performs query-specific retrieval over the structured memory and selectively recalls relevant visual evidence for downstream reasoning. Notably, MEMO is training-free and plug-and-play with existing multimodal large language models. Extensive experiments on StreamingBench and OVO-Bench demonstrate that MEMO consistently improves multiple base models and achieves state-of-the-art performance.
Figures & tables
Figure 1 . Comparison of memory paradigms for streaming video understanding. (a) Internal memory compression is efficient but loses fine-grained details. (b) External memory extension retains longer history but relies on coarse-grained evidence. (c) MEMO decouples lightweight structured indices from high-resolution visual evidence. Comparison of internal memory compression, external memory extension, and MEMO's structured entity-aware memory.
Figure 2 . Overall architecture of MEMO. The streaming pipeline consists of four stages: (1) Multi-Level Entity-Aware Perception extracts global, entity-level, and spatial cues to model semantic continuity; (2) Online Temporal Chunking incrementally partitions the video stream into semantically coherent chunks based on similarity; (3) Structured Memory Construction organizes chunks into a structured memory with lightweight indices and high-resolution visual evidence; and (4) Query-Specific Evidence Retrieval retrieves the most relevant chunks and their associated visual evidence to support downstream reasoning. A left-to-right overview of the MEMO pipeline begins with a continuous video stream. Multi-level perception combines whole-frame CLIP similarity, matched entity features updated by exponential moving average, and spatial cues based on mask overlap and object displacement into a total similarity score. Online temporal chunking compares this score with an adaptive threshold to partition the stream into semantic chunks. For each chunk, high-resolution visual evidence is stored in a CPU pool, while lightweight global and entity indices remain on the GPU. At query time, the text representation retrieves the top-K historical chunks. Their visual evidence is combined with frames from the current working memory, encoded, and passed with the query to the language model to generate an answer.
Method
Frames
OVO-Bench real-time
StreamingBench real-time
OCR
ACR
ATR
STU
FPD
OJR
Avg.
OP
CR
CS
ATP
EU
TR
PR
SU
ACP
CT
Avg.
Proprietary Models
Gemini 1.5 Pro ( Team et al., 2024 )
1 fps
85.9
67.0
79.3
58.4
63.4
62.0
69.3
79.0
80.5
83.5
79.7
80.0
84.7
77.8
64.2
72.0
48.7
75.7
GPT-4o ( Hurst et al., 2024 )
64
69.8
64.2
71.6
51.1
70.3
59.8
64.5
77.1
80.5
83.9
76.5
70.2
83.8
66.7
62.2
69.1
49.2
73.3
Open-source Offline MLLMs
LongVA ( Zhang et al., 2024a )
128
–
–
–
–
–
–
–
70.0
63.3
61.2
70.9
62.7
59.5
61.1
53.7
54.7
34.7
60.0
Table 1 . Performance comparison on OVO-Bench and StreamingBench.
Figure 3 . Comparison of different methods in terms of peak GPU memory usage, response latency, and accuracy. All experiments are conducted on a single NVIDIA A6000 GPU. Three bar charts compare four systems in terms of peak GPU memory, response latency, and StreamingBench accuracy: TimeChat-Online-7B, LLaVA-OneVision-7B with ReKV, Qwen2.5-VL-7B with MEMO, and LLaVA-OneVision-7B with MEMO. Qwen2.5-VL-7B with MEMO achieves the highest accuracy of 78.5 percent while using 21.1 GB of GPU memory and 1.2 seconds of latency. LLaVA-OneVision-7B with MEMO uses 20.5 GB and has the lowest latency of 0.8 seconds while reaching 73.2 percent accuracy. TimeChat-Online has substantially higher latency, whereas ReKV uses the most GPU memory and has the lowest accuracy among the four systems.
Base
Perception
Memory
Acc. (%)
✓
✗
✗
73.2
✓
✗
✓
79.4 +6.2
✓
✓
✗
77.8 +4.6
✓
✓
✓
83.7 +10.5
Table 2 . Ablation on perception and memory synergy.
Vision / Frame
Memory
LLM / Query
Module
ms
Operation
ms
Metric
ms
CLIP
16.56
Update
0.39
TTFT
598.44
Grounding DINO
154.91
Ret. & Recall
8.08
TPOT
55.02
SAM
126.13
Total Gen.
1367.62
Table 3 . Latency breakdown of Qwen2.5-VL-7B + MEMO. Ret. and Gen. denote retrieval and generation, respectively.
Variants
Feature Dependency
Acc. (%)
Chunking Stage
Retrieval Stage
Full Model
Sspatial+Slocal+Sglobal
Mentity+Mglobal
83.69
Entity-only
Sspatial+Slocal
Mentity
81.77
Global-only
Sglobal
Mglobal
81.31
Table 4. Ablation on multi-level memory. “Entity-only” relies exclusively on spatial and local cues for chunking and Mentity for retrieval. “Global-only” relies exclusively on global cues for chunking and Mglobal for retrieval.
Figure 4 . Ablation results for different retrieval strategies (left) and temporal chunking strategies (right). Two bar charts report StreamingBench accuracy for retrieval and temporal chunking ablations. In the left chart, the full semantic retrieval strategy achieves the highest accuracy of approximately 83.7 percent. Nearest-K retrieval ranks second at 82.32 percent, while the local-only, global-only, and random-K alternatives obtain lower scores. In the right chart, adaptive semantic chunking achieves the highest accuracy. Fixed-length windows of 16 and 32 frames perform better than windows of 8 and 64 frames, but all fixed-length variants remain below the adaptive strategy. The 8-frame and 64-frame settings achieve 81.77 and 81.82 percent, respectively.
Streaming video understanding requires multimodal large language models (MLLMs) to preserve relevant evidence from continuously evolving streams under strict causality and bounded memory. Yet existing paradigms remain limited: model-based methods require intrusive backbone updates, while memory-based methods expend substantial visual-encoding computation on temporally redundant content and rely on rigid access to visual history. To address these limitations, we introduce StreamFlow, an efficient visual memory framework that enables dynamic, on-demand access to historical visual information. StreamFlow combines a lightweight, dynamics-aware mid-term memory that filters temporal redundancy before visual encoding with a latent long-term memory that consolidates historical video content into visual latents accessible to subsequent reasoning. During generation, an attention-guided retrieval mechanism injects relevant visual latents when the model's reliance on visual evidence weakens. StreamFlow achieves state-of-the-art streaming video understanding performance, reaching 67.73% overall accuracy on StreamingBench, while also delivering strong performance on offline long-video benchmarks. Relative to the vanilla setting, it improves the visual attention score (VAS) by 59.1% while reducing end-to-end latency and peak memory by 50.4% and 21.1%, respectively, enabling more visually grounded and efficient reasoning.
Muxin Fu, Yifan Zhang, Wentao Zhang +5
Tongji University · Nanyang Technological University · University of Michigan +2
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual inputs and respond to user queries under strict causality and bounded memory. Existing approaches typically compress historical observations into an external memory bank and retrieve query-relevant evidence as additional visual context. Though effective, this store-and-retrieve paradigm keeps historical evidence as external visual context, preventing it from being internalized into a compact, evolving latent memory that can continuously guide streaming reasoning. To bridge this gap, we introduce LatentStream, a progressive latent working memory framework that shifts streaming memory from store-and-retrieve to retrieve-and-internalize. Specifically, LatentStream comprises three coordinated components. First, Query-agnostic Hierarchical Streaming Memory organizes visual history into short-, mid-, and long-term levels under a fixed memory budget through Jenks-guided adaptive consolidation. Once a query arrives, Hierarchical Latent Memory Evolution equips groups of latent memory tokens with progressively expanding memory receptive fields, enabling them to iteratively retrieve historical evidence from their corresponding scopes and internalize it into a compact, fixed-length latent memory. Finally, Progressive Confidence-guided Latent Memory Optimization constructs a hierarchical progression reward from group-wise predictive entropy and jointly refines the latent memory tokens and retrieved evidence, encouraging increasingly confident streaming reasoning. Extensive experiments demonstrate that LatentStream achieves new state-of-the-art results on existing online and offline video benchmarks.
Hongyu Qu, Guangming Yao, Ling Xing +7
Nanjing University of Science and Technology · Ant Group · National University of Singapore +1
In online streaming video understanding, a video stream continues to arrive and queries may be issued at any time. Because streaming frames grow without bound, the system must continuously compress and retain information from the observed video prefix while future frames and future queries remain unknown. The core challenge is deciding what information to retain and how to organize the maintained history: as this history grows with the stream, memory cost increases and many redundant visual details are retained, whereas later queries often depend on specific entities, actions, and their temporal changes. To address this challenge, we introduce FOLIO, a training-free focused semantic memory system that records important parts of the stream in higher detail while keeping surrounding context compact. As the stream arrives, FOLIO updates memory at the segment level, guided by a dynamic focus state, combining a short-term visual buffer with a long-term semantic memory organized around observed entities and linked to a visual-evidence cache. At query time, lightweight hybrid retrieval combines direct matching over the structured memory with semantic query expansion. FOLIO achieves state-of-the-art performance, reaching 82.0/69.1 Perception/Backward accuracy on OVO-Bench with Qwen3-VL-8B and 74.5 overall accuracy on StreamingBench, while substantially reducing the cost of maintaining streaming memory by reserving detailed records for focused entities and storing surrounding context compactly.