Streaming video understanding requires models to process unbounded visual streams while preserving rich visual semantics across vast temporal horizons, posing a fundamental challenge for memory modeling. Existing approaches primarily focus on increasing memory capacity, either by compressing historical information into fixed-size representations or by extending storage beyond GPU memory. However, these methods largely rely on global or coarse-grained representations, inevitably losing fine-grained visual information. In this work, we argue that streaming video memory should explicitly encode structured and semantically meaningful representations, particularly at the entity level. To this end, we propose MEMO, a novel framework that models streaming video through multi-level, entity-aware structured memory. MEMO performs multi-level perception to jointly capture global semantics, entity dynamics, and spatial structures, partitioning streaming video into semantically coherent chunks. Each chunk is organized into a structured memory, where lightweight global and entity-level representations serve as retrieval indices, while the corresponding high-resolution visual content is retained separately for on-demand access. At inference time, MEMO performs query-specific retrieval over the structured memory and selectively recalls relevant visual evidence for downstream reasoning. Notably, MEMO is training-free and plug-and-play with existing multimodal large language models. Extensive experiments on StreamingBench and OVO-Bench demonstrate that MEMO consistently improves multiple base models and achieves state-of-the-art performance.
Figures & tables
Figure 1 . Comparison of memory paradigms for streaming video understanding. (a) Internal memory compression is efficient but loses fine-grained details. (b) External memory extension retains longer history but relies on coarse-grained evidence. (c) MEMO decouples lightweight structured indices from high-resolution visual evidence. Comparison of internal memory compression, external memory extension, and MEMO's structured entity-aware memory.
Figure 2 . Overall architecture of MEMO. The streaming pipeline consists of four stages: (1) Multi-Level Entity-Aware Perception extracts global, entity-level, and spatial cues to model semantic continuity; (2) Online Temporal Chunking incrementally partitions the video stream into semantically coherent chunks based on similarity; (3) Structured Memory Construction organizes chunks into a structured memory with lightweight indices and high-resolution visual evidence; and (4) Query-Specific Evidence Retrieval retrieves the most relevant chunks and their associated visual evidence to support downstream reasoning. A left-to-right overview of the MEMO pipeline begins with a continuous video stream. Multi-level perception combines whole-frame CLIP similarity, matched entity features updated by exponential moving average, and spatial cues based on mask overlap and object displacement into a total similarity score. Online temporal chunking compares this score with an adaptive threshold to partition the stream into semantic chunks. For each chunk, high-resolution visual evidence is stored in a CPU pool, while lightweight global and entity indices remain on the GPU. At query time, the text representation retrieves the top-K historical chunks. Their visual evidence is combined with frames from the current working memory, encoded, and passed with the query to the language model to generate an answer.
Method
Frames
OVO-Bench real-time
StreamingBench real-time
OCR
ACR
ATR
STU
FPD
OJR
Avg.
OP
CR
CS
ATP
EU
TR
PR
SU
ACP
CT
Avg.
Proprietary Models
Gemini 1.5 Pro ( Team et al., 2024 )
1 fps
85.9
67.0
79.3
58.4
63.4
62.0
69.3
79.0
80.5
83.5
79.7
80.0
84.7
77.8
64.2
72.0
48.7
75.7
GPT-4o ( Hurst et al., 2024 )
64
69.8
64.2
71.6
51.1
70.3
59.8
64.5
77.1
80.5
83.9
76.5
70.2
83.8
66.7
62.2
69.1
49.2
73.3
Open-source Offline MLLMs
LongVA ( Zhang et al., 2024a )
128
–
–
–
–
–
–
–
70.0
63.3
61.2
70.9
62.7
59.5
61.1
53.7
54.7
34.7
60.0
Table 1 . Performance comparison on OVO-Bench and StreamingBench.
Figure 3 . Comparison of different methods in terms of peak GPU memory usage, response latency, and accuracy. All experiments are conducted on a single NVIDIA A6000 GPU. Three bar charts compare four systems in terms of peak GPU memory, response latency, and StreamingBench accuracy: TimeChat-Online-7B, LLaVA-OneVision-7B with ReKV, Qwen2.5-VL-7B with MEMO, and LLaVA-OneVision-7B with MEMO. Qwen2.5-VL-7B with MEMO achieves the highest accuracy of 78.5 percent while using 21.1 GB of GPU memory and 1.2 seconds of latency. LLaVA-OneVision-7B with MEMO uses 20.5 GB and has the lowest latency of 0.8 seconds while reaching 73.2 percent accuracy. TimeChat-Online has substantially higher latency, whereas ReKV uses the most GPU memory and has the lowest accuracy among the four systems.
Base
Perception
Memory
Acc. (%)
✓
✗
✗
73.2
✓
✗
✓
79.4 +6.2
✓
✓
✗
77.8 +4.6
✓
✓
✓
83.7 +10.5
Table 2 . Ablation on perception and memory synergy.
Vision / Frame
Memory
LLM / Query
Module
ms
Operation
ms
Metric
ms
CLIP
16.56
Update
0.39
TTFT
598.44
Grounding DINO
154.91
Ret. & Recall
8.08
TPOT
55.02
SAM
126.13
Total Gen.
1367.62
Table 3 . Latency breakdown of Qwen2.5-VL-7B + MEMO. Ret. and Gen. denote retrieval and generation, respectively.
Variants
Feature Dependency
Acc. (%)
Chunking Stage
Retrieval Stage
Full Model
Sspatial+Slocal+Sglobal
Mentity+Mglobal
83.69
Entity-only
Sspatial+Slocal
Mentity
81.77
Global-only
Sglobal
Mglobal
81.31
Table 4. Ablation on multi-level memory. “Entity-only” relies exclusively on spatial and local cues for chunking and Mentity for retrieval. “Global-only” relies exclusively on global cues for chunking and Mglobal for retrieval.
Figure 4 . Ablation results for different retrieval strategies (left) and temporal chunking strategies (right). Two bar charts report StreamingBench accuracy for retrieval and temporal chunking ablations. In the left chart, the full semantic retrieval strategy achieves the highest accuracy of approximately 83.7 percent. Nearest-K retrieval ranks second at 82.32 percent, while the local-only, global-only, and random-K alternatives obtain lower scores. In the right chart, adaptive semantic chunking achieves the highest accuracy. Fixed-length windows of 16 and 32 frames perform better than windows of 8 and 64 frames, but all fixed-length variants remain below the adaptive strategy. The 8-frame and 64-frame settings achieve 81.77 and 81.82 percent, respectively.