Despite their impressive performance on a wide range of video understanding tasks, current Vision Language Models (VLMs) are predominantly designed for offline scenarios and struggle to handle online streaming videos that demand low latency response. Several studies have explored memory and token compression strategies in an attempt to adapt offline VLMs for streaming video understanding tasks. However, through our probing experiment, we identify that most existing works tend to progressively lose long context information as length of input stream increases. To address this, we propose ReMem, a novel training-free adaptation technique that enables VLMs to process streaming videos of arbitrary lengths while improving their long context information retention capability. ReMem exploits memory from two perspectives, implemented as two core components. The Streaming Context Memory (SCM) continuously compresses historical context with query-independent attention. The Retrieved Vision Memory (RVM) then retrieves the most salient, query-relevant context from memory to augment the VLM's input. Comprehensive experiments demonstrate that the proposed ReMem achieves state-of-the-art (SOTA) performance across a variety of widely used benchmarks, spanning both streaming video and general long video understanding tasks.
Figures & tables
Figure 1: Comparison of visual haystacks retrieval performance among Qwen2.5VL base model ( Bai et al. 2025b ) , FluxMem ( Xie et al. 2026 ) , HERMES ( Zhang et al. 2026 ) , OASIS ( Liang et al. 2026 ) and our method. As the haystack size increases, ReMem maintains more stable performance in long visual sequences, indicating that crucial contextual information is better preserved.
Figure 2: Overview of the proposed ReMem framework. The video stream is segmented to video chunks continuously. At each time step, the VLM attends to visual and contextual tokens, selectively storing informative representations into the SCM. Upon receiving a specific question query, we score each memory slot by cosine similarity, select the top vision tokens from the RVM, and integrate them with the ongoing context, yielding a compact representation.
Model
Frames
StreamingBench Contextual Understanding
OVOBench Backward-Tracing
SQA
MCU
ACU
Avg
EPM
ASI
HLD
Avg
Streaming Video VLMs (training-based)
VideoLLM-online-8B
2 fps
30.8
29.2
24.2
28.0
22.2
18.8
12.2
17.7
FlashVStream-7B
1 fps
24.8
25.2
26.8
25.6
39.1
37.2
5.9
27.4
TimeChat-Online-7B
1 fps
41.6
31.6
39.2
37.5
55.9
59.5
9.7
41.7
Dispider
1 fps
34.8
27.7
39.6
34.0
48.5
55.4
4.3
36.1
Table 1: Benchmark comparison on StreamingBench and OVOBench online video evaluation dataset.
Model
Frames
EgoSchema
VideoMME
Subset †
Fullset
short
medium
long
Avg
Streaming Video VLMs (training-based)
VideoLLM-online-8B
2 fps
–
32.8
–
–
–
–
FlashVStream-7B
1 fps
38.40
68.2
72.0
61.1
50.3
61.2
TimeChat-Online-7B
1 fps
–
61.9
67.4
50.5
41.7
53.2
Dispider
1 fps
–
55.6
66.1
53.7
49.7
56.5
Table 2: Benchmark comparison on EgoSchema and VideoMME offline video evaluation dataset. † Full-set evaluation is unavailable due to evaluation server issue; results are reported on the EgoSchema subset.
Base Model
ReMem Components
StreamingBench
SCM
Pos. Emb Mgmt.
RVM
SQA
ACU
Qwen2.5-VL-7B
✗
✗
✗
36.9
34.8
Qwen2.5-VL-7B
✓
✗
✗
41.7
37.6
Qwen2.5-VL-7B
✗
✗
✓
37.2
40.5
Qwen2.5-VL-7B
✓
✓
✗
45.0
38.4
Qwen2.5-VL-7B
✓
✓
✓
45.6
44.8
Table 3: Ablation of memory components on StreamingBench.
Base Model
Method
Chunk Size
StreamingBench
SQA
ACU
Qwen2.5-VL-7B
Official
–
36.9
34.8
Qwen2.5-VL-7B
SCM
1
35.3
31.0
Qwen2.5-VL-7B
SCM
2
38.6
32.3
Qwen2.5-VL-7B
SCM
4
41.8
37.6
Qwen2.5-VL-7B
SCM
16
45.0
38.4
Table 4: Video input chunk size analysis for Streaming Context Memory (SCM)
Number of Clips
SQA
MCU
ACU
1
45.6
36.8
44.8
2
41.6
34.0
41.2
3
41.6
29.2
38.2
Table 5: Influence of the number of retrieved historical clips on overall performance within the RVM design.
Figure 3: Left: response latency of different VLM models under streaming evaluation. Right: Maximum GPU memory consumption of different VLM models under streaming evaluation.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Model
#Frames
OP
CR
CS
ATP
EU
TR
PR
SU
ACP
CT
Avg.
Streaming Video VLMs (training-based)
VideoLLM-online-8B
2 fps
39.1
40.1
34.5
31.1
46.0
32.4
31.5
34.2
42.5
27.9
36.0
FlashVStream-7B
1 fps
25.9
43.6
24.9
23.9
27.3
13.1
18.5
25.2
23.9
48.7
23.2
TimeChat-Online-7B
1 fps
80.2
82.0
79.5
83.3
76.1
78.5
78.7
64.6
69.6
58.0
75.4
Dispider
1 fps
74.9
75.5
74.1
73.1
74.4
59.9
76.1
62.9
62.2
45.8
67.6
Streaming Video VLMs (training-free)
Appendix
Table 6: Detailed performance comparison of VLMs across StreamingBench real-time evaluation. All results are accuracies in percent. The Overall average is computed over all ten tasks. For rows marked with † , each task reports the best result across the evaluated ReMem chunk sizes: {4,8,16,32} for Qwen2.5-VL-7B, {4,8,16} for LLaVA-OV-7B and {4,8} for Qwen3-VL-8B.
Model
#Frames
Real-Time Visual Perception
Backward Tracing
Overall
OCR
ACR
ATR
STU
FPD
OJR
Avg.
EPM
ASI
HLD
Avg.
Avg.
Streaming Video VLMs (training-based)
VideoLLM-online-8B
2 fps
8.1
23.9
12.1
14.0
45.5
21.2
20.8
22.2
18.8
12.2
17.7
19.3
FlashVStream-7B
1 fps
25.5
32.1
29.3
33.7
29.7
28.8
29.9
36.4
33.8
5.9
25.4
27.6
TimeChat-Online-7B
1 fps
75.2
46.8
70.7
47.8
69.3
61.4
61.9
55.9
59.5
9.7
41.7
51.8
Dispider
1 fps
57.7
49.5
62.1
44.9
61.4
51.6
54.6
48.5
55.4
4.3
36.1
45.3
Appendix
Table 7: Detailed performance comparison of VLMs across OVOBench Real-Time Visual Perception and Backward Tracing tasks. All results are accuracies in percent. The Real-Time and Backward averages are computed over six and three tasks, respectively, while the Overall average is computed over all nine tasks. For rows marked with † , each task reports the best result across the evaluated ReMem chunk sizes: {4,8,16,32} for Qwen2.5-VL-7B, {4,8,16} for LLaVA-OV-7B and {4,8} for Qwen3-VL-8B.
Table 11
Figure 4: Distribution of attention score over visual tokens. Top: general query. Bottom: task query. The task query is “How many game scores has the team referred to in the previous question scored so far?” from StreamingBench SQA, and the general query is “Describe the video in detail sequentially from start to finish.”
Figure 5: Case study on Misleading Context Understanding . The video provides a detailed introduction to the rules and scoring of the card game Bridge, including the suit ranking: Clubs, Diamonds, Hearts, Spades, and No Trump (NT). At timestamp 00:01:52 the question asks: “What exactly is on the table now?” with four multiple choice options describing different symbol configurations. The ground truth answer is C , while the Qwen2.5-VL baseline is distracted by earlier frames and incorrectly predicts D . In contrast, our ReMem-enhanced model correctly outputs C , highlighting its improved ability to track evolving object states over time and resist misleading intermediate context.
Figure 6: Case study on Misleading Context Understanding . The video explains the rules and scoring of the card game Bridge, and this example shows several snapshots of the scorecard as the match progresses, ending with the final tally sheet on the tabletop. At timestamp 00:05:46, the question asks: “What’s inside the white curtain on the tabletop?” with four multiple choice options describing different numerical layouts. The ground truth answer is B , while the Qwen2.5-VL baseline is distracted by earlier frames and incorrectly predicts C . In contrast, our ReMem-enhanced model correctly outputs B , highlighting its improved ability to track evolving object states over time and resist misleading intermediate context.
Figure 7: Case study on Misleading Context Understanding . This example shows a sequence of frames from a casino table game where face-up cards are gradually revealed and rearranged. At timestamp 00:07:50, the question asks: “How many face-up playing cards are here?” with four multiple choice options describing different cards arrangement. The ground truth answer is C , while the Qwen2.5-VL baseline is distracted by earlier frames and incorrectly predicts A . In contrast, our ReMem-enhanced model correctly outputs C , highlighting its improved ability to track evolving object states over time and resist misleading intermediate context.
Figure 8: Case study on Anomaly Context Understanding . The video shows a stage magic performance where a magician covers a chair with a black cloth while a young boy stands beside it. At timestamp 00:05:00, the question asks “What unusual event just occurred?” , with four multiple choice options describing different outcomes after the cloth is removed. The ground truth answer is C , while the Qwen2.5-VL baseline incorrectly predicts D , hallucinating that a boy suddenly appears on the chair. In contrast, our ReMem-enhanced model correctly selects C , indicating improved sensitivity to fine-grained, unexpected events and more precise reasoning in dynamic, unpredictable scenes.
Figure 9: Case study on Anomaly Context Understanding . This video presents a dashcam-style traffic video in which rare or unexpected events must be localized and identified. At timestamp 00:01:08, the question asks: “What unusual event just occurred?” , with four multi choice descriptions of possible anomalies. The ground truth answer is D , while the Qwen2.5-VL baseline incorrectly predicts B , being distracted by large moving vehicles in earlier frames. In contrast, our ReMem-enhanced model correctly identifies D , demonstrating stronger temporal grounding and improved resilience to misleading context in fast-changing real-world scenes.
Figure 10: Case study on Anomaly Context Understanding . The video records a stage magic performance in which a magician interacts with guests at a table. At timestamp 00:03:01, the question asks: “What unusual event just occurred?” with four multiple choice options describing possible outcomes. The ground-truth answer is A , while the Qwen2.5-VL baseline incorrectly predicts B , focusing on ripping the signed card, a salient but non-anomalous action. In contrast, our ReMem-enhanced model correctly selects A . This example demonstrates that ReMem better captures temporal dependencies and distinguishes genuinely unexpected events from normal but visually prominent actions, leading to more reliable anomaly reasoning in streaming video.
Figure 11: Case study on Sequential Question Answering . The video records a professional table-tennis match, where the model must answer five temporally linked questions based on the same ongoing rally. The first question at timestamp 00:00:36 asks: “Who is preparing to serve the ball now?” , and all models correctly identify Wang Chuqin (option A). The second question at timestamp 00:03:02 then queries: “How much game score has the person mentioned in the last question scored now?” with four multiple choice options describing different scores. The ground truth answer is C ,while the Qwen2.5-VL baseline fails to propagate the correct entity and its evolving score,incorrectly predicting B . In contrast, our ReMem-enhanced model maintains consistent temporal grounding and accurately outputs C . This example illustrates that ReMem substantially improves long-range entity tracking and state updating in streaming video question answering.
Figure 12: Case study on Sequential Question Answering (SQA) . The video shows a professional basketball game, where the model is queried with five temporally dependent questions about the same ongoing play. The first question at timestamp 00:00:25 asks: “Which team is currently on offense in the game?” , and all models correctly identify LAC (option A). The fourth question at timestamp 00:06:09, asks: “How many game scores does the team referred to in the first question have currently?” . The fifth question at timestamp 00:08:04 further asks: “How many game scores has the team that the third question specifically refers to scored now?” , both under a multi-choice setting. The ground-truth answers are D (85 points) and D (115 points), respectively. The Qwen2.5-VL baseline fails to correctly update and track the scores, predicting A and B , whereas our ReMem-enhanced model consistently outputs the correct options D and D . This example highlights that ReMem more effectively leverages episodic memory to maintain entity references and their evolving states over time, enabling coherent and contextually accurate reasoning across a sequence of related questions in streaming video.
Streaming video understanding requires multimodal large language models (MLLMs) to preserve relevant evidence from continuously evolving streams under strict causality and bounded memory. Yet existing paradigms remain limited: model-based methods require intrusive backbone updates, while memory-based methods expend substantial visual-encoding computation on temporally redundant content and rely on rigid access to visual history. To address these limitations, we introduce StreamFlow, an efficient visual memory framework that enables dynamic, on-demand access to historical visual information. StreamFlow combines a lightweight, dynamics-aware mid-term memory that filters temporal redundancy before visual encoding with a latent long-term memory that consolidates historical video content into visual latents accessible to subsequent reasoning. During generation, an attention-guided retrieval mechanism injects relevant visual latents when the model's reliance on visual evidence weakens. StreamFlow achieves state-of-the-art streaming video understanding performance, reaching 67.73% overall accuracy on StreamingBench, while also delivering strong performance on offline long-video benchmarks. Relative to the vanilla setting, it improves the visual attention score (VAS) by 59.1% while reducing end-to-end latency and peak memory by 50.4% and 21.1%, respectively, enabling more visually grounded and efficient reasoning.
Muxin Fu, Yifan Zhang, Wentao Zhang +5
Tongji University · Nanyang Technological University · University of Michigan +2
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual inputs and respond to user queries under strict causality and bounded memory. Existing approaches typically compress historical observations into an external memory bank and retrieve query-relevant evidence as additional visual context. Though effective, this store-and-retrieve paradigm keeps historical evidence as external visual context, preventing it from being internalized into a compact, evolving latent memory that can continuously guide streaming reasoning. To bridge this gap, we introduce LatentStream, a progressive latent working memory framework that shifts streaming memory from store-and-retrieve to retrieve-and-internalize. Specifically, LatentStream comprises three coordinated components. First, Query-agnostic Hierarchical Streaming Memory organizes visual history into short-, mid-, and long-term levels under a fixed memory budget through Jenks-guided adaptive consolidation. Once a query arrives, Hierarchical Latent Memory Evolution equips groups of latent memory tokens with progressively expanding memory receptive fields, enabling them to iteratively retrieve historical evidence from their corresponding scopes and internalize it into a compact, fixed-length latent memory. Finally, Progressive Confidence-guided Latent Memory Optimization constructs a hierarchical progression reward from group-wise predictive entropy and jointly refines the latent memory tokens and retrieved evidence, encouraging increasingly confident streaming reasoning. Extensive experiments demonstrate that LatentStream achieves new state-of-the-art results on existing online and offline video benchmarks.
Hongyu Qu, Guangming Yao, Ling Xing +7
Nanjing University of Science and Technology · Ant Group · National University of Singapore +1
Streaming video understanding requires models to continuously retain useful visual evidence before future questions are known. Existing approaches primarily manage the growing visual context according to token importance, temporal redundancy, or segment-level relevance, but rarely organize evidence around objects that persist and evolve over time. Thus, in this paper, we introduce ObjectStream, a training-free framework that treats latent objects as memory anchors for streaming video understanding. ObjectStream induces spatially coherent latent objects directly from frozen Video-LLM representations, links them across frames into persistent anchors, and maintains their histories under a bounded memory budget, without requiring external object detectors or segmentation models. Built on these anchors, ObjectStream preserves three complementary forms of evidence: persistent object histories, transient object changes, and recent visual context. This design enables existing Video Large Language Models (Video-LLMs) to reason over object identities, interactions, and state changes while leaving the underlying model unchanged. Extensive experiments on online streaming and offline long-video benchmarks demonstrate both effectiveness and efficiency. In online streaming evaluation, ObjectStream improves Qwen2.5-VL-7B by 10.0 points on OVO-Bench Real-Time Visual Perception, while reducing peak GPU mem-ory and TTFT by approximately 50%. On offline long-video benchmarks, it surpasses the full-token baseline while discarding 82.5% of visual tokens. These results highlight latent objects as a practical and effective organizing principle for compact streaming video memory.