Real-time video understanding requires incrementally maintaining a memory of streaming content, and optimizing this requires dense process signals. On-Policy Self-Distillation (OPSD), which lets one model serve as both teacher and student with the teacher receiving additional privileged information such as the question and ground-truth (GT) answer, can supply such token-level signals. However, applying it directly to streaming video raises two problems. (1) The student cannot be optimized end-to-end, where memory is written before the question arrives, yet the teacher scores it with the question-and-GT privilege, misaligning their preferences. (2) Effective-entity memory collapses, where the question-and-GT privilege makes the teacher favor only question-relevant entities, and token-mean averaging over a memory renders its signal invariant to how many entities that memory covers, both driving memory against the streaming need for diversity. To address these issues, we propose Event-Grounded Self-Distillation (EGSD), which characterizes streaming memory as an incremental update over verifiable Events (key visual entities, actions, and details) and targets the two problems on this basis. For problem (1), we adapt the OPSD signal into a multiplicative weight combined with the outcome reward; for problem (2), we re-weight the teacher with Events as privileged information to counter its question-relevance bias, and add an entity-coverage reward to supply the coverage preference the token-mean teacher lacks. Extensive experiments on mainstream online and offline benchmarks show EGSD achieves strong performance, reaching 79.8% on StreamingBench and 73.4% on the OVO-Bench Real-Time track, while memory analysis shows effective-entity recall rises 17.4% at only 6.8% more memory length.
Figures & tables
Figure 1: Applying OPSD/RLSD to streaming video collapses memory diversity.
Figure 2: Overview of EGSD. The student writes free-text memory per clip without seeing the question, while a frozen large model extracts a per-clip Event fact set offline. The Event-privileged teacher re-weights student tokens into a multiplicative advantage, and an Event-based entity-coverage reward prefers higher-coverage memories, guiding memory toward grounded diversity.
Method
Size
StreamingBench
OVO Real-Time
OP
CR
CS
ATP
EU
TR
PR
SU
ACP
CT
Avg.
OCR
ACR
ATR
STU
FPD
OJR
Avg.
Open-source Offline models
LLaVA-Video [TMLR’25]
7B
–
–
–
–
–
–
–
–
–
–
–
69.1
58.7
68.8
49.4
74.3
59.8
63.4
LLaVA-OV [TMLR’25]
7B
80.4
74.2
76.0
80.7
72.7
71.7
67.6
65.5
65.7
45.1
71.1
66.4
57.8
73.3
53.4
71.3
62.0
64.0
LongVU [ICML’25]
7B
–
–
–
–
–
–
–
–
–
–
–
53.7
53.2
62.9
47.8
68.3
59.8
57.6
LongVA [TMLR’25]
7B
70.0
63.3
61.2
70.9
62.7
59.5
61.1
53.7
54.7
34.7
60.0
–
–
–
–
–
–
–
Table 1: Results on the real-time subtasks of OVO-Bench and StreamingBench. Best overall results are in bold and the best results among training-free methods are underlined.
Method
Size
Online Video
Offline Video
OVO-Bench Backward
OVO-Bench Forward
VideoMME
LVB
VH
EPM
ASI
HLD
Avg.
REC
SSR
CRR
Avg.
Long
Overall
Open-source Offline models
LLaVA-Video [TMLR’25]
7B
56.2
57.4
7.5
40.4
34.1
70.0
60.4
54.8
–
63.3
61.3
–
LLaVA-OV [TMLR’25]
7B
54.2
55.4
21.5
43.7
25.6
67.1
58.8
50.5
–
58.2
56.3
–
LongVA [TMLR’25]
7B
–
–
–
–
–
–
–
–
47.6
54.3
56.3
–
Table 2: Memory-dependent results on the online video understanding benchmarks OVO-Bench Backward and Forward, as well as on the offline benchmarks VideoMME (w/o sub.), LongVideoBench (LVB), and VideoHolmes (VH).
Configuration
OVO-Bench
Video-MME
VideoHolmes
BT
FW
Avg.
Vanilla RLSD
50.8
55.7
53.3
62.6
44.3
+ remove decay rate
52.1
55.3
53.7
63.8
44.9
+ privileged Event
55.3
56.1
55.7
65.6
46.5
+ entity-coverage (EGSD)
56.2
56.7
56.5
66.5
46.2
Table 3: Ablation starting from Vanilla RLSD, enabling components one by one. Backbone Qwen3-VL-8B.
Figure 3: Per-step entity recall (a) and memory character count (b) for four methods sharing the same SFT start.
Figure 4: Case study of per-token teacher guidance on the same clip. (a) EGSD teacher with Event facts; (b) Vanilla RLSD teacher with only the question and ground-truth answer.
Step
10
30
50
70
90
Vanilla RLSD
0.204
0.280
0.237
0.288
0.300
EGSD (ours)
0.366
0.325
0.306
0.380
0.356
Table 4: Reward composition: visual-entity ratio of rewarded tokens across training steps.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Type
Method
Metric
Latency (s)
Offline
Qwen2.5-VL-7B w/CoT
E2E
5.30
Offline
Video-R1 w/CoT
E2E
8.80
Offline
Qwen2.5-VL-7B w/CoT (ours repro)
E2E
5.68
Offline
Qwen3-VL-8B w/CoT (our 8B backbone)
E2E
5.89
Offline
Qwen2.5-VL-7B direct-answer
TTFT
0.54
Online
Dispider-7B
TTFT
1.10
Appendix
Table 5: Inference latency on VideoHolmes (p50), reported for both the Qwen2.5-VL-7B and Qwen3-VL-8B backbones. Offline with-CoT rows report end-to-end QA latency (E2E); direct-answer and online rows report TTFT. Rows shaded green are offline, orange online; numbers not marked “ours” are quoted from the VST paper.
Streaming video understanding requires answering questions that arrive at arbitrary moments over an unbounded video stream. Existing systems primarily focus on what to retain in a bounded memory, yet access that memory using the same fixed-cost procedure for every query, despite substantial variation in the evidence required. We argue that deciding how deeply to access memory for each query is as important as deciding what the memory should store. To this end, we introduce StreamScout, an adaptive inference framework that maintains only a lightweight textual timeline in context as the stream unfolds. At query time, StreamScout progressively augments the timeline with up to three increasingly informative visual views: a glance at recent frames, a uniform look-back over the past stream, and query-salient retrieval. At each stage, the model answers immediately if the available evidence is sufficient; otherwise, it escalates to the next view. To improve this stop-or-escalate policy, we probe the cascade on an auxiliary set and distill the model's empirical competence boundary into supervision for a lightweight LoRA adaptation, yielding StreamScout-S. We further refine the policy through reinforcement learning, allowing the model to explore stopping behaviors beyond imitation of the distilled decisions, yielding StreamScout-R. Across three backbones and three streaming benchmarks, StreamScout and its variants consistently outperform prior streaming methods while substantially reducing inference cost and token consumption; on OVO-Bench, for instance, StreamScout-S improves Qwen3-VL-8B by 14.65 points while using 59% fewer tokens than uniform sampling and answering in 1.04 s on average.
Ce Zhang, Jing Bi, Jinxi He +9
Carnegie Mellon University · TikTok · University of Rochester +2
Streaming video understanding requires models to process unbounded visual streams while preserving rich visual semantics across vast temporal horizons, posing a fundamental challenge for memory modeling. Existing approaches primarily focus on increasing memory capacity, either by compressing historical information into fixed-size representations or by extending storage beyond GPU memory. However, these methods largely rely on global or coarse-grained representations, inevitably losing fine-grained visual information. In this work, we argue that streaming video memory should explicitly encode structured and semantically meaningful representations, particularly at the entity level. To this end, we propose MEMO, a novel framework that models streaming video through multi-level, entity-aware structured memory. MEMO performs multi-level perception to jointly capture global semantics, entity dynamics, and spatial structures, partitioning streaming video into semantically coherent chunks. Each chunk is organized into a structured memory, where lightweight global and entity-level representations serve as retrieval indices, while the corresponding high-resolution visual content is retained separately for on-demand access. At inference time, MEMO performs query-specific retrieval over the structured memory and selectively recalls relevant visual evidence for downstream reasoning. Notably, MEMO is training-free and plug-and-play with existing multimodal large language models. Extensive experiments on StreamingBench and OVO-Bench demonstrate that MEMO consistently improves multiple base models and achieves state-of-the-art performance.
Yinying Li, Yuqian Fu, Yulin Dai +3
East China Normal University · King Abdullah University of Science and Technology
Streaming video understanding models must answer queries at any moment during an ongoing stream, using only what they have observed so far and under fixed memory and computation budgets. Existing methods address this by adding memory banks, retrieval modules, or visual token compression to preserve long-range history. However, strong recent-window baselines show that indiscriminate history injection can dilute current-scene perception, suggesting that the key challenge is not whether to use memory, but how to allocate it selectively. We formulate this as budgeted online latent evidence allocation and propose \textbf{SelectStream}, a selective latent-memory framework that keeps the current observation directly visible to a frozen VLM while exposing historical information only through a compact, query-conditioned evidence budget. Three coordinated mechanisms govern when to write, what to preserve, and how to retrieve: surprise-driven adaptive windowing, priority-preserving consolidation, and query-conditioned graph reasoning over a fixed-capacity latent memory graph. Retrieved evidence is calibrated and injected as latent tokens for answer generation, without replaying frames or growing the context with stream length. Experimental results show that SelectStream achieves strong online streaming performance and preserves general video understanding, reaching 82.67% on StreamingBench, 67.03% on OVO-Bench, and 74.4% average accuracy on offline video benchmarks, while outperforming strong recent-window baselines and prior streaming memory methods.
Haonan Ge, Yiwei Wang, Hang Wu +1
University of California, Santa Barbara · University of California, Merced · The University of Queensland