Streaming video assistance requires models to answer asynchronous questions from an observed prefix under a fixed context budget. Existing approaches model response timing or compress history, but an online state formed before future questions are known can omit visual details before later questions reveal their relevance; the retained state alone cannot recover them. We introduce Watch-Think-Interact (WTI), a closed-loop framework for multi-question streaming video reasoning. WTI maintains compact natural-language memory entries tagged with source-video time ranges; these entries support direct reasoning when sufficient and otherwise anchor selective recall of finer visual evidence. For each question, WTI answers when current context and memory suffice, continues watching when required evidence has not appeared, or recalls a relevant past interval and decides again after incorporating the returned chunks, without replaying the full observed history. To train this behavior, we construct WTI-82K, comprising 82,335 timed questions across 4,812 causally aligned trajectories, and develop Stream-GDPO to optimize complete multi-question streaming rollouts using trajectory-level feedback for response timing, source-video recall, and memory updates. WTI achieves state-of-the-art aggregate performance among the compared open-source streaming baselines, reaching 83.3% on StreamingBench and 73.6% weighted overall accuracy on OVO-Bench.
Figures & tables
Figure 1: Motivation for Watch-Think-Interact. In causal streaming video, asynchronous questions may require current, past, or not-yet-observed evidence. WTI uses a bounded active window for current perception, compact time-indexed memory for direct reasoning and temporal localization, and selective recall to reload a past visual interval.
Figure 2: Overview of Watch-Think-Interact. For each arriving chunk, WTI updates a short thinking state and chooses to remain silent, answer, or recall. Recall returns visual evidence and loops back to the decision at the same time step; controller-triggered memory updates separately compress elapsed observations into a bounded temporal index. Stream-GDPO optimizes the resulting complete interaction trajectories.
Figure 3: WTI-82K construction pipeline. WTI-82K converts videos into streaming-causal trajectories by standardizing clips, aligning chunk-level evidence chains, generating QAs from aligned evidence, placing query and answerable times, and checking timestamp bounds, answer availability, action grammar, and evidence-interval fields. The resulting samples cover Real-Time, Backward Tracing, and Proactive interactions under the same visibility constraints as inference.
Model
Real-Time
Backward
Forward
Overall
OCR
ACR
ATR
STU
FPD
OJR
Avg.
EPM
ASI
HLD
Avg.
REC
SSR
CRR
Avg.
Avg.
Proprietary models
GPT-4o
69.8
64.2
71.6
51.1
70.3
59.8
64.5
57.9
75.7
48.7
60.8
27.6
73.2
59.4
53.4
59.5
Gemini-1.5-Pro
85.9
67.0
79.3
58.4
63.4
62.0
69.3
58.6
76.4
52.6
62.5
35.5
74.2
61.7
57.2
63.0
Open-source offline models
LongVU-7B
53.7
53.2
62.9
47.8
68.3
59.8
57.6
40.7
59.5
4.8
35.0
12.2
69.5
60.8
47.5
46.7
Table 1: Results on OVO-Bench. Scores are grouped by Real-Time, Backward, and Forward tasks; Avg. gives the corresponding group average, and Overall gives the benchmark average.
Table 5
Figure 4: Training-time GSPO/Stream-GDPO diagnostics. Panels report outcome accuracy, BT recall triggering, post-recall outcome accuracy, and memory quality; faint and bold curves show raw and rolling values.
Streaming video understanding requires answering questions that arrive at arbitrary moments over an unbounded video stream. Existing systems primarily focus on what to retain in a bounded memory, yet access that memory using the same fixed-cost procedure for every query, despite substantial variation in the evidence required. We argue that deciding how deeply to access memory for each query is as important as deciding what the memory should store. To this end, we introduce StreamScout, an adaptive inference framework that maintains only a lightweight textual timeline in context as the stream unfolds. At query time, StreamScout progressively augments the timeline with up to three increasingly informative visual views: a glance at recent frames, a uniform look-back over the past stream, and query-salient retrieval. At each stage, the model answers immediately if the available evidence is sufficient; otherwise, it escalates to the next view. To improve this stop-or-escalate policy, we probe the cascade on an auxiliary set and distill the model's empirical competence boundary into supervision for a lightweight LoRA adaptation, yielding StreamScout-S. We further refine the policy through reinforcement learning, allowing the model to explore stopping behaviors beyond imitation of the distilled decisions, yielding StreamScout-R. Across three backbones and three streaming benchmarks, StreamScout and its variants consistently outperform prior streaming methods while substantially reducing inference cost and token consumption; on OVO-Bench, for instance, StreamScout-S improves Qwen3-VL-8B by 14.65 points while using 59% fewer tokens than uniform sampling and answering in 1.04 s on average.
Ce Zhang, Jing Bi, Jinxi He +9
Carnegie Mellon University · TikTok · University of Rochester +2
Streaming video reasoning requires models to operate in a setting where history grows without bound while meaningful evidence remains scarce. In such a landscape, relevant signal is like an oasis-small, critical, and easily lost in a desert of redundancy. Enlarging memory only widens the desert; aggressive compression dries up the oasis. The real difficulty lies in discovering where to look, not how much to remember. We therefore introduce OASIS, a novel framework for streaming video reasoning that tackles this challenge through structured, on-demand retrieval. It organizes streaming history into hierarchical events and performs reasoning as controlled refinement-short-context inference first, followed by semantically grounded retrieval only when uncertainty arises. As the retrieval is driven by high-level intent rather than embedding similarity, the retrieved memory is substantially more accurate and less noisy. Additionally, the mechanism is plug-and-play, training-free, and readily attaches to different streaming MLLM backbones. Experiments across multiple benchmarks and backbones show that OASIS achieves strong gains in long-horizon accuracy and compositional reasoning with bounded token cost and low request delay. Code is available at https://github.com/Solus-sano/OASIS.
Zhijia Liang, Jiaming Li, Weikai Chen +3
Sun Yat-sen University · OPPO AI Center · Shenzhen Loop Area Institute +1
Streaming video understanding requires multimodal large language models (MLLMs) to preserve relevant evidence from continuously evolving streams under strict causality and bounded memory. Yet existing paradigms remain limited: model-based methods require intrusive backbone updates, while memory-based methods expend substantial visual-encoding computation on temporally redundant content and rely on rigid access to visual history. To address these limitations, we introduce StreamFlow, an efficient visual memory framework that enables dynamic, on-demand access to historical visual information. StreamFlow combines a lightweight, dynamics-aware mid-term memory that filters temporal redundancy before visual encoding with a latent long-term memory that consolidates historical video content into visual latents accessible to subsequent reasoning. During generation, an attention-guided retrieval mechanism injects relevant visual latents when the model's reliance on visual evidence weakens. StreamFlow achieves state-of-the-art streaming video understanding performance, reaching 67.73% overall accuracy on StreamingBench, while also delivering strong performance on offline long-video benchmarks. Relative to the vanilla setting, it improves the visual attention score (VAS) by 59.1% while reducing end-to-end latency and peak memory by 50.4% and 21.1%, respectively, enabling more visually grounded and efficient reasoning.
Muxin Fu, Yifan Zhang, Wentao Zhang +5
Tongji University · Nanyang Technological University · University of Michigan +2