cs.AISep 29, 2026

Watch-Think-Interact: Bootstrapping Long-Horizon Multi-Turn Streaming Video Reasoning with Reinforcement Learning

Authors: Ziheng Huang, Yicheng Bao, Xueheng Li, Zhenkun Gao, Bangwei Liu, Kunquan Li, Yuxiang Shen, Bangyan Li, +3 more

Organizations: East China Normal University · University of Science and Technology of China · Xiamen University

Abstract

Streaming video assistance requires models to answer asynchronous questions from an observed prefix under a fixed context budget. Existing approaches model response timing or compress history, but an online state formed before future questions are known can omit visual details before later questions reveal their relevance; the retained state alone cannot recover them. We introduce Watch-Think-Interact (WTI), a closed-loop framework for multi-question streaming video reasoning. WTI maintains compact natural-language memory entries tagged with source-video time ranges; these entries support direct reasoning when sufficient and otherwise anchor selective recall of finer visual evidence. For each question, WTI answers when current context and memory suffice, continues watching when required evidence has not appeared, or recalls a relevant past interval and decides again after incorporating the returned chunks, without replaying the full observed history. To train this behavior, we construct WTI-82K, comprising 82,335 timed questions across 4,812 causally aligned trajectories, and develop Stream-GDPO to optimize complete multi-question streaming rollouts using trajectory-level feedback for response timing, source-video recall, and memory updates. WTI achieves state-of-the-art aggregate performance among the compared open-source streaming baselines, reaching 83.3% on StreamingBench and 73.6% weighted overall accuracy on OVO-Bench.

Figures & tables

Explore similar work

CardsList
  1. StreamScout: Learning When to Look Deeper for Streaming Video Understanding

    Aug 31, 2026Ce Zhang, Jing Bi, Jinxi He +9Streaming Video UnderstandingBounded-Memory

  2. OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning

    Apr 18, 2026Zhijia Liang, Jiaming Li, Weikai Chen +3Streaming Video UnderstandingStreaming

  3. StreamFlow: Dynamic Memory Flows for Streaming Video Understanding

    Aug 11, 2026Muxin Fu, Yifan Zhang, Wentao Zhang +5Streaming Video UnderstandingLong-Video Benchmarks