cs.CVSep 29, 2026

When to Retrieve, When to Stay: Uncertainty-Aware Temporal Evidence Allocation for Streaming Video-LLMs

Authors: Xiang Hu, Jiazuo Yu, Lu Zhang, Yunzhi Zhuge, Huchuan Lu

Organizations: IIAU Lab, Dalian University of Technology

Abstract

Streaming video understanding requires Video Large Language Models (Video-LLMs) to reason over continuous visual streams under causal constraints. As the visual history grows, a bounded visual?processing budget requires evidence selection that balances temporal recency with query relevance. Recent-only selection excludes potentially relevant historical evidence, whereas Semantic-only retrieval can displace useful recent context when relevance scores are ambiguous. We introduce WRWS (When to Retrieve, When to Stay), a training-free framework for uncertainty-adaptive evidence allocation. A lightweight external vision-language encoder scores query relevance across the observed history, while an adaptive allocation module uses the normalized entropy of the similarity distribution as a proxy for retrieval uncertainty. WRWS favors semantic retrieval when relevance cues are reliable and strengthens the recency prior under uncertainty. Following a retrieve-first, encode-later pipeline, WRWS selects evidence before target-model visual encoding, such that only the selected observations are processed by the costly target Video-LLM. Experiments across four Video-LLM families and multiple model scales demonstrate competitive accuracy on StreamingBench and OVO-Bench. In our efficiency evaluation, WRWS reduces average vision-to-answer time to 47.93% of the state-of-the-art method. Code will be released.

Figures & tables

Explore similar work

CardsList
  1. What Should a Streaming Video Model Remember?

    Jun 15, 2026Haonan Ge, Yiwei Wang, Hang Wu +1Streaming Video UnderstandingStreaming

  2. StreamFlow: Dynamic Memory Flows for Streaming Video Understanding

    Aug 11, 2026Muxin Fu, Yifan Zhang, Wentao Zhang +5Streaming Video UnderstandingLong-Video Benchmarks

  3. Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

    Sep 3, 2026Hongyu Qu, Guangming Yao, Ling Xing +7Streaming Video UnderstandingModern Data-Streaming Systems