When to Retrieve, When to Stay: Uncertainty-Aware Temporal Evidence Allocation for Streaming Video-LLMs
Organizations: IIAU Lab, Dalian University of Technology
Abstract
Streaming video understanding requires Video Large Language Models (Video-LLMs) to reason over continuous visual streams under causal constraints. As the visual history grows, a bounded visual?processing budget requires evidence selection that balances temporal recency with query relevance. Recent-only selection excludes potentially relevant historical evidence, whereas Semantic-only retrieval can displace useful recent context when relevance scores are ambiguous. We introduce WRWS (When to Retrieve, When to Stay), a training-free framework for uncertainty-adaptive evidence allocation. A lightweight external vision-language encoder scores query relevance across the observed history, while an adaptive allocation module uses the normalized entropy of the similarity distribution as a proxy for retrieval uncertainty. WRWS favors semantic retrieval when relevance cues are reliable and strengthens the recency prior under uncertainty. Following a retrieve-first, encode-later pipeline, WRWS selects evidence before target-model visual encoding, such that only the selected observations are processed by the costly target Video-LLM. Experiments across four Video-LLM families and multiple model scales demonstrate competitive accuracy on StreamingBench and OVO-Bench. In our efficiency evaluation, WRWS reduces average vision-to-answer time to 47.93% of the state-of-the-art method. Code will be released.
Figures & tables
| Model | #Frames | StreamingBench | OVO-Bench | |||
| Real-Time | Backward | Real-Time | Forward | Avg. | ||
| Human | – | 91.46 | 92.33 | 93.20 | 92.90 | 92.81 |
| Proprietary MLLMs | ||||||
| Gemini 1.5 Pro ( Team et al., 2024 ) | 1 fps | 75.69 | 62.54 | 69.32 | 57.15 | 63.00 |
| GPT-4o ( Hurst et al., 2024 ) | 64 | 73.28 | 60.75 | 64.46 | 53.40 | 59.54 |
| Claude 3.5 Sonnet ( Anthropic, 2024 ) | 20 | 72.44 | – | – | – | – |
| Allocation | Qwen2.5 | Qwen3 |
| (Semantic-only) | 48.46 | 51.76 |
| 52.84 | 56.15 | |
| (Recent-only) | 55.73 | 58.10 |
| Adaptive † | 55.98 | 58.59 |
| Metric | Target-VLM visual encoding | Other pipeline overhead | Inference TTFT | Total VTAT |
| Time (s) | 7.596 | 4.342 | 0.243 | 12.181 |
| Fraction (%) | 62.36 | 35.65 | 1.99 | 100.00 |
| Model | Method | Streaming | Backward | Real-Time | Forward | B+R | All |
| HERMES | 81.32 | 46.78 | 73.21 | – | 60.00 | – | |
| Qwen3-VL-8B | WRWS | 80.39 | 50.76 | 80.60 | 44.17 | 65.68 | 58.51 |
| HERMES | 78.40 | 54.00 | 71.90 | – | 62.95 | – | |
| Qwen3-VL-4B | WRWS | 80.07 | 52.24 | 78.64 | 41.26 | 65.44 | 57.38 |
| Model | Method | Streaming | Backward | Real-Time | Forward | B+R | All | |
| WRWS | 2 | 78.83 | 60.46 | 83.03 | 42.91 | 71.74 | 62.13 | |
| WRWS | 4 | 83.51 | 62.17 | 85.33 | 45.26 | 73.75 | 64.25 | |
| WRWS | 6 | 83.71 | 62.35 | 84.94 | 45.77 | 73.59 | 64.35 | |
| WRWS | 8 | 83.67 | 62.35 | 84.62 | 47.33 | 73.49 | 64.77 | |
| 30B-A3B | WRWS | 16 | 84.31 | 64.53 | 81.29 | 50.41 | 72.91 | 65.41 |
| Model | Method | Streaming | Backward | Real-Time | Forward | B+R | All | |
| Qwen2.5-VL | ||||||||
| SimpleStream | 4 | 77.83 | 48.50 | 74.64 | 43.19 | 61.57 | 55.44 | |
| 3B | WRWS | 4 | 78.07 | 47.27 | 74.56 | 45.21 | 60.91 | 55.68 |
| SimpleStream | 4 | 80.31 | 49.03 | 80.58 | 42.05 | 64.81 | 57.22 | |
| 32B | WRWS | 4 | 80.23 | 49.28 | 80.13 | 42.25 | 64.65 | 57.22 |
| SimpleStream | 16 | 84.31 | 61.46 | 79.52 | 42.07 | 70.49 | 61.02 | |
| Model | Allocation | Backward | Real-Time | Forward | B+R | All |
| (Semantic-only) | 47.12 | 59.08 | 39.17 | 53.10 | 48.46 | |
| 47.97 | 64.46 | 40.72 | 56.21 | 51.05 | ||
| 50.22 | 68.34 | 39.96 | 59.28 | 52.84 | ||
| 49.97 | 73.95 | 38.58 | 61.96 | 54.16 | ||
| (Recent-only) | 50.43 | 77.20 | 39.56 | 63.82 | 55.73 | |
| Qwen2.5-VL-7B | Adaptive | 50.59 | 77.13 | 40.23 | 63.86 | 55.98 |
| Group | # Samples | VTAT (s) | Inference TTFT (s) | GPU Memory (GB) | |||||
| WRWS | SimpleStream | Ratio | WRWS | SimpleStream | Ratio | WRWS | SimpleStream | ||
| Overall | 3035 | 5.84 | 12.18 | 47.93% | 0.2431 | 0.2429 | 100.08% | 20.76 | 17.28 |
| Backward | 631 | 6.37 | 15.56 | 40.92% | 0.2788 | 0.2793 | 99.82% | ||
| Real-Time | 837 | 7.17 | 16.50 | 43.44% | 0.2863 | 0.2866 | 99.91% | ||
| Forward | 1567 | 4.92 | 8.51 | 57.74% | 0.2057 | 0.2050 | 100.35% | ||