Streaming video understanding is a critical capability for real-world applications, including embodied intelligence, autonomous driving, industrial monitoring, surveillance and early warning, and wearable assistants. However, processing continuous video streams with multimodal large language models (MLLMs) is computationally expensive. Existing efforts have explored reducing streaming overhead through visual token pruning, token merging, quantization, on-demand frame retrieval, and context offloading. However, most existing methods overlook the dimension of model depth. Repeatedly executing full-depth MLLM prefill over incoming frames is prohibitively expensive, incurring substantial computational overhead and causing the KV cache to grow at a rate directly proportional to the prefill depth. To address these challenges, we propose ShallowStream, a novel framework that leverages the shallow layers of an MLLM to simultaneously perform frame encoding and retrieval index building. During stream processing, ShallowStream maintains an always-on lightweight index using the KV cache of shallow layers. During query-time answering, we leverage the attention scores generated by the shallow layers to score context frames and employ a diversity-aware selection strategy to retrieve precise and comprehensive evidence. ShallowStream achieves performance on par with the strongest existing streaming methods, while reducing per-frame prefill latency and 10-second end-to-end latency by up to 52.1x and 11.9x, respectively. Our code is available at https://github.com/CURRENTF/ShallowStream.
Figures & tables
Figure 1 : Continuous-stream efficiency. Per-frame prefill and end-to-end latency with one query every 20s. Latencies and their 95% confidence intervals are scaled by the corresponding ShallowStream mean and shown on a logarithmic scale. Performance is the OVO-Bench average.
Figure 2 : Layer-wise retrieval quality and stream-time prefill cost on LVBench. Results for Qwen3-VL-8B ( left ) and LLaVA-OneVision-7B ( right ) as the retrieval layer varies, with all other settings fixed. QA measures full-depth answer accuracy using units selected at that layer, while latency measures stream-time prefill per video unit up to that layer.
Figure 3 : Overview of ShallowStream. Stage I: Query-Agnostic Shallow Index Construction (Section 4.1 ). Incoming video units traverse only the shallow MLLM layers, whose KVs form a lightweight index that can optionally be compressed into fixed-size historical clusters. Stage II: Query-Time Answering (Section 4.2 ). A text-only query-logit gate determines whether to retrieve history; if activated, shallow-layer attention, cross-layer unit voting, and max-min diversity selection identify complementary evidence for full-depth answering alongside recent context.
Model / Method
# Frames
Real-Time Visual Perception
Backward Tracing
Avg. ↑
OCR
ACR
ATR
STU
FPD
OJR
Avg. ↑
EPM
ASI
HLD
Avg. ↑
Open-source Online MLLMs
VideoLLM-online-8B [ 5 ]
2 fps
8.1
23.9
12.1
14.0
45.5
21.2
20.8
22.2
18.8
12.2
17.7
19.3
Flash-VStream-7B [ 52 ]
1 fps
25.5
32.1
29.3
33.7
29.7
28.8
29.9
36.4
33.8
5.9
25.4
27.6
Dispider-7B [ 29 ]
1 fps
57.7
49.5
62.1
44.9
61.4
51.6
54.6
48.5
55.4
4.3
36.1
45.3
TimeChat-Online-7B [ 49 ]
1 fps
75.2
46.8
70.7
47.8
69.3
61.4
61.9
55.9
59.5
9.7
41.7
51.8
Table 1: Main results on OVO-Bench. Baseline results marked with ‡ are from our reruns. Results marked with † enable long-cluster compression. Dashes indicate unreported entries.
Model / Method
# Frames
Real-Time Visual Understanding
Avg. ↑
OP
CR
CS
ATP
EU
TR
PR
SU
ACP
CT
Open-source Online MLLMs
Flash-VStream-7B [ 52 ]
–
25.9
43.6
24.9
23.9
27.3
13.1
18.5
25.2
23.9
48.7
23.2
VideoLLM-online-8B [ 5 ]
2 fps
39.1
40.1
34.5
31.1
46.0
32.4
31.5
34.2
42.5
27.9
36.0
Dispider-7B [ 29 ]
1 fps
74.9
75.5
74.1
73.1
74.4
59.9
76.1
62.9
62.2
45.8
67.6
TimeChat-Online-7B [ 49 ]
1 fps
80.2
82.0
79.5
83.3
76.1
78.5
78.7
64.6
69.6
58.0
75.4
Table 2: Main results on real-time visual understanding subset of StreamingBench. Baseline results marked with ‡ are from our controlled reruns. Results marked with † enable long-cluster compression. Dashes indicate unreported results.
Figure 4 : Memory scaling and real-time compute demand. Left: Peak GPU allocation from 64 to 1,024 frames, averaged over five videos; both ShallowStream variants use the calibrated Gate and token-vote retriever. Right: Total compute for a query interval x , obtained by combining mean query computation with stream-time prefill accumulated over x seconds at 1 FPS. The dashed boundary y=x separates configurations that keep pace with the stream (below) from those whose compute exceeds the available interval (above).
Figure 7
Figure 7 : Qualitative evidence retrieval on OVO-Bench EPM. For the question “Where is the red and white checkered rug?” (ground-truth answer: E), the rows show the eight units selected by pooled shallow Q-K, pooled shallow Q-K with max-min diversity, and the final token-vote retriever with max-min diversity. Red outlines indicate frames containing the rug beneath the dining table. The final retriever selects complementary views and produces the correct answer.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone
τg
Precision
95% lower bound
Recall
Qwen3-VL-8B
11.0
96.04%
91.47%
97.00%
LLaVA-OneVision-7B
0.875
97.78%
90.63%
44.00%
Appendix
Table 3: Query-logit Gate calibration on the benchmark-independent query-only set.
Setting
Qwen3-VL-8B
LLaVA-OneVision-7B
Backbone and Stream Processing
Pruning boundary P
5
4
Sampling rate (OVO / StreamingBench)
1 / 1 FPS
1 / 1 FPS
Video unit
2 frames
1 frame
Stream prefill window
64 units, 128 frames
128 units, 128 frames
Query Routing
Appendix
Table 4: Main implementation settings for the two evaluated backbones.
Measurement
Setting or Component
Cost
Stream-time prefill
P=1
9.47 ms/frame
P=5
10.72 ms/frame
P=19
15.53 ms/frame
Query-time computation
Full query
1.759 s/query
Gate
92.7 ms/query
Evidence selection
106.0 ms/query
Appendix
Table 5: Cost summary under the final calibrated Gate and token-vote retriever. Stream-time prefill is averaged over the same five long videos; query components use the matched full-history LC-off setting.
Figure 8 : Sensitivity to the pruning boundary P on OVO-Bench Backward under the final calibrated Gate and token-vote retriever. Routing and evidence selection are held fixed across depths.
Figure 9 : Query-time scaling with retained history on an NVIDIA RTX 5090. Results span six history lengths from 64 to 1,024 frames. (a) Total query latency and time to first token (TTFT). (b) Evidence-selection and Gate latency. (c) Decoding latency. Solid circles and dashed squares denote full-history LC-off and the matched final LC-on setting, respectively. Curves show means over five retrieval queries; shaded regions indicate 95% confidence intervals.