Video Large Language Models (VideoLLMs) have achieved strong video understanding capabilities but incur substantial inference overhead due to the large number of visual tokens. Existing VideoLLM token compression methods largely rely on selection-independent scoring, overlooking cross-frame complementarity and consequently retaining redundant evidence across frames. Instead, we view video token selection as a progressive evidence accumulation process. It aims to retain visual evidence that is individually informative and collectively complementary under a limited token budget. Building on this insight, we introduce GleanVID, a training-free inference acceleration framework for VideoLLMs. Specifically, GleanVID first allocates the global token budget across frames according to temporal novelty and then selects tokens by jointly considering local representativeness and subspace complementarity, thereby preserving richer and less redundant visual evidence. Extensive experiments across diverse VideoLLMs and benchmarks demonstrate that GleanVID consistently achieves state-of-the-art performance. Notably, with only 25% of visual tokens, GleanVID preserves 98.6% of Qwen3-VL's original performance while reducing its prefill latency by 44.7%. On LLaVA-OV-7B, GleanVID at a 25% retention ratio even slightly surpasses the original model.
Figures & tables
Figure 1: Motivation of GleanVID. (a) Existing methods typically prune each frame independently, retaining redundant evidence, whereas GleanVID progressively selects tokens complementary to previously retained evidence. (b) Only candidates outside the retained subspace Q ( e.g. , x2 ) contribute new information. (c) Under tighter budgets, effective dimension collapses in existing methods but remains high for GleanVID.
Figure 2: Overview of GleanVID. (a) GleanVID compresses the visual tokens of T frames for the VideoLLM. (b) TNB measures the novelty of each neighborhood prototype against its history subspace and allocates per-frame budgets {kt} with a Gini-adaptive weight. (c) CTS selects the top- kt tokens by combining LRS with SCS, measured against the basis Qt−1 of retained tokens.
Figure 3: Qualitative comparison of frame-budget allocation under the same global token budget. Curves show the per-frame budget relative to the uniform allocation B/T (dashed line); highlighted frames (✓) contain newly introduced visual content.
Method
MVBench
LongVideo Bench
MLVU
VideoMME
Average
Overall
Short
Medium
Long
Score
%
Max Input Frames = 64
Qwen3-VL-8B-Instruct
69.2
62.8
68.9
66.9
78.8
66.2
55.8
67.0
100.0
Retention Ratio = 25%
VisionZip (CVPR’25)
64.8
58.9
63.4
63.1
73.7
60.3
55.4
62.6
93.4
VidCom 2 (EMNLP’25)
67.5
59.6
64.0
64.9
75.4
63.4
55.9
64.0
95.5
Table 1: Performance comparison with existing baselines on Qwen3-VL-8B-Instruct across multiple benchmarks and input settings. “Average” denotes the mean performance across benchmarks.
Method
MVBench
LongVideo Bench
MLVU
VideoMME
Average
Overall
Short
Medium
Long
Score
%
LLaVA-OV-7B
58.3
56.6
63.1
58.4
69.9
56.7
48.8
59.1
100.0
Retention Ratio = 25%
FastV (ECCV’24)
55.5
53.3
59.6
55.3
65.0
53.8
47.0
55.9
94.6
SparseVLM (ICML’25)
56.4
53.9
60.7
57.3
68.4
55.2
48.1
57.1
96.6
VisionZip (CVPR’25)
56.9
56.0
62.9
58.0
68.9
57.4
47.6
58.5
99.0
Table 2: Performance comparison with existing baselines on LLaVA-OV-7B across different benchmarks. We use the default 32-frame input setting.
Method
MVBench
LongVideo Bench
MLVU
VideoMME
Average
Overall
Short
Medium
Long
Score
%
LLaVA-Video-7B
60.4
58.9
67.3
64.4
77.3
62.4
53.4
62.8
100.0
Retention Ratio = 25%
FastV (ECCV’24)
52.1
54.8
57.8
58.6
68.7
58.4
48.7
55.8
88.9
SparseVLM (ICML’25)
55.4
54.2
58.9
60.1
71.1
59.1
50.1
57.2
91.1
VisionZip (CVPR’25)
57.9
56.3
62.6
62.5
73.6
62.3
51.9
59.8
95.2
Table 3: Performance comparison with existing baselines on LLaVA-Video-7B across different benchmarks. We use the default 64-frame input setting.
Method
Prefill Latency ↓ (s)
LLM Generation ↓ Latency (s)
Total Latency ↓ (s)
GPU Peak Memory ↓ (MB)
Throughput ↑ (items/s)
Performance ↑
Qwen3-VL-8B-Instruct
239.9
280.2
1369.5
22478.0
1.97
64.5
VidCom 2 (EMNLP’25)
119.0 ( ↓ 50.4%)
154.1 ( ↓ 45.0%)
1033.6 ( ↓ 24.5%)
19547.5 ( ↓ 13.0%)
2.61 (1.32 × )
62.4 ( ↓ 2.1)
HoliTom (NeurIPS’25)
134.3 ( ↓ 44.0%)
169.6 ( ↓ 39.5%)
1063.2 ( ↓ 22.4%)
19974.7 ( ↓ 11.1%)
2.54 (1.29 × )
60.2 ( ↓ 4.3)
FlashVID (ICLR’26)
145.0 ( ↓ 39.6%)
179.0 ( ↓ 36.1%)
1132.2 ( ↓ 17.3%)
40199.2 ( ↑ 78.8%)
2.38 (1.21 × )
62.3 ( ↓ 2.2)
V-CAST (2026’03)
121.2 ( ↓ 49.5%)
159.9 ( ↓ 42.9%)
1039.7 ( ↓ 24.1%)
19547.5 ( ↓ 13.0%)
2.60 (1.32 × )
63.5 ( ↓ 1.0)
GleanVID (Ours)
132.6 ( ↓ 44.7%)
165.7 ( ↓ 40.9%)
1040.1 ( ↓ 24.1%)
19547.5 ( ↓ 13.0%)
2.60 (1.32 × )
63.6 ( ↓ 0.9)
Table 4: Efficiency comparison on VideoMME using Qwen3-VL-8B-Instruct with a retention ratio of ρ=25% and a maximum of 32 input frames. “Prefill Latency” denotes the prompt-to-first-token time; “LLM Generation Latency” denotes the first-to-last-token decoding time; “Total Latency” denotes the end-to-end wall-clock time in our experimental setup; and “Throughput” is measured in items per second.
Variant
MVBench
LVB
MLVU
VideoMME
Rel. Perf. (%)
(a) Model Components
w/o TNB
65.5
58.0
60.4
61.2
95.4 ( ↓ 1.6)
w/o CTS
65.8
56.3
59.5
59.5
93.9 ( ↓ 3.1)
GleanVID
66.0
59.0
61.4
62.6
97.0
(b) Token-Selection Signals
LRS only
64.8
55.5
59.5
62.0
94.2 ( ↓ 2.8)
Table 5: Component and token-selection ablations. Relative performance is reported in % and its drop in points.
Variant
MVBench
LVB
MLVU
VideoMME
Rel. Perf. (%)
(a) Model Components
w/o TNB
65.5
58.0
60.4
61.2
95.4 ( ↓ 1.6)
w/o CTS
65.8
56.3
59.5
59.5
93.9 ( ↓ 3.1)
GleanVID
66.0
59.0
61.4
62.6
97.0
(b) Token-Selection Signals
LRS only
64.8
55.5
59.5
62.0
94.2 ( ↓ 2.8)
Table 5: Component and token-selection ablations. Relative performance is reported in % and its drop in points.
Variant
MVBench
LVB
MLVU
VideoMME
Rel. Perf. (%)
(a) Frame-Budget Allocation
Random
65.5
58.0
60.4
61.2
95.4 ( ↓ 1.6)
Uniform
64.8
58.3
60.6
62.1
95.8 ( ↓ 1.2)
Fixed ( α=0.5 )
66.1
58.5
61.1
62.4
96.6 ( ↓ 0.4)
Adaptive
66.0
59.0
61.4
62.6
97.0
(b) Complementarity Measure
Table 6: Frame-budget and complementarity ablations. Relative performance is reported in % and its drop in points.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Method
#Frames
Retention Ratio ρ
LongVideo Bench
MLVU
VideoMME
Average
Overall
Short
Medium
Long
Score
%
Vanilla
16 (1 × )
100%
57.1
58.7
60.6
70.0
57.8
54.0
58.8
100.0
V-CAST (2026’03)
61.8
65.9
66.1
77.7
64.1
56.4
64.6
109.9
GleanVID (Ours)
80 (5 × )
20%
61.9
66.6
66.3
77.7
64.2
56.9
64.9
110.4
V-CAST (2026’03)
61.0
67.3
65.9
78.0
64.4
55.3
64.7
110.1
GleanVID (Ours)
160 (10 × )
10%
62.2
68.8
67.2
77.2
66.6
57.8
66.1
112.4
Appendix
Table 7: Performance comparison under a fixed LLM-side visual-token budget on Qwen3-VL-8B-Instruct. The uncompressed 16-frame baseline, 80-frame inputs at 20% retention, and 160-frame inputs at 10% retention contain the same number of visual tokens after compression.
Variant
LongVideoBench
VideoMME
Overall
Short
Medium
Long
V-CAST
57.6
62.0
71.9
58.3
55.9
TNB → V-CAST budgeting
58.0
62.1
72.3
58.9
55.1
CTS → V-CAST scoring
58.1
61.8
72.6
58.4
54.3
GleanVID
59.0
62.6
73.6
59.6
54.6
Appendix
Table 8: Ablation study on component replacement on Qwen3-VL-8B-Instruct with 32 input frames and a 15% retention ratio. Each variant replaces one GleanVID component with its V-CAST counterpart; V-CAST corresponds to replacing both.
Rank rd
LongVideoBench
VideoMME
Overall
Short
Medium
Long
0
58.6
62.4
73.2
59.1
54.8
1†
59.0
62.6
73.6
59.6
54.6
2
58.8
62.5
73.3
59.4
54.8
4
58.8
62.6
73.6
60.0
54.1
Appendix
Table 9: Sensitivity analysis of the Context Debiasing rank on Qwen3-VL-8B-Instruct with 32 input frames and a 15% retention ratio. A dagger ( † ) marks the default setting.
Rank rd
LongVideoBench
VideoMME
Overall
Short
Medium
Long
0
58.6
62.4
73.2
59.1
54.8
1†
59.0
62.6
73.6
59.6
54.6
2
58.8
62.5
73.3
59.4
54.8
4
58.8
62.6
73.6
60.0
54.1
Appendix
Table 9: Sensitivity analysis of the Context Debiasing rank on Qwen3-VL-8B-Instruct with 32 input frames and a 15% retention ratio. A dagger ( † ) marks the default setting.
History h
LongVideoBench
VideoMME
Overall
Short
Medium
Long
1
58.6
62.1
72.7
59.9
53.9
2
58.9
61.7
72.7
58.9
53.7
4
58.9
62.1
72.4
59.6
54.3
all†
59.0
62.6
73.6
59.6
54.6
Appendix
Table 10: Sensitivity analysis of the temporal history length on Qwen3-VL-8B-Instruct with 32 input frames and a 15% retention ratio. A dagger ( † ) marks the default setting.
Neighborhood
LongVideoBench
VideoMME
Overall
Short
Medium
Long
1×1
57.9
62.2
72.3
59.3
55.0
2×2†
59.0
62.6
73.6
59.6
54.6
4×4
58.1
62.4
73.2
59.6
54.3
Appendix
Table 11: Sensitivity to the neighborhood size used for temporal prototypes in TNB. A dagger ( † ) marks the default setting.
Neighborhood
LongVideoBench
VideoMME
Overall
Short
Medium
Long
1×1
57.9
62.2
72.3
59.3
55.0
2×2†
59.0
62.6
73.6
59.6
54.6
4×4
58.1
62.4
73.2
59.6
54.3
Appendix
Table 11: Sensitivity to the neighborhood size used for temporal prototypes in TNB. A dagger ( † ) marks the default setting.
Neighborhood
LongVideoBench
VideoMME
Overall
Short
Medium
Long
1×1
58.8
62.5
73.4
59.7
54.3
2×2†
59.0
62.6
73.6
59.6
54.6
4×4
58.4
62.4
73.2
59.6
54.3
Appendix
Table 12: Sensitivity to the neighborhood size used for local prototypes in LRS. A dagger ( † ) marks the default setting.
Temperature τ
LongVideoBench
VideoMME
Overall
Short
Medium
Long
1
58.3
61.6
72.3
58.7
53.8
2
58.6
62.2
73.0
58.9
54.7
4†
59.0
62.6
73.6
59.6
54.6
8
58.6
62.8
73.6
59.9
55.0
Appendix
Table 13: Sensitivity to the LRS temperature τ . A dagger ( † ) marks the default setting.
Temperature τ
LongVideoBench
VideoMME
Overall
Short
Medium
Long
1
58.3
61.6
72.3
58.7
53.8
2
58.6
62.2
73.0
58.9
54.7
4†
59.0
62.6
73.6
59.6
54.6
8
58.6
62.8
73.6
59.9
55.0
Appendix
Table 13: Sensitivity to the LRS temperature τ . A dagger ( † ) marks the default setting.
[αmin,αmax]
LongVideoBench
VideoMME
Overall
Short
Medium
Long
[0.2,0.4]
58.9
62.6
73.7
59.3
54.9
[0.2,0.6]†
59.0
62.6
73.6
59.6
54.6
[0.2,0.8]
58.7
62.4
73.2
59.8
54.2
Appendix
Table 14: Sensitivity to the adaptive allocation bounds. A dagger ( † ) marks the default setting.
Method
Prefill Latency ↓ (s)
LLM Generation Latency ↓ (s)
GPU Peak Memory ↓ (MB)
Performance ↑
LLaVA-OneVision-7B
99.0
180.8
21,969.0
58.4
VidCom 2 (EMNLP’25)
26.5 ( ↓ 73.2%)
108.6 ( ↓ 39.9%)
19,512.6 ( ↓ 11.2%)
58.4
FastVID (NeurIPS’25)
78.8 ( ↓ 20.4%)
160.4 ( ↓ 11.3%)
20,803.5 ( ↓ 5.3%)
58.3 ( ↓ 0.1)
FlashVID (ICLR’26)
28.4 ( ↓ 71.3%)
146.1 ( ↓ 19.2%)
19,059.9 ( ↓ 13.2%)
58.7 ( ↑ 0.3)
V-CAST (2026’03)
27.0 ( ↓ 72.7%)
113.8 ( ↓ 37.0%)
19,512.6 ( ↓ 11.2%)
58.2 ( ↓ 0.2)
GleanVID (Ours)
26.7 ( ↓ 73.1%)
122.8 ( ↓ 32.1%)
19,512.6 ( ↓ 11.2%)
59.7 ( ↑ 1.3)
Appendix
Table 15: Efficiency and performance comparison on VideoMME using LLaVA-OneVision-7B with 32 input frames and a retention ratio of ρ=25% . Efficiency reductions and performance changes are computed relative to the uncompressed model.
Figure 4: Cross-frame evidence complementarity analysis on 100 LongVideoBench videos with 32 input frames and a 15% retention ratio. All variants use identical per-frame token budgets. Higher Effective Rank and Residual Energy indicate richer and more complementary retained evidence, while lower Coverage Error indicates better representation of discarded tokens.
Figure 5: Qualitative comparison on LLaVA-OneVision-7B. Vanilla denotes the uncompressed model, while the remaining methods apply visual-token compression to the same model. The examples cover physical reasoning, action prediction, fine-grained grounding, and long-range video understanding. Correct and incorrect predictions are shown in green and red, respectively. GleanVID preserves task-relevant temporal and fine-grained evidence under a constrained visual-token budget.
BNRist, THUIBCS, BLBCI, School of Software, Tsinghua University · Yangtze Delta Region Institute, Tsinghua University · Beijing University of Posts and Telecommunications, Beijing, China