GleanVID: Complementary Token Selection for Efficient Video Large Language Models
Organizations: Beihang University · Communication University of China
Abstract
Video Large Language Models (VideoLLMs) have achieved strong video understanding capabilities but incur substantial inference overhead due to the large number of visual tokens. Existing VideoLLM token compression methods largely rely on selection-independent scoring, overlooking cross-frame complementarity and consequently retaining redundant evidence across frames. Instead, we view video token selection as a progressive evidence accumulation process. It aims to retain visual evidence that is individually informative and collectively complementary under a limited token budget. Building on this insight, we introduce GleanVID, a training-free inference acceleration framework for VideoLLMs. Specifically, GleanVID first allocates the global token budget across frames according to temporal novelty and then selects tokens by jointly considering local representativeness and subspace complementarity, thereby preserving richer and less redundant visual evidence. Extensive experiments across diverse VideoLLMs and benchmarks demonstrate that GleanVID consistently achieves state-of-the-art performance. Notably, with only 25% of visual tokens, GleanVID preserves 98.6% of Qwen3-VL's original performance while reducing its prefill latency by 44.7%. On LLaVA-OV-7B, GleanVID at a 25% retention ratio even slightly surpasses the original model.
Figures & tables
| Method | MVBench | LongVideo Bench | MLVU | VideoMME | Average | ||||
| Overall | Short | Medium | Long | Score | % | ||||
| Max Input Frames = 64 | |||||||||
| Qwen3-VL-8B-Instruct | 69.2 | 62.8 | 68.9 | 66.9 | 78.8 | 66.2 | 55.8 | 67.0 | 100.0 |
| Retention Ratio = 25% | |||||||||
| VisionZip (CVPR’25) | 64.8 | 58.9 | 63.4 | 63.1 | 73.7 | 60.3 | 55.4 | 62.6 | 93.4 |
| VidCom 2 (EMNLP’25) | 67.5 | 59.6 | 64.0 | 64.9 | 75.4 | 63.4 | 55.9 | 64.0 | 95.5 |
| Method | MVBench | LongVideo Bench | MLVU | VideoMME | Average | ||||
| Overall | Short | Medium | Long | Score | % | ||||
| LLaVA-OV-7B | 58.3 | 56.6 | 63.1 | 58.4 | 69.9 | 56.7 | 48.8 | 59.1 | 100.0 |
| Retention Ratio = 25% | |||||||||
| FastV (ECCV’24) | 55.5 | 53.3 | 59.6 | 55.3 | 65.0 | 53.8 | 47.0 | 55.9 | 94.6 |
| SparseVLM (ICML’25) | 56.4 | 53.9 | 60.7 | 57.3 | 68.4 | 55.2 | 48.1 | 57.1 | 96.6 |
| VisionZip (CVPR’25) | 56.9 | 56.0 | 62.9 | 58.0 | 68.9 | 57.4 | 47.6 | 58.5 | 99.0 |
| Method | MVBench | LongVideo Bench | MLVU | VideoMME | Average | ||||
| Overall | Short | Medium | Long | Score | % | ||||
| LLaVA-Video-7B | 60.4 | 58.9 | 67.3 | 64.4 | 77.3 | 62.4 | 53.4 | 62.8 | 100.0 |
| Retention Ratio = 25% | |||||||||
| FastV (ECCV’24) | 52.1 | 54.8 | 57.8 | 58.6 | 68.7 | 58.4 | 48.7 | 55.8 | 88.9 |
| SparseVLM (ICML’25) | 55.4 | 54.2 | 58.9 | 60.1 | 71.1 | 59.1 | 50.1 | 57.2 | 91.1 |
| VisionZip (CVPR’25) | 57.9 | 56.3 | 62.6 | 62.5 | 73.6 | 62.3 | 51.9 | 59.8 | 95.2 |
| Method | Prefill Latency (s) | LLM Generation Latency (s) | Total Latency (s) | GPU Peak Memory (MB) | Throughput (items/s) | Performance |
| Qwen3-VL-8B-Instruct | 239.9 | 280.2 | 1369.5 | 22478.0 | 1.97 | 64.5 |
| VidCom 2 (EMNLP’25) | 119.0 ( 50.4%) | 154.1 ( 45.0%) | 1033.6 ( 24.5%) | 19547.5 ( 13.0%) | 2.61 (1.32 ) | 62.4 ( 2.1) |
| HoliTom (NeurIPS’25) | 134.3 ( 44.0%) | 169.6 ( 39.5%) | 1063.2 ( 22.4%) | 19974.7 ( 11.1%) | 2.54 (1.29 ) | 60.2 ( 4.3) |
| FlashVID (ICLR’26) | 145.0 ( 39.6%) | 179.0 ( 36.1%) | 1132.2 ( 17.3%) | 40199.2 ( 78.8%) | 2.38 (1.21 ) | 62.3 ( 2.2) |
| V-CAST (2026’03) | 121.2 ( 49.5%) | 159.9 ( 42.9%) | 1039.7 ( 24.1%) | 19547.5 ( 13.0%) | 2.60 (1.32 ) | 63.5 ( 1.0) |
| GleanVID (Ours) | 132.6 ( 44.7%) | 165.7 ( 40.9%) | 1040.1 ( 24.1%) | 19547.5 ( 13.0%) | 2.60 (1.32 ) | 63.6 ( 0.9) |
| Variant | MVBench | LVB | MLVU | VideoMME | Rel. Perf. (%) |
| (a) Model Components | |||||
| w/o TNB | 65.5 | 58.0 | 60.4 | 61.2 | 95.4 ( 1.6) |
| w/o CTS | 65.8 | 56.3 | 59.5 | 59.5 | 93.9 ( 3.1) |
| GleanVID | 66.0 | 59.0 | 61.4 | 62.6 | 97.0 |
| (b) Token-Selection Signals | |||||
| LRS only | 64.8 | 55.5 | 59.5 | 62.0 | 94.2 ( 2.8) |
| Variant | MVBench | LVB | MLVU | VideoMME | Rel. Perf. (%) |
| (a) Model Components | |||||
| w/o TNB | 65.5 | 58.0 | 60.4 | 61.2 | 95.4 ( 1.6) |
| w/o CTS | 65.8 | 56.3 | 59.5 | 59.5 | 93.9 ( 3.1) |
| GleanVID | 66.0 | 59.0 | 61.4 | 62.6 | 97.0 |
| (b) Token-Selection Signals | |||||
| LRS only | 64.8 | 55.5 | 59.5 | 62.0 | 94.2 ( 2.8) |
| Variant | MVBench | LVB | MLVU | VideoMME | Rel. Perf. (%) |
| (a) Frame-Budget Allocation | |||||
| Random | 65.5 | 58.0 | 60.4 | 61.2 | 95.4 ( 1.6) |
| Uniform | 64.8 | 58.3 | 60.6 | 62.1 | 95.8 ( 1.2) |
| Fixed ( ) | 66.1 | 58.5 | 61.1 | 62.4 | 96.6 ( 0.4) |
| Adaptive | 66.0 | 59.0 | 61.4 | 62.6 | 97.0 |
| (b) Complementarity Measure | |||||
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | #Frames | Retention Ratio | LongVideo Bench | MLVU | VideoMME | Average | ||||
| Overall | Short | Medium | Long | Score | % | |||||
| Vanilla | 16 (1 ) | 100% | 57.1 | 58.7 | 60.6 | 70.0 | 57.8 | 54.0 | 58.8 | 100.0 |
| V-CAST (2026’03) | 61.8 | 65.9 | 66.1 | 77.7 | 64.1 | 56.4 | 64.6 | 109.9 | ||
| GleanVID (Ours) | 80 (5 ) | 20% | 61.9 | 66.6 | 66.3 | 77.7 | 64.2 | 56.9 | 64.9 | 110.4 |
| V-CAST (2026’03) | 61.0 | 67.3 | 65.9 | 78.0 | 64.4 | 55.3 | 64.7 | 110.1 | ||
| GleanVID (Ours) | 160 (10 ) | 10% | 62.2 | 68.8 | 67.2 | 77.2 | 66.6 | 57.8 | 66.1 | 112.4 |
| Variant | LongVideoBench | VideoMME | |||
| Overall | Short | Medium | Long | ||
| V-CAST | 57.6 | 62.0 | 71.9 | 58.3 | 55.9 |
| TNB V-CAST budgeting | 58.0 | 62.1 | 72.3 | 58.9 | 55.1 |
| CTS V-CAST scoring | 58.1 | 61.8 | 72.6 | 58.4 | 54.3 |
| GleanVID | 59.0 | 62.6 | 73.6 | 59.6 | 54.6 |
| Rank | LongVideoBench | VideoMME | |||
| Overall | Short | Medium | Long | ||
| 58.6 | 62.4 | 73.2 | 59.1 | 54.8 | |
| 59.0 | 62.6 | 73.6 | 59.6 | 54.6 | |
| 58.8 | 62.5 | 73.3 | 59.4 | 54.8 | |
| 58.8 | 62.6 | 73.6 | 60.0 | 54.1 | |
| Rank | LongVideoBench | VideoMME | |||
| Overall | Short | Medium | Long | ||
| 58.6 | 62.4 | 73.2 | 59.1 | 54.8 | |
| 59.0 | 62.6 | 73.6 | 59.6 | 54.6 | |
| 58.8 | 62.5 | 73.3 | 59.4 | 54.8 | |
| 58.8 | 62.6 | 73.6 | 60.0 | 54.1 | |
| History | LongVideoBench | VideoMME | |||
| Overall | Short | Medium | Long | ||
| 58.6 | 62.1 | 72.7 | 59.9 | 53.9 | |
| 58.9 | 61.7 | 72.7 | 58.9 | 53.7 | |
| 58.9 | 62.1 | 72.4 | 59.6 | 54.3 | |
| 59.0 | 62.6 | 73.6 | 59.6 | 54.6 | |
| Neighborhood | LongVideoBench | VideoMME | |||
| Overall | Short | Medium | Long | ||
| 57.9 | 62.2 | 72.3 | 59.3 | 55.0 | |
| 59.0 | 62.6 | 73.6 | 59.6 | 54.6 | |
| 58.1 | 62.4 | 73.2 | 59.6 | 54.3 | |
| Neighborhood | LongVideoBench | VideoMME | |||
| Overall | Short | Medium | Long | ||
| 57.9 | 62.2 | 72.3 | 59.3 | 55.0 | |
| 59.0 | 62.6 | 73.6 | 59.6 | 54.6 | |
| 58.1 | 62.4 | 73.2 | 59.6 | 54.3 | |
| Neighborhood | LongVideoBench | VideoMME | |||
| Overall | Short | Medium | Long | ||
| 58.8 | 62.5 | 73.4 | 59.7 | 54.3 | |
| 59.0 | 62.6 | 73.6 | 59.6 | 54.6 | |
| 58.4 | 62.4 | 73.2 | 59.6 | 54.3 | |
| Temperature | LongVideoBench | VideoMME | |||
| Overall | Short | Medium | Long | ||
| 58.3 | 61.6 | 72.3 | 58.7 | 53.8 | |
| 58.6 | 62.2 | 73.0 | 58.9 | 54.7 | |
| 59.0 | 62.6 | 73.6 | 59.6 | 54.6 | |
| 58.6 | 62.8 | 73.6 | 59.9 | 55.0 | |
| Temperature | LongVideoBench | VideoMME | |||
| Overall | Short | Medium | Long | ||
| 58.3 | 61.6 | 72.3 | 58.7 | 53.8 | |
| 58.6 | 62.2 | 73.0 | 58.9 | 54.7 | |
| 59.0 | 62.6 | 73.6 | 59.6 | 54.6 | |
| 58.6 | 62.8 | 73.6 | 59.9 | 55.0 | |
| LongVideoBench | VideoMME | ||||
| Overall | Short | Medium | Long | ||
| 58.9 | 62.6 | 73.7 | 59.3 | 54.9 | |
| 59.0 | 62.6 | 73.6 | 59.6 | 54.6 | |
| 58.7 | 62.4 | 73.2 | 59.8 | 54.2 | |
| Method | Prefill Latency (s) | LLM Generation Latency (s) | GPU Peak Memory (MB) | Performance |
| LLaVA-OneVision-7B | 99.0 | 180.8 | 21,969.0 | 58.4 |
| VidCom 2 (EMNLP’25) | 26.5 ( 73.2%) | 108.6 ( 39.9%) | 19,512.6 ( 11.2%) | 58.4 |
| FastVID (NeurIPS’25) | 78.8 ( 20.4%) | 160.4 ( 11.3%) | 20,803.5 ( 5.3%) | 58.3 ( 0.1) |
| FlashVID (ICLR’26) | 28.4 ( 71.3%) | 146.1 ( 19.2%) | 19,059.9 ( 13.2%) | 58.7 ( 0.3) |
| V-CAST (2026’03) | 27.0 ( 72.7%) | 113.8 ( 37.0%) | 19,512.6 ( 11.2%) | 58.2 ( 0.2) |
| GleanVID (Ours) | 26.7 ( 73.1%) | 122.8 ( 32.1%) | 19,512.6 ( 11.2%) | 59.7 ( 1.3) |