Video Large Language Models (Video LLMs) have achieved remarkable progress in video understanding, but their inference efficiency is constrained by the large number of visual tokens produced by long videos. Recent video token compression methods increasingly introduce sophisticated strategies for token selection, pruning, and merging. This raises a fundamental question: how much of compression performance can be obtained by simply preserving the structure encoded in the visual representations? We investigate this question with SimpleCluster, a simple and training-free baseline that performs position-aware cross-frame clustering in the visual feature space and represents each cluster using the mean of its original visual features. Extensive experiments across four video understanding benchmarks and three representative Video LLMs show that SimpleCluster achieves competitive or superior performance over recent compression methods across a wide range of token retention ratios, with particularly strong robustness under extremely low retention rates (e.g., 1%). To understand this behavior, we analyze the feature space preserved by different compression methods in terms of local approximation fidelity and global coverage. The results show that stronger downstream performance is consistently associated with better preservation of the original visual feature distribution, especially its global coverage. These findings highlight feature-space preservation as an important consideration for video token compression under highly constrained token budgets. Our code is available at https://github.com/xiaozhang79/SimpleCluster.
Figures & tables
Figure 1: Performance comparison on LLaVA-OneVision-7B, LLaVA-Video-7B, and InternVL3-2B. SimpleCluster remains competitive across different token retention ratios and is particularly robust under aggressive compression.
Figure 2: Overall pipeline of SimpleCluster. SimpleCluster compresses visual tokens in Video LLMs by performing cross-frame clustering in the visual feature space and representing each cluster with the mean of its original projected features.
Method
Before LLM Retained Ratio
MVBench ↑
EgoSchema ↑
LongVideo Bench ↑
VideoMME ↑
Avg. ↑
Score
%
LLaVA-OV-7B
100%
58.3
60.4
56.4
58.6
58.4
100.0
DyCoke [CVPR′25]
25%
53.1
59.5
49.5
54.3
54.1
92.6
VisionZip [CVPR′25]
25%
57.9
60.3
56.5
58.2
58.2
99.7
FastVID [NeurIPS′25]
25%
56.5
58.2
56.3
58.0
57.3
98.1
EarlyTom [CVPR′26]
25%
57.4
60.5
56.3
58.5
58.2
99.7
Table 1: Comparison with state-of-the-art methods on LLaVA-OneVision-7B. “Avg. Score” and “Avg. %” denote the average score across benchmarks and the relative performance to the unpruned model. FastVID and AOT are omitted at 1% because their structural constraints require more tokens than the available budget. The best and second-best results are in bold and underlined , respectively.
Method
Before LLM Retained Ratio
MVBench ↑
EgoSchema ↑
LongVideo Bench ↑
VideoMME ↑
Avg. ↑
Score
%
LLaVA-Video-7B
100%
60.4
57.2
58.9
64.3
60.2
100.0
VisionZip [CVPR′25]
25%
56.7
54.7
54.7
60.7
56.7
94.2
EarlyTom [CVPR′26]
25%
59.1
55.3
57.3
61.9
58.4
97.0
AOT [CVPR′26]
25%
58.8
55.4
56.2
62.4
58.2
96.7
SimpleCluster
25%
60.5
54.8
58.1
62.7
59.0
98.0
Table 2: Comparison with state-of-the-art methods on LLaVA-Video-7B. “Avg. Score” and “Avg. %” denote the average score across benchmarks and the relative performance to the unpruned model. AOT is omitted at 1% because its structural constraints require more tokens than the available budget. The best and second-best results are in bold and underlined , respectively.
Method
Before LLM Retained Ratio
MVBench ↑
EgoSchema ↑
LongVideo Bench ↑
VideoMME ↑
Avg. ↑
Score
%
InternVL3-2B
100%
68.6
49.1
53.7
57.2
57.2
100.0
VisionZip [CVPR′25]
25%
67.3
47.9
53.3
56.9
56.3
98.4
EarlyTom [CVPR′26]
25%
68.8
47.8
53.5
55.0
55.8
97.6
AOT [CVPR′26]
25%
64.9
47.8
52.5
55.6
55.2
96.5
SimpleCluster
25%
67.1
47.6
54.9
57.8
56.9
99.5
Table 3: Comparison with state-of-the-art methods on InternVL3-2B. “Avg. Score” and “Avg. %” denote the average score across benchmarks and the relative performance to the unpruned model. AOT is omitted at 1% because its structural constraints require more tokens than the available budget. The best and second-best results are in bold and underlined , respectively.
Figure 3: Visualization of cross-frame token clustering in SimpleCluster. Different colors indicate different clusters.
Method
Before LLM Retained Ratio
MVBench ↑
EgoSchema ↑
LongVideo Bench ↑
VideoMME ↑
Avg. ↑
Spatial AvgPooling
10%
53.3
57.7
54.2
54.1
54.8
Spatiotemporal AvgPooling
10%
54.1
57.4
54.4
53.5
54.9
Spatial Grid Sampling
10%
53.9
58.5
54.8
55.0
55.6
Spatiotemporal Grid Sampling
10%
55.0
58.8
53.1
54.8
55.4
SimpleCluster
10%
56.8
59.9
56.5
56.0
57.3
Table 4: Comparison with simple token compression strategies on LLaVA-OneVision-7B.
Setting
Before LLM Retained Ratio
MVBench ↑
EgoSchema ↑
LongVideo Bench ↑
VideoMME ↑
Avg. ↑
None
10%
55.8
59.3
55.4
55.8
56.6
3D RoPE
10%
56.8
59.9
56.5
56.0
57.3
Per-frame
10%
55.0
59.3
54.1
55.8
56.1
Cross-frame
10%
56.8
59.9
56.5
56.0
57.3
Nearest
10%
57.0
59.7
55.7
56.4
57.2
Mean
10%
56.8
59.9
56.5
56.0
57.3
Table 5: Ablation studies of 3D RoPE encoding, clustering scope, and cluster representation on LLaVA-OneVision-7B. The default configuration is highlighted in gray.
Method
Prefilling FLOPs (T) ↓
TTFT (ms) ↓
Avg. ↑
Score
%
LLaVA-OneVision-7B
102.5
943.3
58.4
100.0
AOT
28.1
696.5
57.0
97.6
SimpleCluster
27.5
514.1
57.3
98.1
Table 6: Efficiency analysis of SimpleCluster on LLaVA-OneVision-7B with 10% token retention.
Method
Before LLM Retention
L2-NQE ↓
Cosine Coverage@0.90 ↑
VisionZip
10%
0.554
19.3%
FastVID
10%
0.408
35.0%
EarlyTom
10%
0.407
28.6%
AOT
10%
0.410
22.1%
SimpleCluster
10%
0.281
59.7%
VisionZip
5%
0.690
8.2%
Table 7: Comparison of local fidelity and global coverage of the original visual feature space across different visual token compression methods on LLaVA-OneVision-7B.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Before LLM Retained Ratio
MVBench ↑
EgoSchema ↑
LongVideo Bench ↑
VideoMME ↑
Avg. ↑
Score
%
VisionZip [CVPR′25]
20%
57.7
59.8
55.2
57.9
57.7
98.8
FastVID [NeurIPS′25]
20%
56.3
57.9
57.1
57.9
57.3
98.1
EarlyTom [CVPR′26]
20%
57.8
60.6
55.6
58.0
58.1
99.3
AOT [CVPR′26]
20%
58.1
61.3
56.2
57.2
58.2
99.7
SimpleCluster
20%
57.0
59.8
56.6
57.6
57.8
98.8
Appendix
Table 8: Comparison with state-of-the-art methods on LLaVA-OneVision-7B. “Avg. Score” and “Avg. %” denote the average score across benchmarks and the relative performance to the unpruned model. The best and second-best results are in bold and underlined , respectively.
Method
Before LLM Retained Ratio
MVBench ↑
EgoSchema ↑
LongVideo Bench ↑
VideoMME ↑
Avg. ↑
Score
%
LLaVA-Video-7B
100%
60.4
57.2
58.9
64.3
60.2
100.0
VisionZip [CVPR′25]
20%
60.5
55.5
57.1
62.0
58.8
97.7
EarlyTom [CVPR′26]
20%
59.2
54.5
54.4
61.3
57.4
95.3
AOT [CVPR′26]
20%
61.0
55.2
56.5
62.0
58.7
97.5
SimpleCluster
20%
60.3
54.1
57.7
61.6
58.4
97.0
Appendix
Table 9: Comparison with state-of-the-art methods on LLaVA-Video-7B. “Avg. Score” and “Avg. %” denote the average score across benchmarks and the relative performance to the unpruned model. The best and second-best results are in bold and underlined , respectively.
Method
Before LLM Retained Ratio
MVBench ↑
EgoSchema ↑
LongVideo Bench ↑
VideoMME ↑
Avg. ↑
Score
%
InternVL3-2B
100%
68.6
49.1
53.7
57.2
57.2
100.0
VisionZip [CVPR′25]
20%
65.2
47.4
52.5
57.0
55.5
97.0
EarlyTom [CVPR′26]
20%
65.7
47.1
50.7
54.9
54.6
95.5
AOT [CVPR′26]
20%
65.6
44.5
49.8
51.3
52.8
92.3
SimpleCluster
20%
66.7
47.2
53.4
57.4
56.2
98.3
Appendix
Table 10: Comparison with state-of-the-art methods on InternVL3-2B. “Avg. Score” and “Avg. %” denote the average score across benchmarks and the relative performance to the unpruned model. The best and second-best results are in bold and underlined , respectively.
Method
Prefilling FLOPs (T)
TTFT (ms)
Avg.
Score
%
LLaVA-Video-7B
218.1
1722.3
60.2
100.0
SimpleCluster
62.8
1082.6
57.3
95.2
Appendix
Table 11: Efficiency analysis of SimpleCluster on LLaVA-Video-7B with 10% token retention.
Method
Prefilling FLOPs (T)
TTFT (ms)
Avg.
Score
%
LLaVA-Video-7B
218.1
1722.3
60.2
100.0
SimpleCluster
62.8
1082.6
57.3
95.2
Appendix
Table 11: Efficiency analysis of SimpleCluster on LLaVA-Video-7B with 10% token retention.
Method
Prefilling FLOPs (T)
TTFT (ms)
Avg.
Score
%
InternVL3-2B
227.3
1079.2
57.2
100.0
SimpleCluster
76.2
862.6
55.1
96.3
Appendix
Table 12: Efficiency analysis of SimpleCluster on InternVL3-2B with 10% token retention.
Figure 4: Cross-frame cluster assignments produced by SimpleCluster on four videos.
Figure 5: Dense-feature coverage under nominal 5% token retention. Each heatmap cell shows the maximum cosine similarity between an original dense visual token and the compressed tokens produced by the corresponding method. All methods and frames share the same color scale, with brighter colors indicating better feature-space coverage. The maps measure representational proximity rather than attention or task relevance.
Video Large Language Models (Video-LLMs) achieve strong performance in video understanding, but their excessive visual tokens bring substantial computational overhead. Existing training-free compression methods improve inference efficiency by reducing visual tokens, yet they often rely on local adjacent-frame similarity for temporal redundancy estimation or allocate token budgets mainly according to segment length. Such designs are sensitive to frame-level noise and fail to capture the non-uniform information distribution of real-world videos. To address these challenges, we propose InfoMerge, a training-free visual token compression method that improves token utilization through robust redundancy estimation and content-aware budget allocation. Specifically, we propose the Temporal Fingerprint Difference: a segment-level second-order temporal redundancy estimation strategy, which models the temporal similarity structure of tokens at the same spatial positions within each segment. We further introduce Content-Aware Budget Allocation (CABA), which dynamically allocates segment-level token budgets based on segment uniqueness and spectral-entropy-based representational richness. By reducing repeated preservation of redundant static regions and allocating more tokens to informative segments, InfoMerge makes better use of the limited token budget while maintaining strong performance. Extensive experiments show that InfoMerge achieves strong efficiency--accuracy trade-offs across multiple benchmarks and backbones, with more pronounced advantages under aggressive compression. On LLaVA-OneVision-7B, InfoMerge retains 98.8% of the original average performance while reducing 85% of visual tokens and achieving a 4.24-fold speedup in the prefill stage.
Xinxin Liu, Shiwei Gan, Xiao Liu +3
State Key Laboratory of Novel Software Technology, Nanjing University Nanjing 210023, China
Video large language models (Video-LLMs) have demonstrated strong capabilities in video understanding tasks. However, their practical deployment is still hindered by the inefficiency introduced by processing massive amounts of visual tokens. Although recent approaches achieve extremely low token retention ratios while maintaining accuracy comparable to full-token baselines, most of them perform compression only at the late stage of prefilling, leaving the efficiency of the vision encoder unoptimized. In this paper, we first show that vision encoding contributes a large portion to the time-to-first-token (TTFT). Therefore, instead of compressing visual tokens only after the vision encoder, performing compression inside the encoder still leaves substantial room for exploration. Based on this insight, we propose EarlyTom, a training-free token compression framework that performs early-stage visual token compression inside the vision encoder, enabling significantly better TTFT reduction and higher throughput. In addition, we introduce a decoupled spatial token selection strategy that improves the overall compression effectiveness. EarlyTom reduces TTFT by up to 2.65x and FLOPs by up to 61% on a single NVIDIA A100 GPU for the LLaVA-OneVision-7B model, while maintaining accuracy comparable to the full-token baseline. These improvements substantially enhance the practicality of deploying Video-LLMs in real-world production scenarios.
Hesong Wang, Xin Jin, Lu Lu +4
Zhejiang University · Westlake University · Alibaba Cloud Computing
Video large language models (Video-LLMs) represent videos as dense sequences of visual tokens, whose length grows with the temporal and spatial extent of the input. These tokens often contain substantial redundancy arising from repeated visual patterns, leading to unnecessary computation in the subsequent language-model processing. Existing token compression methods, including pruning and merging, perform compression online during inference, repeatedly incurring additional computation for each input video and often relying on model-specific designs that limit their generality, we instead rethink this paradigm by shifting the costly compression process offline. We propose \textbf{ONCE}, a plug-in video token compression framework that introduces an offline-to-online paradigm: a frequency-aware global codebook is learned once in the visual feature space and reused for lightweight online compression through codebook lookup and aggregation, reducing repeated per-video computation and the need for model-specific compression designs. Extensive experiments across multiple video understanding benchmarks and against diverse compression baselines demonstrate that our approach achieves a strong accuracy-efficiency trade-off, maintaining competitive performance while achieving the lowest inference latency among compared methods.
Jiayang He, Tianling Xu, Diancheng Kang +4
Southern University of Science and Technology · Imperial College London · Jilin University