Video Large Language Models (Video LLMs) have achieved remarkable progress in video understanding, but their inference efficiency is constrained by the large number of visual tokens produced by long videos. Recent video token compression methods increasingly introduce sophisticated strategies for token selection, pruning, and merging. This raises a fundamental question: how much of compression performance can be obtained by simply preserving the structure encoded in the visual representations? We investigate this question with SimpleCluster, a simple and training-free baseline that performs position-aware cross-frame clustering in the visual feature space and represents each cluster using the mean of its original visual features. Extensive experiments across four video understanding benchmarks and three representative Video LLMs show that SimpleCluster achieves competitive or superior performance over recent compression methods across a wide range of token retention ratios, with particularly strong robustness under extremely low retention rates (e.g., 1%). To understand this behavior, we analyze the feature space preserved by different compression methods in terms of local approximation fidelity and global coverage. The results show that stronger downstream performance is consistently associated with better preservation of the original visual feature distribution, especially its global coverage. These findings highlight feature-space preservation as an important consideration for video token compression under highly constrained token budgets. Our code is available at https://github.com/xiaozhang79/SimpleCluster.
Figures & tables
Figure 1: Performance comparison on LLaVA-OneVision-7B, LLaVA-Video-7B, and InternVL3-2B. SimpleCluster remains competitive across different token retention ratios and is particularly robust under aggressive compression.
Figure 2: Overall pipeline of SimpleCluster. SimpleCluster compresses visual tokens in Video LLMs by performing cross-frame clustering in the visual feature space and representing each cluster with the mean of its original projected features.
Method
Before LLM Retained Ratio
MVBench ↑
EgoSchema ↑
LongVideo Bench ↑
VideoMME ↑
Avg. ↑
Score
%
LLaVA-OV-7B
100%
58.3
60.4
56.4
58.6
58.4
100.0
DyCoke [CVPR′25]
25%
53.1
59.5
49.5
54.3
54.1
92.6
VisionZip [CVPR′25]
25%
57.9
60.3
56.5
58.2
58.2
99.7
FastVID [NeurIPS′25]
25%
56.5
58.2
56.3
58.0
57.3
98.1
EarlyTom [CVPR′26]
25%
57.4
60.5
56.3
58.5
58.2
99.7
Table 1: Comparison with state-of-the-art methods on LLaVA-OneVision-7B. “Avg. Score” and “Avg. %” denote the average score across benchmarks and the relative performance to the unpruned model. FastVID and AOT are omitted at 1% because their structural constraints require more tokens than the available budget. The best and second-best results are in bold and underlined , respectively.
Method
Before LLM Retained Ratio
MVBench ↑
EgoSchema ↑
LongVideo Bench ↑
VideoMME ↑
Avg. ↑
Score
%
LLaVA-Video-7B
100%
60.4
57.2
58.9
64.3
60.2
100.0
VisionZip [CVPR′25]
25%
56.7
54.7
54.7
60.7
56.7
94.2
EarlyTom [CVPR′26]
25%
59.1
55.3
57.3
61.9
58.4
97.0
AOT [CVPR′26]
25%
58.8
55.4
56.2
62.4
58.2
96.7
SimpleCluster
25%
60.5
54.8
58.1
62.7
59.0
98.0
Table 2: Comparison with state-of-the-art methods on LLaVA-Video-7B. “Avg. Score” and “Avg. %” denote the average score across benchmarks and the relative performance to the unpruned model. AOT is omitted at 1% because its structural constraints require more tokens than the available budget. The best and second-best results are in bold and underlined , respectively.
Method
Before LLM Retained Ratio
MVBench ↑
EgoSchema ↑
LongVideo Bench ↑
VideoMME ↑
Avg. ↑
Score
%
InternVL3-2B
100%
68.6
49.1
53.7
57.2
57.2
100.0
VisionZip [CVPR′25]
25%
67.3
47.9
53.3
56.9
56.3
98.4
EarlyTom [CVPR′26]
25%
68.8
47.8
53.5
55.0
55.8
97.6
AOT [CVPR′26]
25%
64.9
47.8
52.5
55.6
55.2
96.5
SimpleCluster
25%
67.1
47.6
54.9
57.8
56.9
99.5
Table 3: Comparison with state-of-the-art methods on InternVL3-2B. “Avg. Score” and “Avg. %” denote the average score across benchmarks and the relative performance to the unpruned model. AOT is omitted at 1% because its structural constraints require more tokens than the available budget. The best and second-best results are in bold and underlined , respectively.
Figure 3: Visualization of cross-frame token clustering in SimpleCluster. Different colors indicate different clusters.
Method
Before LLM Retained Ratio
MVBench ↑
EgoSchema ↑
LongVideo Bench ↑
VideoMME ↑
Avg. ↑
Spatial AvgPooling
10%
53.3
57.7
54.2
54.1
54.8
Spatiotemporal AvgPooling
10%
54.1
57.4
54.4
53.5
54.9
Spatial Grid Sampling
10%
53.9
58.5
54.8
55.0
55.6
Spatiotemporal Grid Sampling
10%
55.0
58.8
53.1
54.8
55.4
SimpleCluster
10%
56.8
59.9
56.5
56.0
57.3
Table 4: Comparison with simple token compression strategies on LLaVA-OneVision-7B.
Setting
Before LLM Retained Ratio
MVBench ↑
EgoSchema ↑
LongVideo Bench ↑
VideoMME ↑
Avg. ↑
None
10%
55.8
59.3
55.4
55.8
56.6
3D RoPE
10%
56.8
59.9
56.5
56.0
57.3
Per-frame
10%
55.0
59.3
54.1
55.8
56.1
Cross-frame
10%
56.8
59.9
56.5
56.0
57.3
Nearest
10%
57.0
59.7
55.7
56.4
57.2
Mean
10%
56.8
59.9
56.5
56.0
57.3
Table 5: Ablation studies of 3D RoPE encoding, clustering scope, and cluster representation on LLaVA-OneVision-7B. The default configuration is highlighted in gray.
Method
Prefilling FLOPs (T) ↓
TTFT (ms) ↓
Avg. ↑
Score
%
LLaVA-OneVision-7B
102.5
943.3
58.4
100.0
AOT
28.1
696.5
57.0
97.6
SimpleCluster
27.5
514.1
57.3
98.1
Table 6: Efficiency analysis of SimpleCluster on LLaVA-OneVision-7B with 10% token retention.
Method
Before LLM Retention
L2-NQE ↓
Cosine Coverage@0.90 ↑
VisionZip
10%
0.554
19.3%
FastVID
10%
0.408
35.0%
EarlyTom
10%
0.407
28.6%
AOT
10%
0.410
22.1%
SimpleCluster
10%
0.281
59.7%
VisionZip
5%
0.690
8.2%
Table 7: Comparison of local fidelity and global coverage of the original visual feature space across different visual token compression methods on LLaVA-OneVision-7B.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Before LLM Retained Ratio
MVBench ↑
EgoSchema ↑
LongVideo Bench ↑
VideoMME ↑
Avg. ↑
Score
%
VisionZip [CVPR′25]
20%
57.7
59.8
55.2
57.9
57.7
98.8
FastVID [NeurIPS′25]
20%
56.3
57.9
57.1
57.9
57.3
98.1
EarlyTom [CVPR′26]
20%
57.8
60.6
55.6
58.0
58.1
99.3
AOT [CVPR′26]
20%
58.1
61.3
56.2
57.2
58.2
99.7
SimpleCluster
20%
57.0
59.8
56.6
57.6
57.8
98.8
Appendix
Table 8: Comparison with state-of-the-art methods on LLaVA-OneVision-7B. “Avg. Score” and “Avg. %” denote the average score across benchmarks and the relative performance to the unpruned model. The best and second-best results are in bold and underlined , respectively.
Method
Before LLM Retained Ratio
MVBench ↑
EgoSchema ↑
LongVideo Bench ↑
VideoMME ↑
Avg. ↑
Score
%
LLaVA-Video-7B
100%
60.4
57.2
58.9
64.3
60.2
100.0
VisionZip [CVPR′25]
20%
60.5
55.5
57.1
62.0
58.8
97.7
EarlyTom [CVPR′26]
20%
59.2
54.5
54.4
61.3
57.4
95.3
AOT [CVPR′26]
20%
61.0
55.2
56.5
62.0
58.7
97.5
SimpleCluster
20%
60.3
54.1
57.7
61.6
58.4
97.0
Appendix
Table 9: Comparison with state-of-the-art methods on LLaVA-Video-7B. “Avg. Score” and “Avg. %” denote the average score across benchmarks and the relative performance to the unpruned model. The best and second-best results are in bold and underlined , respectively.
Method
Before LLM Retained Ratio
MVBench ↑
EgoSchema ↑
LongVideo Bench ↑
VideoMME ↑
Avg. ↑
Score
%
InternVL3-2B
100%
68.6
49.1
53.7
57.2
57.2
100.0
VisionZip [CVPR′25]
20%
65.2
47.4
52.5
57.0
55.5
97.0
EarlyTom [CVPR′26]
20%
65.7
47.1
50.7
54.9
54.6
95.5
AOT [CVPR′26]
20%
65.6
44.5
49.8
51.3
52.8
92.3
SimpleCluster
20%
66.7
47.2
53.4
57.4
56.2
98.3
Appendix
Table 10: Comparison with state-of-the-art methods on InternVL3-2B. “Avg. Score” and “Avg. %” denote the average score across benchmarks and the relative performance to the unpruned model. The best and second-best results are in bold and underlined , respectively.
Method
Prefilling FLOPs (T)
TTFT (ms)
Avg.
Score
%
LLaVA-Video-7B
218.1
1722.3
60.2
100.0
SimpleCluster
62.8
1082.6
57.3
95.2
Appendix
Table 11: Efficiency analysis of SimpleCluster on LLaVA-Video-7B with 10% token retention.
Method
Prefilling FLOPs (T)
TTFT (ms)
Avg.
Score
%
LLaVA-Video-7B
218.1
1722.3
60.2
100.0
SimpleCluster
62.8
1082.6
57.3
95.2
Appendix
Table 11: Efficiency analysis of SimpleCluster on LLaVA-Video-7B with 10% token retention.
Method
Prefilling FLOPs (T)
TTFT (ms)
Avg.
Score
%
InternVL3-2B
227.3
1079.2
57.2
100.0
SimpleCluster
76.2
862.6
55.1
96.3
Appendix
Table 12: Efficiency analysis of SimpleCluster on InternVL3-2B with 10% token retention.
Figure 4: Cross-frame cluster assignments produced by SimpleCluster on four videos.
Figure 5: Dense-feature coverage under nominal 5% token retention. Each heatmap cell shows the maximum cosine similarity between an original dense visual token and the compressed tokens produced by the corresponding method. All methods and frames share the same color scale, with brighter colors indicating better feature-space coverage. The maps measure representational proximity rather than attention or task relevance.