Multimodal large language models (MLLMs) incur high inference latency from long visual token sequences. Existing pruning methods commonly use attention maps or output features to estimate token importance or redundancy. Several recent approaches also exploit representation changes, but when and how these changes reflect foreground saliency and semantic consistency remain insufficiently understood. We analyze visual token representation dynamics across encoder depth and uncover two findings. First, the relationship between token update magnitudes and foreground saliency is layer-dependent: large token updates concentrate on foreground regions in two depth intervals, separated by several sink-dominated layers at intermediate depths. Second, similarities between token update directions better distinguish same-class from different-class tokens than those between encoder output features. Building on these findings, we propose MSDG-Prune, a training-free method that uses update magnitudes and directions to preserve salient and diverse visual information. Specifically, we group tokens by update-direction similarity and use query-weighted saliency derived from update magnitudes across a chosen depth window for group-wise token pruning. Extensive experiments across four MLLMs demonstrate the effectiveness and generalizability of MSDG-Prune. On LLaVA-NeXT, it retains 91.9% of uncompressed performance on average with only 5.6% of visual tokens, while achieving a 7.8x prefilling speedup. Code is available at https://github.com/liweixuan-hitsz/MSDG-Prune.
Figures & tables
Figure 1: Output features (OF) versus update directions (UD) for semantic similarity in LLaVA-1.5. Panels (a) and (b) visualize the top 40 tokens ranked by cosine similarity to the airplane anchor token ( red ), using OF and UD, respectively. Panels (c) and (d) show cosine similarity distributions on COCO 2017 computed using OF and UD, respectively. For each anchor token, we compute cosine similarities to other tokens and group the resulting pairs by whether they share the same semantic class as the anchor token. Blue and gray curves represent same-class and different-class pairs, respectively. The horizontal axis indicates cosine similarity, and the vertical axis shows the percentage of pairs within each similarity bin. Compared with OF, UD shows less overlap between the two groups, indicating better semantic discrimination.
Figure 2: Layer-wise token update magnitudes in the vision encoder of LLaVA-1.5. Each map overlays patch-token update magnitudes across a Transformer block on the input image. L00-01 denotes the change from token states entering Layer 0 to those entering Layer 1. Because LLaVA-1.5 uses the penultimate-layer output of its CLIP ViT-L/14@336 ( Radford et al., 2021 ; Liu et al., 2024a ) vision encoder as input to the visual projector, the visualization ends at L22-23 and omits L23-24. The maps illustrate depth-dependent spatial patterns in token updates, including foreground enhancement before and after the sink-dominated stage.
Figure 3: Layer-wise analysis of token update magnitudes in LLaVA-1.5 on COCO 2017 validation images. Top: foreground recall of the top 30% of tokens ranked by update magnitude, with random selection as the dashed baseline. Bottom: the max-to-median update magnitude ratio, maxiviℓ/medianiviℓ , with a dashed diagnostic threshold of 10 . Both metrics are computed per image and then averaged across images. The highlighted Layers 11–12 exhibit sharp ratio peaks, marking the sink-dominated stage in the corresponding vision encoder.
Figure 4: Overview of MSDG-Prune. After sink filtering, update magnitudes over a selected foreground-enhanced window provide visual saliency scores, while update-direction similarity organizes tokens into semantic groups. Visual saliency and query relevance jointly guide budget allocation across groups and token selection within each group.
Method
GQA
MMB
MME
POPE
SQA
VQA v2
VQA Text
SEED I
VizWiz
RelAcc.
Uncompressed baseline (100%)
Vanilla CVPR 2024
61.9
64.7
1862
85.9
69.5
78.5
58.2
60.5
54.3
100%
Token retention: 33.3%
FastV ECCV 2024
52.7
61.2
1612
64.8
67.3
67.1
52.5
57.1
50.8
89.1%
SparseVLM ICML 2025
57.6
62.5
1721
83.6
69.1
75.6
56.1
55.8
50.5
95.2%
DART EMNLP 2025
60.0
63.6
1856
82.8
69.8
76.7
57.4
51.5
54.9
97.1%
Table 1: Performance comparison on LLaVA-1.5-7B. Vanilla refers to the uncompressed baseline. The subscript after each method gives the publication venue.
Table 6
Method
Tokens
Inference Time ↓
Prefilling Time ↓
POPE (F1) ↑
LLaVA-NeXT
2880
254 ms
218 ms
86.4
SparseVLM
160
199 ms
119 ms
69.3
VisionZip
160
84 ms
27.8 ms
74.8
PruneSID
160
89 ms
27.8 ms
76.9
Ours
160
96 ms
27.8 ms
87.1
Table 5: Inference latency and POPE F1 score on LLaVA-NeXT-7B. Latencies are averaged per sample using a single NVIDIA A800-80 GB GPU.
Method
Retain 33.3%
Retain 22.2%
Retain 11.1%
GQA
MME
POPE
SQA
RelAcc.
GQA
MME
POPE
SQA
RelAcc.
GQA
MME
POPE
SQA
RelAcc.
Ours
58.5
2332
84.7
87.8
98.5%
57.7
2297
84.0
87.3
97.4%
55.4
2178
81.3
86.3
94.1%
Output features
58.1
2286
84.0
87.5
97.5%
56.9
2254
82.8
87.1
96.2%
54.8
2178
80.1
86.0
93.5%
Attention
57.9
2319
86.1
87.9
98.5%
56.9
2268
84.4
87.6
97.0%
53.8
2176
81.3
86.5
93.5%
Accumulated updates
58.5
2317
84.4
87.9
98.3%
56.5
2272
83.2
87.3
96.4%
54.3
2129
80.1
85.9
92.7%
Keep sink tokens
58.4
2330
84.7
87.7
98.4%
57.1
2264
83.1
87.2
96.5%
54.8
2143
81.2
86.3
93.5%
Table 6: Component ablations and staged token selection on Qwen2.5-VL-7B. Ours is the full method; Output features uses encoder-output features for grouping. Query (global) ranks by query relevance without grouping; Query × Visual (global) ranks by the product of query relevance and visual saliency, still without grouping.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Algorithm 1 MSDG-Prune Visual Token Selection
Require: encoder states {hiℓ:1≤i≤N,0≤ℓ≤L} , projected visual tokens {xi}i=1N , where N is the number of visual tokens before pruning and L is the number of encoder layers, query embeddings {tj}j∈Q , where Q indexes the query tokens, budget B , group count K , ℓ0 and ℓd (start and end layers of the update direction di ), ℓ⋆ (peak concentration layer; Appendix C.1 ), ℓs and ℓe (start and end of the saliency window for the visual saliency score), diagnostic coordinate ds , threshold τ , and integer initialization seed s ; unit(z)=z/max{∥z∥,10−12} . TopIdxb returns the indices of the b largest scores, breaking ties by original token order.
1:
V←{i:∣(hiℓ⋆+1)ds∣≤τ} ; require B≤∣V∣ , 1≤K≤∣V∣ , and Q=∅
Figure 6: Cross-image sink signature on COCO 2017 validation images on LLaVA-NeXT and Qwen2.5-VL. Top: image-mean layer-wise foreground recall when selecting the top 30% of tokens by viℓ . Bottom: mean concentration ratio maxiviℓ/medianiviℓ ; the dashed line marks the threshold ( 5 on LLaVA-NeXT and 10 on Qwen2.5-VL). A ratio spike at a fixed intermediate depth indicates isolated tokens with large update magnitudes, consistent with the sink-dominated stage.
Figure 7: Image-mean Foreground Recall@30% for windows of width w=5 on COCO 2017 validation images. At each input-state start ℓs , patches are ranked by ∥hiℓs+5−hiℓs∥ . Horizontal dashed lines show the random selection expectation (approximately 30% ); vertical dotted lines mark the default starts used in the main experiments.
Figure 8: Spatial maps of update magnitudes viℓ across LLaVA-1.5 and Mini-Gemini vision encoder layers for the tennis scene.
Figure 9: Spatial maps of update magnitudes viℓ across LLaVA-NeXT vision encoder layers for the tennis scene.
Figure 10: Spatial maps of update magnitudes viℓ across LLaVA-NeXT vision encoder layers for the kite scene.
Figure 12: Spatial maps of update magnitudes viℓ across Qwen2.5-VL vision encoder layers for the tennis scene.
Figure 16
Method
GQA
MMB
MME
POPE
SQA
VQA v2
VQA Text
SEED I
VizWiz
RelAcc.
Uncompressed baseline (100%)
Vanilla
63.2
67.7
1818
85.9
72.8
80.0
61.3
66.9
56.6
100%
Token retention: 33.3%
VisionZip
59.6
65.9
1770
86.4
72.8
78.0
58.6
65.2
54.9
97.5%
PruneSID
59.6
65.9
1770
86.4
72.8
78.0
58.6
65.2
56.0
97.7%
Ours
60.1
67.4
1801
86.1
73.7
76.8
59.4
65.3
56.0
98.3%
Appendix
Table 9: Performance comparison on LLaVA-1.5-13B. RelAcc. averages the relative scores across the nine listed benchmarks.
Method
GQA
MMB
MME
POPE
SQA
VQA v2
VQA Text
SEED I
VizWiz
RelAcc.
Uncompressed baseline (100%)
Vanilla
65.4
70.0
1858
86.2
73.5
81.8
64.3
71.9
64.0
100%
Token retention: 22.2%
PruneSID
62.4
67.0
1817
85.6
70.1
79.1
60.2
68.5
60.2
95.9%
Ours
63.6
66.6
1820
87.0
72.8
78.6
62.0
68.5
60.1
96.9%
Token retention: 11.1%
Appendix
Table 10: Performance comparison on LLaVA-NeXT-13B.
Method
GQA
MMB
MME
POPE
SQA
MMMU
VQA Text
RelAcc.
Uncompressed baseline (100%)
Vanilla
59.3
86.3
2428
84.3
91.5
61.1
76.5
100%
Token retention: 33.3%
VisionZip
57.3
85.1
2307
82.1
88.9
57.7
72.1
96.2%
PruneSID
57.0
81.6
2209
82.7
85.4
57.6
70.6
94.2%
Ours
56.4
84.4
2396
81.6
89.1
58.3
72.9
96.6%
Appendix
Table 11: Performance comparison on Qwen2.5-VL-32B.
Method
Retain 33.3%
Retain 22.2%
Retain 11.1%
GQA
MME
POPE
SQA
RelAcc.
GQA
MME
POPE
SQA
RelAcc.
GQA
MME
POPE
SQA
RelAcc.
Ours
60.1
1806
86.9
68.7
98.5%
58.9
1778
86.7
68.7
97.6%
57.7
1708
85.6
68.4
95.8%
Grouping
Output features
60.0
1771
86.7
68.4
97.8%
58.9
1712
86.6
68.7
96.7%
57.8
1680
84.4
68.4
95.1%
Saliency scoring
Attention
58.7
1795
84.3
68.7
97.1%
57.9
1775
83.3
68.4
96.1%
56.4
1675
81.2
68.7
93.6%
Appendix
Table 12: Grouping, saliency scoring, and update aggregation on LLaVA-1.5-7B. Ours uses unit update directions and endpoint displacement. Output features replaces the grouping representation with hio . Attention replaces the visual saliency score ui with [CLS] attention from the output layer. Accumulated replaces endpoint displacement with accumulated update magnitude.
Retain 33.3%
Retain 22.2%
Retain 11.1%
ℓ0
GQA
MMB
MME
POPE
SQA
RelAcc.
GQA
MMB
MME
POPE
SQA
RelAcc.
GQA
MMB
MME
POPE
SQA
RelAcc.
Avg.
0
58.1
81.7
2323
84.9
88.0
98.2%
57.0
80.2
2300
83.7
87.1
96.8%
54.5
77.6
2165
81.3
86.4
93.4%
96.1%
1
58.2
82.0
2318
84.7
87.7
98.1%
57.3
80.8
2290
84.1
87.3
97.1%
54.8
77.7
2212
81.4
86.3
94.0%
96.4%
2
58.5
82.6
2332
84.7
87.8
98.5%
57.7
81.2
2297
84.0
87.3
97.3%
55.4
78.4
2178
81.3
86.3
94.0%
96.6%
3
58.2
82.0
2329
85.0
87.8
98.3%
57.5
80.7
2294
84.6
87.0
97.1%
55.4
77.1
2161
81.3
86.4
93.6%
96.3%
Appendix
Table 13: Ablation of the early endpoint ℓ0 on Qwen2.5-VL-7B. Avg. is the mean RelAcc. over the three retention ratios.
Retain 33.3%
Retain 22.2%
Retain 11.1%
K
GQA
MMB
MME
POPE
SQA
RelAcc.
GQA
MMB
MME
POPE
SQA
RelAcc.
GQA
MMB
MME
POPE
SQA
RelAcc.
Avg.
8
57.6
81.2
2276
83.3
87.6
97.0%
56.2
79.8
2263
81.8
86.7
95.5%
53.7
76.7
2105
78.8
85.3
91.6%
94.7%
12
57.8
80.8
2286
83.6
87.7
97.1%
56.2
79.7
2249
82.7
86.8
95.6%
54.4
75.8
2131
80.0
85.6
92.2%
95.0%
16
58.0
81.3
2302
84.4
87.8
97.7%
56.8
79.7
2230
82.5
87.2
95.7%
54.6
76.6
2130
80.8
85.7
92.6%
95.3%
20
58.5
82.6
2332
84.7
87.8
98.5%
57.7
81.2
2297
84.0
87.3
97.3%
55.4
78.4
2178
81.3
86.3
94.0%
96.6%
24
58.5
81.9
2361
85.3
87.8
98.7%
57.5
81.1
2311
83.8
87.3
97.3%
55.3
78.0
2170
81.4
86.3
93.8%
96.6%
Appendix
Table 15: Ablation of the group count K on Qwen2.5-VL.
Table 23
Window
GQA
MME
POPE
SQA
RelAcc.
Uncompressed baseline (100%)
Vanilla
60.9
2310
86.3
88.9
100%
Token retention: 22.2%
19→22 ( w=3 )
57.8
2252
84.2
87.1
97.0%
19→24 ( w=5 )
57.7
2297
84.0
87.3
97.4%
19→26 ( w=7 )
57.6
2273
84.0
87.4
97.2%
Appendix
Table 18: Saliency window width ablation on Qwen2.5-VL-7B with the start fixed at Layer 19 and all other components unchanged. An interval a→b covers updates from Layer a through Layer b−1 , so varying b changes the width while preserving the starting layer.
Method
GQA
MME
POPE
SQA
RelAcc.
Token retention: 22.2%
α Alloc., u Rank
57.4
2307
84.2
87.1
97.4%
u Alloc., α Rank
55.0
2264
81.1
86.6
94.9%
Ours
57.7
2297
84.0
87.3
97.4%
Token retention: 11.1%
α Alloc., u Rank
55.0
2166
81.5
86.5
94.0%
Appendix
Table 19: Ablation of signal assignment on Qwen2.5-VL-7B. Allocation and ranking refer to budget allocation across groups and token selection within each group, respectively.
Figure 16: Failure cases of MSDG-Prune on Qwen2.5-VL-7B at 11.1% target token retention. Each case pairs the original image (left) with the retained-token visualization (right). GT denotes the ground-truth answer and Pred denotes the answer after pruning; the uncompressed model answers both questions correctly. Discarded patches are covered by a translucent gray mask.
Figure 17: Qualitative comparison at a retention ratio of 11.1%. Columns show the original image and the tokens retained by VisionZip, PruneSID, and MSDG-Prune, with the corresponding query below each example. Discarded patches are covered by a translucent gray mask.
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · Nanyang Technological University · University of Electronic Science and Technology of China +1