Visual token pruning speeds up multimodal large language models (MLLMs) by keeping a small subset of the visual tokens, and pruning methods are compared by the accuracy they retain. We study what pruning does to the confidence of these models, across common selectors, several MLLMs, and different output formats. Pruning errors concentrate on questions whose evidence the selector removed, and confidence does not register the removal. When the queried object loses all of its tokens, accuracy on these questions drops from 59% to 17%, while confidence stays at the unpruned level. Returning a few object tokens to the kept set recovers most of the lost accuracy. Selectors that keep the most attended tokens remove such evidence most often and produce confident errors, which temperature scaling cannot re-rank. Selectors that avoid keeping redundant tokens stay close to the calibration of the unpruned model. We then use the confidence of the pruned model to set a per-question token budget. The model answers with few tokens first and again with all tokens when its confidence is low. Conformal risk control sets the threshold to bound the expected deviation from the unpruned model. With coverage-based selection, this cascade needs about a third of the prefill tokens of the unpruned model, while with FastV it needs more than the unpruned model. The savings come mainly from how well the confidence ranks the answers that differ from the unpruned ones.
Figures & tables
Family
Selector
Ranking signal
Stage
Attention
CLS attention
class-token attention
input
VisionZip
class-token attn., merging
input
FastV
last-token attn., layer 2
LLM
PDrop
last-token attn., 3 stages
LLM
Coverage / diversity
Coverage ( α=0 )
facility location, Eq. 1
input
SCOPE ( α=1 )
coverage gain × attn.
input
Table 1: The nine token selectors, grouped by the signal that ranks tokens. “Input” selectors act before the language model; “LLM” selectors remove tokens after the first layers of the language model.
POPE (first-token confidence)
GQA (sequence confidence)
K=64
K=128
K=64
K=128
Family
Selector
Acc
ECE ↓
AUROC
Acc
ECE ↓
AUROC
Acc
ECE ↓
AUROC
Acc
ECE ↓
AUROC
Unpruned ( N=576 )
87.0
.040
.801
87.0
.040
.801
62.0
.053
.801
62.0
.053
.801
CLS attention
80.5
.090
.764
84.5
.050
.805
55.1
.088
.765
57.9
.072
.779
VisionZip
80.7
.075
.776
84.8
.046
.807
55.1
.087
.767
57.6
.073
.778
FastV
71.1
.214
.674
77.8
.146
.715
51.6
.115
.763
55.7
.098
.773
Table 2: Selectors on LLaVA-1.5-7B with 64 and 128 of the 576 visual tokens: accuracy (%), ECE, and AUROC. POPE uses the first-token confidence and GQA the probability of the whole answer. Bold: best pruned selector per column.
Figure 2 : Removed evidence on POPE (LLaVA-1.5-7B, questions about present objects). (a) Share of questions whose object loses all of its tokens. (b) Accuracy and mean confidence on these questions; ticks: questions whose object keeps some tokens. (c) Mean change of confidence on the same question from the unpruned to the pruned model.
Model
Unpr.
Cov.
DivP.
Attn.
FastV
Rand.
POPE, ECE ↓
LLaVA-1.5-13B
.032
.017
.017
.054
.107
.052
Qwen2.5-VL-3B
.067
.074
.063
.080
.094
.078
Qwen2.5-VL-7B
.082
.084
.086
.097
.126
.107
Qwen2.5-VL-32B
.085
.109
.102
.103
.188
.125
GQA, ECE ↓
Table 3: Other models at a keep ratio of 22%. For LLaVA-1.5-13B, Attn. is CLS attention; for Qwen2.5-VL, it is the attention received in the last vision block. GQA uses the probability of the whole answer. Bold: best pruned selector in each row.
ECE ↓
AUROC ↑
Selector
raw
Tunpr.
Town
raw
+answer
Unpruned
.040
.014
.014
.801
.848
Coverage
.025
.023
.017
.810
.843
DivPrune
.009
.038
.011
.812
.835
CDPruner
.007
.039
.008
.819
.821
CLS attention
.090
.049
.038
.764
.829
Table 4: Post-hoc recalibration on POPE (LLaVA-1.5-7B, 64 tokens), fitted on half of the images and evaluated on the other half. Tunpr. : temperature fitted to the unpruned model; Town : temperature fitted to each pruned model; both leave AUROC unchanged. +answer: logistic recalibration that also receives the predicted answer.
Figure 3 : Confidence-gated token budget with two stages. Conformal risk control chooses τ^ once on calibration images. Each question is answered with K1 visual tokens and, if its confidence is below τ^ , again with all tokens, reusing the visual tokens of the first pass.
Model
Bench.
Cov.
DivP.
CDP.
Attn.
FastV
Rand.
LLaVA-1.5-7B
POPE
0.36
0.34
0.32
0.73
1.05
0.65
GQA
0.80
0.80
0.72
0.89
1.08
0.93
LLaVA-1.5-13B
POPE
0.38
0.36
0.27
0.72
0.97
0.65
GQA
0.87
0.88
0.78
0.94
1.03
0.97
Qwen2.5-VL-3B
POPE
0.45
0.50
–
0.62
0.85
0.68
GQA
0.86
0.97
–
1.01
1.11
1.04
Table 5: Prefill cost (unpruned model =1 ) of the two-stage cascade at ε=2% , with 11% of the visual tokens in the first stage. Attn.: CLS attention on LLaVA-1.5, last-block attention on Qwen2.5-VL. Bold: cheapest selector.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : Reliability diagrams on POPE (LLaVA-1.5-7B, 64 tokens). Bars: accuracy per confidence bin; gray: share of answers per bin; dashed: perfect calibration.
Figure 5 : POPE questions answered by LLaVA-1.5-7B with all tokens and with 64 tokens kept by CLS attention or coverage selection. Red outline: queried object; dimmed cells: pruned tokens. Below each panel: answer and confidence (green: correct, red: wrong).
Figure 6 : Prefill cost at which the cascade keeps the expected deviation from the unpruned LLaVA-1.5-7B below ε (first stage: 64 tokens). Solid: coverage- and diversity-based selectors; dashed: attention-ranked selectors; dotted: random.
Selector
Kept set
n
Retention
Acc
Conf.
“no”
CLS attention
recorded
504
0.0
12.5
.816
87.5
restored
90.7
61.7
.770
38.3
control
0.0
13.3
.812
86.7
FastV
recorded
1260
0.0
5.0
.910
95.0
restored
87.0
64.6
.825
35.4
control
0.0
5.2
.898
94.8
Appendix
Table 6: Object-token intervention on POPE (LLaVA-1.5-7B, 64 tokens, questions about present objects). Retention: share of the object area on kept tokens (%). Acc and “no”: share of correct and of “no” answers (%).
ECE
AUROC
Condition
first
seq.
agree 10
first
seq.
agree 10
CLS attention, 128
.083
.072
.090
.768
.773
.736
CLS attention, 64
.100
.088
.104
.753
.760
.733
Coverage, 128
.066
.056
.093
.762
.771
.730
Coverage, 64
.076
.064
.086
.764
.772
.733
DivPrune, 128
.062
.053
.075
.771
.778
.748
Appendix
Table 7: Confidence scores on the first 3,000 GQA questions (LLaVA-1.5-7B): first-token probability, probability of the greedy answer, and agreement of ten answers sampled at temperature one with the greedy answer.
Latency (ms)
Selection (ms)
Selector
64
128
64
128
Unpruned
52.7
–
CLS attention
53.6
52.4
0.2
0.2
VisionZip
46.0
45.3
–
–
FastV
45.4
45.3
–
–
PDrop
48.6
48.3
–
–
Appendix
Table 8: Median time to answer one POPE question with LLaVA-1.5-7B at batch size one on an H200 GPU, including the vision encoder and decoding, and mean time of the token selection itself. Selection time is listed for the input-level selectors that run a separate selection step.
Visual tokens
Batch
64
128
192
576
1
44.9
41.6
37.6
20.6
8
97.5
70.5
53.3
21.9
32
110.5
73.0
55.5
22.1
64
112.3
74.8
56.1
22.2
Appendix
Table 9: Prefill throughput (questions per second) of the language model of LLaVA-1.5-7B on an A100 GPU for 61 text tokens and a given number of visual tokens. The vision encoder is excluded because pruning does not change its cost.
Pipeline
Vision
Select
Prefill
Prefill esc
Total
q/s
Unpruned
40
0
436
0
477
18.9
64 tokens
40
6
84
0
130
69.0
Cascade
40
6
84
76
207
43.5
Appendix
Table 10: End-to-end time (s) for the 9000 POPE questions with LLaVA-1.5-7B in batches of 32 on an A100 GPU. Select: coverage selection of 64 tokens; Prefill esc : second pass of the escalated questions with all tokens; q/s: questions per second.
AUROC agr
Cost
Selector
conf.
learned
conf.
learned
CLS attention
.777
.818
0.78
0.68
VisionZip
.807
.837
0.70
0.63
FastV
.667
.802
1.08
0.91
PDrop
.747
.801
0.99
0.91
Coverage ( α=0 )
.921
.921
0.39
0.39
Appendix
Table 11: Label-free learned gate on POPE (LLaVA-1.5-7B, 64 tokens, ε=2% ). The gate is fitted on a quarter of the images, calibrated on another quarter, and evaluated on the remaining half, averaged over 200 splits; the raw confidence uses the same calibration and test images.
Run
Mentions
Recall
CHAIR i
CHAIR s
ECE
AUROC
CDPruner, 128
2326
76.8
28.5
66.8
.194
.696
CDPruner, 64
2304
76.1
28.5
67.2
.198
.687
CLS attention, 128
2032
71.7
23.6
54.4
.242
.706
CLS attention, 64
1864
65.0
24.5
54.1
.241
.692
Coverage, 128
2284
76.3
27.6
65.4
.209
.694
Coverage, 64
2139
72.3
26.7
62.8
.202
.696
Appendix
Table 12: Object mentions in detailed descriptions of the 500 POPE images (LLaVA-1.5-7B). Recall: share of ground-truth objects mentioned. CHAIR i : share of mentions that are hallucinated; CHAIR s : share of descriptions with at least one hallucination. ECE and AUROC of the first-token probability of each mention.
POPE
GQA
cost with score
cost with score
Model
Selector
D
AUROC agr
conf.
random
perfect
share
D
AUROC agr
conf.
random
perfect
share
LLaVA-1.5-7B
CLS attention
11.8
.777
0.73
1.07
0.31
44.8
24.7
.777
0.89
1.14
0.45
35.7
VisionZip
11.4
.807
0.65
1.06
0.31
55.1
24.6
.775
0.89
1.15
0.44
35.9
FastV
21.0
.667
1.05
1.18
0.46
17.7
32.0
.757
1.08
1.20
0.57
19.2
PDrop
23.3
.747
0.96
1.13
0.43
24.5
36.4
.750
1.04
1.16
0.56
20.2
Appendix
Table 13 : Two-stage cascade at ε=2% (first stage: 11% of the visual tokens). D : share of questions (%) on which the first stage disagrees with the unpruned model; AUROC agr : AUROC of the confidence for agreement. Cost with the confidence, with an uninformative random score, and with a gate that escalates exactly the disagreements; share: part of the saving between the random and the perfect score that the confidence reaches (%).
POPE
GQA
ScienceQA-IMG
Selector
11%
22%
M
IA
11%
22%
M
IA
11%
22%
M
IA
CLS attention
0.73
0.55
0.95
0.80
0.89
0.86
1.32
0.94
0.65
0.66
0.93
0.80
VisionZip
0.65
0.52
0.83
–
0.89
0.87
1.33
–
0.66
0.63
0.93
–
FastV
1.05
1.04
1.57
–
1.08
1.07
1.69
–
0.73
0.62
1.04
–
PDrop
0.96
0.70
1.12
–
1.04
1.01
1.53
–
0.82
0.80
1.25
–
Coverage ( α=0 )
0.36
0.40
0.44
0.77
0.80
0.78
1.18
0.93
0.68
0.72
1.03
0.93
Appendix
Table 14: Prefill cost relative to the unpruned LLaVA-1.5-7B at ε=2% . 11%, 22%: two-stage cascade whose first stage keeps this share of the visual tokens; M: all three budgets in increasing order before the unpruned model; IA: input-adaptive budget chosen from the coverage of the kept set.
keep 11%
keep 22%
keep 33%
Selector
Acc
ECE
AUROC
CER .9
Acc
ECE
AUROC
CER .9
Acc
ECE
AUROC
CER .9
POPE, first-token confidence
Unpruned
87.0
.040
.801
.057
87.0
.040
.801
.057
87.0
.040
.801
.057
CLS attention
80.5
.090
.764
.097
84.5
.050
.805
.063
86.4
.038
.805
.055
VisionZip
80.7
.075
.776
.088
84.8
.046
.807
.056
86.5
.033
.812
.048
FastV
71.1
.214
.674
.237
77.8
.146
.715
.158
81.4
.109
.729
.127
Appendix
Table 15 : LLaVA-1.5-7B: every selector and keep ratio (POPE, first-token confidence; GQA, probability of the whole answer). Accuracy in %, ECE, AUROC, and the share of answers with confidence of at least 0.9 that are wrong (CER .9 ). The unpruned row is repeated in every column group. Cell colours mark the selector families.
keep 11%
keep 22%
keep 33%
Selector
Acc
ECE
AUROC
CER .9
Acc
ECE
AUROC
CER .9
Acc
ECE
AUROC
CER .9
GQA, first-token confidence
Unpruned
62.0
.065
.792
.079
62.0
.065
.792
.079
62.0
.065
.792
.079
CLS attention
55.1
.100
.758
.127
57.9
.084
.770
.110
59.4
.075
.778
.090
VisionZip
55.1
.099
.760
.127
57.6
.085
.769
.112
59.1
.075
.777
.093
FastV
51.6
.126
.756
.183
55.7
.110
.765
.137
57.9
.097
.775
.119
Appendix
Table 16 : LLaVA-1.5-7B: every selector and keep ratio (GQA, first-token confidence; ScienceQA-IMG, first-token confidence). Accuracy in %, ECE, AUROC, and the share of answers with confidence of at least 0.9 that are wrong (CER .9 ). The unpruned row is repeated in every column group. Cell colours mark the selector families.
keep 11%
keep 22%
Selector
Acc
ECE
AUROC
CER .9
Acc
ECE
AUROC
CER .9
POPE, first-token confidence
Unpruned
87.1
.032
.830
.047
87.1
.032
.830
.047
CLS attention
80.2
.101
.765
.109
84.4
.054
.805
.065
VisionZip
80.2
.099
.783
.104
84.7
.048
.809
.061
FastV
76.3
.164
.726
.167
81.5
.107
.760
.112
Appendix
Table 17 : LLaVA-1.5-13B: every selector and keep ratio. Accuracy in %, ECE, AUROC, and the share of answers with confidence of at least 0.9 that are wrong (CER .9 ). The unpruned row is repeated in every column group. Cell colours mark the selector families.
keep 11%
keep 22%
keep 33%
Selector
Acc
ECE
AUROC
CER .9
Acc
ECE
AUROC
CER .9
Acc
ECE
AUROC
CER .9
POPE, first-token confidence
Unpruned
87.9
.067
.826
.067
87.9
.067
.826
.067
87.9
.067
.826
.067
CLS attention
82.7
.106
.778
.111
85.7
.080
.818
.083
87.3
.070
.824
.074
FastV
79.5
.149
.748
.157
85.0
.094
.801
.095
87.2
.077
.807
.082
Coverage ( α=0 )
83.7
.086
.821
.081
85.9
.074
.822
.076
87.2
.068
.823
.069
Appendix
Table 18 : Qwen2.5-VL-3B: every selector and keep ratio. Accuracy in %, ECE, AUROC, and the share of answers with confidence of at least 0.9 that are wrong (CER .9 ). The unpruned row is repeated in every column group. Cell colours mark the selector families.
keep 11%
keep 22%
keep 33%
Selector
Acc
ECE
AUROC
CER .9
Acc
ECE
AUROC
CER .9
Acc
ECE
AUROC
CER .9
POPE, first-token confidence
Unpruned
87.8
.082
.822
.083
87.8
.082
.822
.083
87.8
.082
.822
.083
CLS attention
82.3
.124
.775
.126
85.9
.097
.809
.094
87.0
.088
.821
.083
FastV
77.3
.185
.728
.189
83.5
.126
.771
.125
85.6
.101
.806
.100
Coverage ( α=0 )
84.3
.097
.802
.092
86.3
.084
.809
.090
87.2
.082
.804
.084
Appendix
Table 19 : Qwen2.5-VL-7B: every selector and keep ratio. Accuracy in %, ECE, AUROC, and the share of answers with confidence of at least 0.9 that are wrong (CER .9 ). The unpruned row is repeated in every column group. Cell colours mark the selector families.
keep 11%
keep 22%
Selector
Acc
ECE
AUROC
CER .9
Acc
ECE
AUROC
CER .9
POPE, first-token confidence
Unpruned
86.8
.085
.810
.083
86.8
.085
.810
.083
CLS attention
78.7
.151
.748
.162
84.0
.103
.799
.103
FastV
64.8
.265
.721
.266
75.7
.188
.760
.184
Coverage ( α=0 )
81.8
.134
.787
.129
83.5
.109
.809
.103
Appendix
Table 20 : Qwen2.5-VL-32B: every selector and keep ratio. Accuracy in %, ECE, AUROC, and the share of answers with confidence of at least 0.9 that are wrong (CER .9 ). The unpruned row is repeated in every column group. Cell colours mark the selector families.
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · Nanyang Technological University · University of Electronic Science and Technology of China +1