Vision language models have achieved strong performance across a wide range of multimodal applications, yet their substantial computational and memory costs hinder efficient deployment. Visual token pruning and post-training quantization reduce inference overhead along two complementary dimensions, namely sequence length and numerical precision. Existing workflows typically optimize these techniques independently or apply them sequentially. Their distinct optimization objectives leave critical interactions unaddressed and constrain the achievable compression performance. We revisit these designs and present P4Q, a practical co-design framework that jointly optimizes visual token pruning and low-bit quantization for efficient VLM inference. First, P4Q introduces a quantization-aware visual token selection strategy before the LLM. It applies fake quantization to copies of the features produced by the projector and selects visual tokens using statistics computed from these fake-quantized features, thereby conditioning the selector's feature-based decisions on simulated low-bit perturbations. Second, P4Q introduces a pruning-aware quantization calibration strategy. It uses the same selection strategy as pruning to calibrate the quantized model on the retained-token distribution, thereby aligning the calibration process with the pruned execution path used during deployment. By coupling these two components, P4Q achieves substantial inference speedups while maintaining comparable task performance, resulting in a better efficiency-accuracy trade-off than independently optimized pipelines. For instance, on LLaVA-NeXT, P4Q achieves an average end-to-end inference speedup of 2.8x across eight distinct test sets, while retaining higher accuracy than prior compression and quantization methods.
Figures & tables
Figure 1: Comparison of sequential and joint pipelines for low-bit calibration and visual-token pruning, together with a visualization of the visual-token importance ranking shifts induced by low-bit quantization across LLM layers. (a) Prune-then-quantize suffers from quantization-induced importance-rank shifts, whereas quantize-then-prune creates a calibration-inference distribution shift. P4Q jointly prunes inference and calibration sequences using the same selector and token budget, matching their retained-token support and sequence length. (b) W4A4 induces layer-dependent visual token rank shifts relative to FP16 across representative in-LLM pruning layers in LLaVA-NeXT-7B.
Figure 2: Overview of quantization-aware pruning in P4Q. It first retains query-relevant tokens to capture task semantics, then selects visually salient tokens based on vision-tower CLS-to-patch attention, and finally applies max-min diversity sampling to reduce redundancy and broaden visual coverage. The query relevance and visual diversity stages use fake-quantized feature copies, while visual saliency uses vision tower attention.
Methods
MMMU
VizWiz
SQA
MME
GQA
MMB CN
POPE
SEED
Avg. ↑
E2E ↑
Mem. ↓
LLaVA-NeXT-7B
Upper Bound, 2880 Tokens (100%)
Vanilla
36.7
58.7
69.4
1725.1
63.9
58.8
86.5
69.1
100%
1 ×
100%
Pruning
Retain 320 Tokens ( ↓ 88.9%)
VisPruner (ICCV25)
35.3
56.2
66.8
1585.8
57.7
52.5
81.3
60.9
92.7%
1.63 ×
96.9%
VisionZip (CVPR25)
35.9
57.5
66.4
1616.9
58.1
54.3
79.3
62.9
93.9%
1.95 ×
98.6%
SpecFlow (ICML26)
36.6
57.8
67.7
1651.6
58.6
54.9
79.9
63.6
95.1%
1.62 ×
98.8%
Table 1: Main comparison of pruning, quantization and our joint method. E2E denotes the end-to-end speedup, and Mem. denotes the peak memory footprint. The best performance is marked in red .
Methods
MMMU
VizWiz
SQA
MMB CN
Avg. ↑
Qwen2.5-VL-7B
Upper Bound, 576 Tokens (100%)
Vanilla
51.4
71.2
87.7
81.6
100%
Retain 192 Tokens ( ↓ 66.7%)
SpecFlow (ICML26)
49.0
69.1
82.7
77.8
95.5%
P4Q (Ours)
50.7
69.4
83.5
78.4
96.8%
Retain 64 Tokens ( ↓ 88.9%)
Table 2: P4Q Generalization on Qwen2.5-VL-7B.
Methods
MME
GQA
POPE
SEED
Avg. ↑
Upper Bound, 576 Tokens (100%)
Vanilla
1650.7
60.6
81.8
63.8
100%
Retain 20% Tokens ( ↓ 80.0%) & W4A4
QUOTA (arXiv26)
90.82%
93.53%
97.56%
95.09%
94.3%
Retain 10% Tokens ( ↓ 90.0%) & W4A4
QUOTA (arXiv26)
90.32%
90.36%
96.01%
93.02%
92.4%
Table 3: Comparison on LLaVA-1.5-7B. QUOTA ( Li et al., 2026a ) reports performance retention (%) from its original paper.
Row
Quantization
Pruning
Avg. Accuracy Retention (%)
W4A4
PC
Pruning
FQP
1 (Vanilla)
–
–
–
–
100.0%
2
✓
–
✓
–
92.1%
3
✓
–
✓
✓
92.5%
4
✓
✓
✓
–
93.2%
5 (Ours)
✓
✓
✓
✓
95.5%
Table 4: Ablation study of quantization-aware token selection and pruning-aware calibration. PC is pruned calibration. FQP is fake quant for pruning.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Qualitative visualization of the W4A4-induced visual-token selection changes for a specific sample. Each column compares the independently selected FP16 and W4A4 Top- K sets at the same decoder layer. All settings retain 11.1% visual patch tokens. Bright regions correspond to retained visual tokens, whereas dark regions correspond to masked tokens. The values below the W4A4 panels report the fraction of the FP16 Top- K set replaced under W4A4: 10.0% , 28.1% , 33.1% , and 53.1% at Layers 2, 4, 8, and 16, respectively.
Row
Token-Budget
Benchmark
Avg. Accuracy Retention (%)
α
β
γ
MMMU
SQA
MME
Vanilla
–
–
–
34.8
66.3
1650.7
100.0%
1
0.42
0.28
0.30
34.3
65.5
1600.8
98.1%
2
0.20
0.30
0.50
34.9
66.0
1564.5
98.2%
3
0.40
0.10
0.50
34.7
65.7
1574.7
98.1%
4
0.12
0.18
0.70
33.3
66.3
1586.1
97.3%
Appendix
Table 5: Sensitivity analysis of the token-budget allocation hyperparameters on LLaVA-1.5-7B at an 88.9% visual-token pruning ratio. α , β , and γ denote the proportions allocated to text-guided anchors, visual-saliency anchors, and diversity completion.
Metric
Method
MMMU
VizWiz
SQA
MME
GQA
MMB CN
POPE
SEED
Avg.
LLaVA-NeXT-7B
E2E Latency
Vanilla
2144.33
2031.77
635.79
886.13
378.12
1044.38
552.55
367.08
1.00 ×
(ms/sample)
P4Q
707.36
806.95
246.78
271.57
165.86
367.40
179.10
130.63
2.80 ×
Peak Memory
Vanilla
20.31
14.88
14.95
14.90
14.90
15.14
14.89
14.95
100.0%
(GB)
P4Q
14.64
7.22
7.08
7.02
6.62
8.01
7.29
7.32
51.3%
LLaVA-1.5-7B
Appendix
Table 6: End-to-end inference latency and peak GPU memory of P4Q on LLaVA-NeXT-7B and LLaVA-1.5-7B. Latency is reported as the mean processing time per sample in milliseconds, and peak memory is reported in GB. For latency, the Avg. column reports the average dataset-wise speedup over Vanilla. For memory, it reports the average memory consumption relative to Vanilla. Lower latency and memory consumption are better.
VCIP, College of Computer Science, Nankai University · School of Computer Science and Engineering, Tianjin University of Technology · OPPO Research Institute +3