Vision language models have achieved strong performance across a wide range of multimodal applications, yet their substantial computational and memory costs hinder efficient deployment. Visual token pruning and post-training quantization reduce inference overhead along two complementary dimensions, namely sequence length and numerical precision. Existing workflows typically optimize these techniques independently or apply them sequentially. Their distinct optimization objectives leave critical interactions unaddressed and constrain the achievable compression performance. We revisit these designs and present P4Q, a practical co-design framework that jointly optimizes visual token pruning and low-bit quantization for efficient VLM inference. First, P4Q introduces a quantization-aware visual token selection strategy before the LLM. It applies fake quantization to copies of the features produced by the projector and selects visual tokens using statistics computed from these fake-quantized features, thereby conditioning the selector's feature-based decisions on simulated low-bit perturbations. Second, P4Q introduces a pruning-aware quantization calibration strategy. It uses the same selection strategy as pruning to calibrate the quantized model on the retained-token distribution, thereby aligning the calibration process with the pruned execution path used during deployment. By coupling these two components, P4Q achieves substantial inference speedups while maintaining comparable task performance, resulting in a better efficiency-accuracy trade-off than independently optimized pipelines. For instance, on LLaVA-NeXT, P4Q achieves an average end-to-end inference speedup of 2.8x across eight distinct test sets, while retaining higher accuracy than prior compression and quantization methods.
Figures & tables
Figure 1: Comparison of sequential and joint pipelines for low-bit calibration and visual-token pruning, together with a visualization of the visual-token importance ranking shifts induced by low-bit quantization across LLM layers. (a) Prune-then-quantize suffers from quantization-induced importance-rank shifts, whereas quantize-then-prune creates a calibration-inference distribution shift. P4Q jointly prunes inference and calibration sequences using the same selector and token budget, matching their retained-token support and sequence length. (b) W4A4 induces layer-dependent visual token rank shifts relative to FP16 across representative in-LLM pruning layers in LLaVA-NeXT-7B.
Figure 2: Overview of quantization-aware pruning in P4Q. It first retains query-relevant tokens to capture task semantics, then selects visually salient tokens based on vision-tower CLS-to-patch attention, and finally applies max-min diversity sampling to reduce redundancy and broaden visual coverage. The query relevance and visual diversity stages use fake-quantized feature copies, while visual saliency uses vision tower attention.
Methods
MMMU
VizWiz
SQA
MME
GQA
MMB CN
POPE
SEED
Avg. ↑
E2E ↑
Mem. ↓
LLaVA-NeXT-7B
Upper Bound, 2880 Tokens (100%)
Vanilla
36.7
58.7
69.4
1725.1
63.9
58.8
86.5
69.1
100%
1 ×
100%
Pruning
Retain 320 Tokens ( ↓ 88.9%)
VisPruner (ICCV25)
35.3
56.2
66.8
1585.8
57.7
52.5
81.3
60.9
92.7%
1.63 ×
96.9%
VisionZip (CVPR25)
35.9
57.5
66.4
1616.9
58.1
54.3
79.3
62.9
93.9%
1.95 ×
98.6%
SpecFlow (ICML26)
36.6
57.8
67.7
1651.6
58.6
54.9
79.9
63.6
95.1%
1.62 ×
98.8%
Table 1: Main comparison of pruning, quantization and our joint method. E2E denotes the end-to-end speedup, and Mem. denotes the peak memory footprint. The best performance is marked in red .
Methods
MMMU
VizWiz
SQA
MMB CN
Avg. ↑
Qwen2.5-VL-7B
Upper Bound, 576 Tokens (100%)
Vanilla
51.4
71.2
87.7
81.6
100%
Retain 192 Tokens ( ↓ 66.7%)
SpecFlow (ICML26)
49.0
69.1
82.7
77.8
95.5%
P4Q (Ours)
50.7
69.4
83.5
78.4
96.8%
Retain 64 Tokens ( ↓ 88.9%)
Table 2: P4Q Generalization on Qwen2.5-VL-7B.
Methods
MME
GQA
POPE
SEED
Avg. ↑
Upper Bound, 576 Tokens (100%)
Vanilla
1650.7
60.6
81.8
63.8
100%
Retain 20% Tokens ( ↓ 80.0%) & W4A4
QUOTA (arXiv26)
90.82%
93.53%
97.56%
95.09%
94.3%
Retain 10% Tokens ( ↓ 90.0%) & W4A4
QUOTA (arXiv26)
90.32%
90.36%
96.01%
93.02%
92.4%
Table 3: Comparison on LLaVA-1.5-7B. QUOTA ( Li et al., 2026a ) reports performance retention (%) from its original paper.
Row
Quantization
Pruning
Avg. Accuracy Retention (%)
W4A4
PC
Pruning
FQP
1 (Vanilla)
–
–
–
–
100.0%
2
✓
–
✓
–
92.1%
3
✓
–
✓
✓
92.5%
4
✓
✓
✓
–
93.2%
5 (Ours)
✓
✓
✓
✓
95.5%
Table 4: Ablation study of quantization-aware token selection and pruning-aware calibration. PC is pruned calibration. FQP is fake quant for pruning.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Qualitative visualization of the W4A4-induced visual-token selection changes for a specific sample. Each column compares the independently selected FP16 and W4A4 Top- K sets at the same decoder layer. All settings retain 11.1% visual patch tokens. Bright regions correspond to retained visual tokens, whereas dark regions correspond to masked tokens. The values below the W4A4 panels report the fraction of the FP16 Top- K set replaced under W4A4: 10.0% , 28.1% , 33.1% , and 53.1% at Layers 2, 4, 8, and 16, respectively.
Row
Token-Budget
Benchmark
Avg. Accuracy Retention (%)
α
β
γ
MMMU
SQA
MME
Vanilla
–
–
–
34.8
66.3
1650.7
100.0%
1
0.42
0.28
0.30
34.3
65.5
1600.8
98.1%
2
0.20
0.30
0.50
34.9
66.0
1564.5
98.2%
3
0.40
0.10
0.50
34.7
65.7
1574.7
98.1%
4
0.12
0.18
0.70
33.3
66.3
1586.1
97.3%
Appendix
Table 5: Sensitivity analysis of the token-budget allocation hyperparameters on LLaVA-1.5-7B at an 88.9% visual-token pruning ratio. α , β , and γ denote the proportions allocated to text-guided anchors, visual-saliency anchors, and diversity completion.
Metric
Method
MMMU
VizWiz
SQA
MME
GQA
MMB CN
POPE
SEED
Avg.
LLaVA-NeXT-7B
E2E Latency
Vanilla
2144.33
2031.77
635.79
886.13
378.12
1044.38
552.55
367.08
1.00 ×
(ms/sample)
P4Q
707.36
806.95
246.78
271.57
165.86
367.40
179.10
130.63
2.80 ×
Peak Memory
Vanilla
20.31
14.88
14.95
14.90
14.90
15.14
14.89
14.95
100.0%
(GB)
P4Q
14.64
7.22
7.08
7.02
6.62
8.01
7.29
7.32
51.3%
LLaVA-1.5-7B
Appendix
Table 6: End-to-end inference latency and peak GPU memory of P4Q on LLaVA-NeXT-7B and LLaVA-1.5-7B. Latency is reported as the mean processing time per sample in milliseconds, and peak memory is reported in GB. For latency, the Avg. column reports the average dataset-wise speedup over Vanilla. For memory, it reports the average memory consumption relative to Vanilla. Lower latency and memory consumption are better.
Deploying Vision-Language Models (VLMs) under aggressive low-bit inference remains challenging because inference cost is dominated by the long visual-token prefix during prefill and the growing KV cache during autoregressive decoding. Token pruning and low-bit quantization are complementary for reducing these costs, yet naive stage-wise combinations are often brittle due to a mismatch between quantization calibration and pruning execution. We present a collaborative quantization-and-pruning framework that unifies low-bit inference and deterministic visual-token pruning in a single deployable pipeline. The framework introduces the \textbf{Q}uantization \textbf{U}nified \textbf{O}ffline \textbf{T}oken \textbf{A}llocator (\textbf{QUOTA}), which converts low-bit calibration signals into a layer-wise token allocation schedule and materializes it as a pruning recipe. Token importance is evaluated under deployed W4A4 operators with a quantized KV cache by combining activation magnitude, attention cues, and an explicit low-bit risk signal, enabling consistent budgeted top-k selection. Experiments on standard VLM benchmarks show improved robustness over stage-wise baselines under the same low-bit regime, achieving 95.65% average retention while retaining only 30% of visual tokens, compared with about 94.3% retention for representative stage-wise combinations. The code will be released.
Xinqing Li, Xin He, Xindong Zhang +3
VCIP, College of Computer Science, Nankai University · School of Computer Science and Engineering, Tianjin University of Technology · OPPO Research Institute +3
Visual token pruning is a promising approach to reducing the inference cost of large vision-language models (LVLMs), yet aggressive token reduction often causes substantial performance degradation. We identify three key factors behind this degradation: text-guided selection bias, information loss from discarded tokens, and positional distortion caused by sequence compaction. Based on these observations, we propose \textbf{VPRune}, a training-free pre-LLM pruning framework consisting of visual-only diversity selection, similarity-guided token recycling, and position-preserving restoration. Experiments on FastVLM-1.5B across multiple vision-language benchmarks demonstrate that VPRune achieves a favorable accuracy--compression trade-off, with particularly pronounced advantages under aggressive compression. Furthermore, evaluations on edge-device show that VPRune effectively reduces end-to-end inference latency while maintaining superior task performance, demonstrating its practicality for resource-constrained LVLM deployment.
Guangchuan Lv, Dianxing Shi, Dingjie Fu
Northeastern University · Beihang University · Huazhong University of Science and Technology
Large language models (LLMs) have demonstrated remarkable capabilities across diverse language tasks, motivating their extension to vision-language models (VLMs) for multimodal understanding. However, billion-parameter VLMs incur substantial memory and computational costs that hinder deployment in resource-constrained settings. Post-training quantization (PTQ) compresses models and accelerates inference without retraining, yet remains underexplored for VLMs. We identify two intrinsic VLM activation properties in PTQ: (1) visual over-representation, where vision tokens are excessive and often redundant, and (2) the modality gap separating text and vision tokens in the latent feature space. Prior methods largely overlook these properties, leading to quantization performance degradation. To address this mismatch, we propose SalQ-VLM, an importance-aware PTQ framework that prioritizes salient tokens and suppresses redundant vision tokens during calibration. We derive a gradient-driven importance factor that captures token-level importance variance and is theoretically grounded in the relationship among loss perturbation, activation errors, and output gradients. SalQ-VLM obtains this factor through a single lightweight block-wise gradient-caching pass and incorporates it into the layer-wise reconstruction objective. Because SalQ-VLM modifies only calibration, it adds no inference-time operations and remains compatible with existing high-performance kernels. Extensive evaluations across benchmarks and backbones show that SalQ-VLM consistently outperforms strong PTQ baselines, especially under ultra-low-bit quantization. Notably, it improves MME-RealWorld accuracy by 16.45% under INT2g128 quantization.
Yufei Xue, Yushi Huang, Lunjie Zhu +2
Institute of Artificial Intelligence (TeleAI), China Telecom · The Hong Kong University of Science and Technology