cs.CVSep 28, 2026

P4Q: Co-designing Token Pruning and Quantization for Vision-Language Model Acceleration

Authors: Haizhao Jing, Zhenhao Shang, Haokui Zhang, Rong Xiao, Peng Wang

Organizations: Northwest Polytechnical University · Intellifusion

Abstract

Vision language models have achieved strong performance across a wide range of multimodal applications, yet their substantial computational and memory costs hinder efficient deployment. Visual token pruning and post-training quantization reduce inference overhead along two complementary dimensions, namely sequence length and numerical precision. Existing workflows typically optimize these techniques independently or apply them sequentially. Their distinct optimization objectives leave critical interactions unaddressed and constrain the achievable compression performance. We revisit these designs and present P4Q, a practical co-design framework that jointly optimizes visual token pruning and low-bit quantization for efficient VLM inference. First, P4Q introduces a quantization-aware visual token selection strategy before the LLM. It applies fake quantization to copies of the features produced by the projector and selects visual tokens using statistics computed from these fake-quantized features, thereby conditioning the selector's feature-based decisions on simulated low-bit perturbations. Second, P4Q introduces a pruning-aware quantization calibration strategy. It uses the same selection strategy as pruning to calibrate the quantized model on the retained-token distribution, thereby aligning the calibration process with the pruned execution path used during deployment. By coupling these two components, P4Q achieves substantial inference speedups while maintaining comparable task performance, resulting in a better efficiency-accuracy trade-off than independently optimized pipelines. For instance, on LLaVA-NeXT, P4Q achieves an average end-to-end inference speedup of 2.8x across eight distinct test sets, while retaining higher accuracy than prior compression and quantization methods.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Towards Joint Quantization and Token Pruning of Vision-Language Models

    Apr 19, 2026Xinqing Li, Xin He, Xindong Zhang +3Visual Token PruningLarge Language Model Quantization

  2. VPRune: Efficient Training-free Pre-LLM Visual Token Pruning

    Sep 21, 2026Guangchuan Lv, Dianxing Shi, Dingjie FuVisual Token PruningRecent Vision-Language Models

  3. SalQ-VLM: Fine-Grained Saliency-Guided Quantization for Vision-Language Models

    Aug 5, 2025Yufei Xue, Yushi Huang, Lunjie Zhu +2Recent Vision-Language ModelsQuantization-Aware Training