Post-training quantization (PTQ) enables efficient deployment of large vision-language models (LVLMs), but is typically calibrated on a small set while expected to generalize across diverse downstream tasks. Although recent PTQ methods for LVLMs incorporate sensitivity signals, they still minimize reconstruction loss with respect to the full-precision model, potentially over-preserving FP behavior and calibration-specific bias. Rather than treating quantization solely as an error to be minimized, we observe that it can also provide beneficial regularization for certain layers and modalities. Motivated by this observation, we propose Balanced Fitting, a quantization effect-based framework that balances precision and regularization beyond reconstruction-based optimization. By measuring layer- and component-wise quantization effects for weights, vision activations, and text activations, Balanced Fitting combines fine-grained fitting for sensitive components with coarser fitting to exploit potential regularization benefits. Experiments on multiple LVLMs show that our method consistently outperforms prior PTQ approaches under both weight-only and weight-activation quantization, while lower reconstruction loss does not reliably translate into better downstream performance. The source code is publicly available at https://github.com/kmc3661/BFQ
Figures & tables
Figure 1: Accuracy drop from the full-precision (FP16) model across multiple LVLMs under low-bit settings (W3A16 and W4A8). Results are reported on the MMMU benchmark (left) and the average across five multimodal benchmarks (right). Our method consistently achieves the lowest accuracy drop across all models and configurations among state-of-the-art PTQ methods for LVLMs.
Figure 2: Motivation of our approach. (a) Component-wise quantization effects across layers reveal heterogeneous sensitivity across weights and modalities, where some components benefit from quantization while others are sensitive. (b) Conventional PTQ uniformly minimizes reconstruction error, leading to suboptimal solutions, whereas our Balanced Fitting adapts fitting strength using component-wise effects, improving generalization.
Figure 3: Overview of the proposed balanced fitting framework. We first estimate layer- and component-wise quantization effects by selectively quantizing weights and activations. These effects are then used to guide a layer-wise budget allocation policy that adaptively controls the search granularity. The resulting allocation balances precise fitting for quantization-sensitive components and coarser fitting for others, improving generalization.
Model
Bitwidth
Method
MMMU
VizWiz
ScienceQA
ChartQA
AI2D
Avg.
Δ FP
LLaVA-OV-7B
FP16
-
46.56
58.71
95.84
80.00
81.28
72.48
—
W3A16
RTN
42.22
57.00
94.55
68.88
78.98
68.33
−4.15
Q-VLM
44.22
59.72
94.40
76.76
78.63
70.75
−1.73
MBQ
43.11
59.88
94.74
77.00
78.27
70.60
−1.88
QIG
43.89
59.54
94.65
77.12
78.47
70.73
−1.75
Ours
46.11
60.28
94.94
77.20
79.02
71.51
−0.97
Table 1: Quantitative results across three LVLMs and five benchmarks. Δ FP: Avg. − FP Avg. (pp).
Figure 4: Qualitative comparison on MMMU and VizWiz.
Figure 5: Layer-wise budget allocation and downstream performance on InternVL2-8B.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Component-wise quantization effects across layers for three vision-language models (InternVL2-8B, LLaVA-OV-7B, and Qwen2-VL-7B) under two quantization settings (W3A16 and W4A8). The results reveal heterogeneous sensitivity across weights and modalities, where some components benefit from quantization while others are more sensitive.
Figure 7: Layer-wise budget allocation and downstream performance on LLaVA-OV-7B and Qwen2-VL-7B. For the W3A16 setting, the minimum grid size is set to 20, making the high-only configuration equivalent to the full setting and the low-only configuration equivalent to the baseline; thus, they are excluded from comparison.
Figure 8: Stability of layer-wise quantization effects estimated from 16, 32, and 64 calibration samples on Qwen2-VL-7B, LLaVA-OV-7B, and InternVL2-8B under W3A16 (top) and W4A8 (bottom). W4A8 aggregates weight, visual-token, and textual-token effects. The reported Pearson correlations compare the 16- and 32-sample profiles with the 64-sample reference.
Figure 9: Analysis of the proposed Quantization Effect-Guided Allocation Rule. Each row corresponds to a model and setting, and each column sweeps one allocation dimension (base grid, target mean, or γ ) while keeping the others fixed. Numbers in parentheses indicate the 5-task average performance, and curves report MMMU performance. The dashed vertical line indicates the reference allocation in the sweep.
Figure 10: Additional qualitative comparison on MMMU and VizWiz.
Figure 11: Additional qualitative comparison on ScienceQA, ChartQA, and AI2D.
Large language models (LLMs) have demonstrated remarkable capabilities across diverse language tasks, motivating their extension to vision-language models (VLMs) for multimodal understanding. However, billion-parameter VLMs incur substantial memory and computational costs that hinder deployment in resource-constrained settings. Post-training quantization (PTQ) compresses models and accelerates inference without retraining, yet remains underexplored for VLMs. We identify two intrinsic VLM activation properties in PTQ: (1) visual over-representation, where vision tokens are excessive and often redundant, and (2) the modality gap separating text and vision tokens in the latent feature space. Prior methods largely overlook these properties, leading to quantization performance degradation. To address this mismatch, we propose SalQ-VLM, an importance-aware PTQ framework that prioritizes salient tokens and suppresses redundant vision tokens during calibration. We derive a gradient-driven importance factor that captures token-level importance variance and is theoretically grounded in the relationship among loss perturbation, activation errors, and output gradients. SalQ-VLM obtains this factor through a single lightweight block-wise gradient-caching pass and incorporates it into the layer-wise reconstruction objective. Because SalQ-VLM modifies only calibration, it adds no inference-time operations and remains compatible with existing high-performance kernels. Extensive evaluations across benchmarks and backbones show that SalQ-VLM consistently outperforms strong PTQ baselines, especially under ultra-low-bit quantization. Notably, it improves MME-RealWorld accuracy by 16.45% under INT2g128 quantization.
Yufei Xue, Yushi Huang, Lunjie Zhu +2
Institute of Artificial Intelligence (TeleAI), China Telecom · The Hong Kong University of Science and Technology
Low-bit post-training quantization (PTQ) is a pivotal technique for deploying Vision-Language Models (VLMs) on resource-constrained devices. However, existing PTQ methods often degrade VLMs' accuracy due to the heterogeneous activation distributions of text and vision modalities during quantization. We find that this cross-modal heterogeneity is distributed unevenly across channels: a small subset of channels contains most modality-specific outliers, and these outliers typically reside in different channels for each modality. Motivated by this, we propose SplitQ, a channel-Splitting-driven post-training Quantization framework. At its core, SplitQ introduces a novel Modality-specific Outlier Channel Decoupling (MOCD) module that effectively isolates salient modality-specific outlier channels with minimal overhead. To further address the remaining cross-modal distribution discrepancies, we design an Adaptive Cross-Modal Calibration (ACC) module that employs dual lightweight learnable branches to dynamically mitigate modality-induced quantization errors. Extensive experiments on popular VLMs demonstrate that SplitQ significantly outperforms existing approaches across 6 popular multi-modal datasets under all evaluated quantization settings, including W4A8, W4A4, W3A3, and W3A2. Notably, SplitQ preserves 93.5% of FP16 performance under the challenging W3A3 setting (69.5 vs. 74.3), pushing the efficiency frontier for deploying advanced VLMs. Our code is available at https://github.com/EMVision-NK/SplitQ
Yi Zhong, Haotong Qin, Xindong Zhang +2
VCIP, College of Computer Science, Nankai University · D-ITET, ETH Zürich · OPPO Research Institute +1
Post-training quantization (PTQ) is an effective approach for deploying large language models (LLMs) under memory and latency constraints. Most existing PTQ methods determine quantization parameters by minimizing a layer-wise reconstruction error on a predetermined calibration dataset, typically optimized via either scale search or Gram-based methods. However, from the perspective of generalization risk, existing PTQ calibration objectives based solely on empirical reconstruction error over limited or unrepresentative calibration data may move the quantized weights away from the original floating-point weights, potentially degrading downstream performance. To address this issue, we propose \emph{Regularized Quantization Calibration} (RQC), a unified framework that augments standard PTQ objectives with a regularizer that explicitly controls weight deviation from the original weights. We further generalize this framework to incorporate a saliency-aware regularizer, resulting in \emph{Saliency-Aware Regularized Quantization Calibration} (SARQC). The proposed regularization encourages quantized weights to remain close to the original weights during calibration, leading to improved generalization at inference time. SARQC integrates seamlessly into existing PTQ pipelines and enhances both scale-search-based and Gram-based methods under a unified formulation. Extensive experiments on dense and Mixture-of-Experts LLMs demonstrate consistent improvements in perplexity and zero-shot accuracy, without introducing additional inference overhead.
Yanlong Zhao, Xiaoyuan Cheng, Huihang Liu +6
University of Science and Technology of China · University College London · 3Shanghai University of Finance and Economics +5