cs.CVSep 28, 2026

Beyond Reconstruction Loss in Post-Training Quantization: Balanced Fitting for Large Vision-Language Models

Authors: Minchan Kang, Kyeonghye Park, Seungyeon Sa, Seoyoung Cho, Daeshik Kim, Yucheol Cho

Organizations: Korea Advanced Institute of Science and Technology (KAIST) · Hanbat National University

Abstract

Post-training quantization (PTQ) enables efficient deployment of large vision-language models (LVLMs), but is typically calibrated on a small set while expected to generalize across diverse downstream tasks. Although recent PTQ methods for LVLMs incorporate sensitivity signals, they still minimize reconstruction loss with respect to the full-precision model, potentially over-preserving FP behavior and calibration-specific bias. Rather than treating quantization solely as an error to be minimized, we observe that it can also provide beneficial regularization for certain layers and modalities. Motivated by this observation, we propose Balanced Fitting, a quantization effect-based framework that balances precision and regularization beyond reconstruction-based optimization. By measuring layer- and component-wise quantization effects for weights, vision activations, and text activations, Balanced Fitting combines fine-grained fitting for sensitive components with coarser fitting to exploit potential regularization benefits. Experiments on multiple LVLMs show that our method consistently outperforms prior PTQ approaches under both weight-only and weight-activation quantization, while lower reconstruction loss does not reliably translate into better downstream performance. The source code is publicly available at https://github.com/kmc3661/BFQ

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SalQ-VLM: Fine-Grained Saliency-Guided Quantization for Vision-Language Models

    Aug 5, 2025Yufei Xue, Yushi Huang, Lunjie Zhu +2Recent Vision-Language ModelsQuantization-Aware Training

  2. Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models

    May 19, 2026Yi Zhong, Haotong Qin, Xindong Zhang +2Cross-ModalModalities

  3. Saliency-Aware Regularized Quantization Calibration for Large Language Models

    May 7, 2026Yanlong Zhao, Xiaoyuan Cheng, Huihang Liu +6Large Language Model QuantizationPost-Training Quantization