cs.CVSep 28, 2026

SubRot: Signed Gradient Subspace Calibration for VLM Rotation Quantization

Authors: Zhenhao Shang, Haizhao Jing, Haokui Zhang, Guoting Wei, Rong Xiao, Jianqing Gao, Peng Wang

Organizations: Northwest Polytechnical University · Nanjing University of Science and Technology · Intellifusion · iFLYTEK CO., LTD

Abstract

Post-training quantization reduces the deployment cost of vision-language models (VLMs), but preserving multimodal capabilities at low bit widths remains challenging. Existing methods rely on modality- or token-level gradient statistics, which are susceptible to cross-sample variations in visual-to-textual token ratios and the positions of visual information, limiting statistical stability. Moreover, overly coarse aggregation through absolute values and averaging discards gradient signs and channel-wise differences, limiting the separation of modality-specific sensitivities. In contrast, the channel space provides a shared coordinate system across samples, making it a more natural basis for capturing stable task-sensitive structures. We therefore propose SubRot, a signed gradient subspace calibration method for VLM rotation quantization. Through eigendecomposition of the empirical Fisher matrix of activation gradients, SubRot identifies a sensitive channel subspace with three properties: cross-sample stability, clear sensitivity separation, and consistent signed effects on the autoregressive loss along certain directions. Guided by a local Taylor expansion, SubRot combines signed first-order guidance along sign-stable directions with second-order constraints along the remaining sensitive directions, while retaining MSE for overall reconstruction quality. This objective steers quantization errors toward loss-decreasing directions while controlling their magnitude. Experiments on five VLMs across five benchmarks show consistent average-score improvements over FlatQuant under W4A6 and W4A4, reaching 1.4 percentage points on LLaVA-NeXT-7B. Under W4A4, average accuracy degradation from FP16 remains within 1.4 percentage points across all evaluated models, while LLaVA-v1.5-13B exceeds its FP16 average score by 0.4 percentage points.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RGSQ: Riemannian Geometry-Sensitive Quantization for Large Vision-Language Models

    Sep 21, 2026Zhiping Wu, Dongdong Ren, Yangchengyu Zhou +4Large Language Model QuantizationPost-Training Quantization

  2. ThinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLMs

    Sep 28, 2026Mehdi Makni, Ryan Lucas, Rahul MazumderLarge Language Model QuantizationQantis

  3. SalQ-VLM: Fine-Grained Saliency-Guided Quantization for Vision-Language Models

    Aug 5, 2025Yufei Xue, Yushi Huang, Lunjie Zhu +2Recent Vision-Language ModelsQuantization-Aware Training