cs.LGMay 31, 2026

GPTQ-intrinsic LoRA: A Near-optimal Algorithm for Low-precision Quantization with Low-rank Adaptation

Authors: Shihao ZhangRayan Saab

Abstract

Post-training quantization is widely used for compressing large neural networks, but aggressive low-bit quantization can significantly degrade model quality. A common remedy is to augment the quantized weights with a low-rank correction, leading to approximations of the form WQ+LRW\approx Q+LR. In this paper, we study this low-precision plus low-rank representation through the layer-wise reconstruction objective XWX(Q+LR)F2\|XW-X(Q+LR)\|_F^2, where XX is a calibration matrix. We establish, to our knowledge, the first information-theoretic lower bounds for this problem under finite-alphabet and bounded low-rank compensation constraints. We then propose GPTQ-intrinsic LoRA, a training-free algorithm that incorporates the low-rank correction directly into a GPTQ-style quantization pass by appropriately augmenting the calibration Hessian. For the choice L=VrL=V_r, where VrV_r contains the top right singular vectors of XX, we prove layer-wise reconstruction error bounds in which the usual GPTQ dependence on XF2\|X\|_F^2 is replaced by the rank-rr residual XXrF2\|X-X_r\|_F^2, up to regularization terms. Under natural structural assumptions, these bounds match the information-theoretic lower bounds in their dominant scaling, up to constants and mild factors. We also introduce Bid-Up, a fixed-grid quantization refinement step that can be alternated with optimal low-rank compensation with guaranteed non-increasing layer-wise reconstruction error. Experiments on Qwen3 language models and DeiT vision transformers show that GPTQ-intrinsic LoRA improves over GPTQ and GPTQ followed by low-rank compensation, with additional gains from refinement loops.

Explore similar work

CardsList
  1. LoRaQ: Optimized Low Rank Approximation for 4-bit Quantization

    Apr 20, 2026Yann Bouquet, Alireza Khodamoradi, Sophie Yáng Shen +2QuantizationLow-Rank Factorization