cs.AIOct 8, 2026

When Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM Quantization

Authors: Yanlong Zhao, Xiaoyuan Cheng, Huihang Liu, Baihua He, Xinyu Zhang, Harrison Bo Hua Zhu, Wenlong Chen, Li Zeng, +1 more

Organizations: University of Science and Technology of China · University College London · Shanghai University of Finance and Economics · AMSS, Chinese Academy of Sciences · University of Copenhagen · Imperial College London · Technical University of Denmark · Shenzhen Research Institute of Big Data

Abstract

Weight-only post-training quantization (PTQ) relies heavily on reconstruction loss minimization to preserve model quality at low precision. We show that the weights favored by minimizing this loss need not yield better model performance on new tasks. In fact, we find that lower reconstruction loss can even degrade model performance on the same calibration data. Our analysis further shows that weights with lower reconstruction loss on calibration data can have higher loss than other weights when the distribution of input activations changes. Motivated by these observations and our analysis, we propose Distributionally Robust Quantization (DRQ), a post-hoc refinement process that minimizes worst-case reconstruction loss over a constrained set of input activation distributions. DRQ refines the integer codes representing quantized weights within the existing quantization grid, keeping quantization parameters and inference operators unchanged. Extensive experiments show that DRQ improves models quantized by six representative PTQ methods, including AWQ, GPTQ, and ParoQuant, and delivers gains across both dense and mixture-of-experts large language models. These results establish DRQ as a general post-hoc refinement framework for weight-only PTQ, achieving better downstream performance without adding inference overhead.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization

    Aug 7, 2026Yongge Ma, Guoan Wang, Feiyu Wang +5LLM QuantizationPost-Training Quantization

  2. Beyond Reconstruction Loss in Post-Training Quantization: Balanced Fitting for Large Vision-Language Models

    Sep 28, 2026Minchan Kang, Kyeonghye Park, Seungyeon Sa +3VLM QuantizationLarge Vision-Language Models

  3. QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction

    Aug 14, 2026Vincent Counathe, Ben Athiwaratkun, Christopher De Sa +1LLM QuantizationQuantization-Aware Training