When Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM Quantization
Authors: Yanlong Zhao, Xiaoyuan Cheng, Huihang Liu, Baihua He, Xinyu Zhang, Harrison Bo Hua Zhu, Wenlong Chen, Li Zeng, +1 more
Organizations: University of Science and Technology of China · University College London · Shanghai University of Finance and Economics · AMSS, Chinese Academy of Sciences · University of Copenhagen · Imperial College London · Technical University of Denmark · Shenzhen Research Institute of Big Data
Weight-only post-training quantization (PTQ) relies heavily on reconstruction loss minimization to preserve model quality at low precision. We show that the weights favored by minimizing this loss need not yield better model performance on new tasks. In fact, we find that lower reconstruction loss can even degrade model performance on the same calibration data. Our analysis further shows that weights with lower reconstruction loss on calibration data can have higher loss than other weights when the distribution of input activations changes. Motivated by these observations and our analysis, we propose Distributionally Robust Quantization (DRQ), a post-hoc refinement process that minimizes worst-case reconstruction loss over a constrained set of input activation distributions. DRQ refines the integer codes representing quantized weights within the existing quantization grid, keeping quantization parameters and inference operators unchanged. Extensive experiments show that DRQ improves models quantized by six representative PTQ methods, including AWQ, GPTQ, and ParoQuant, and delivers gains across both dense and mixture-of-experts large language models. These results establish DRQ as a general post-hoc refinement framework for weight-only PTQ, achieving better downstream performance without adding inference overhead.
Figures & tables
Figure 1: Lower reconstruction loss need not improve quantized model performance. Results are from Llama-3.2-1B-Instruct with 3-bit quantization. (a) Among 32 modifications to AWQ-quantized weights, 18 reduce reconstruction loss but increase model PPL on the same WT2 calibration data. (b) Among 560 modified weight matrices per method that reduce reconstruction loss on calibration data, 413 (73.8%) for AWQ and 264 (47.1%) for GPTQ increase model PPL on held-out WT2 data. (c) Further minimizing reconstruction loss increases held-out C4 PPL by 0.69%, whereas DRQ accepts slightly higher reconstruction loss on calibration data and lowers C4 PPL by 0.32%. Each evaluation modifies one linear layer, keeping all other weights fixed. In (a,c), reconstruction loss is divided by its initial value, and PPL changes are relative to the initial quantized model.
Figure 2: An illustrative example of integer-code refinement with DRQ on Llama-3.2-1B-Instruct quantized with 2-bit GPTQ. We select two integer codes, q1 and q2 , from the quantized weights and plot the loss contours over different pairs of code values. The percentages indicate changes in reconstruction loss relative to the initial quantized weights. Panel (a) shows reconstruction loss on the WikiText-2 (WT2) calibration data, panel (b) illustrates how DRQ refines the two integer codes, and panel (c) shows reconstruction loss on C4. Although DRQ increases reconstruction loss on the calibration data by 1.78%, it reduces reconstruction loss on C4 by 15.11%, both relative to the initial quantized weights.
Quantization
Llama-3.2-1B
Llama-3.2-3B
Llama-3.1-8B
Llama-3.1-70B
Bits
Method
WT2
C4
ACC
WT2
C4
ACC
WT2
C4
ACC
WT2
C4
ACC
FP
Unquantized
13.16
20.82
58.68
11.05
16.19
66.48
7.22
11.24
74.40
3.78
8.24
80.64
W2
GPTQ
1719.88
1.03e4
34.75
141.54
645.35
34.58
139.90
681.48
35.49
13.92
31.77
48.23
+ DRQ
1402.87
7233.14
35.93
130.02
585.86
35.03
98.67
390.36
36.64
12.48
27.11
54.56
AWQ
1.08e5
1.22e5
35.39
2.42e5
5.57e5
35.38
1.63e6
1.84e6
37.80
1.42e6
1.25e6
37.99
+ DRQ
1.31e4
1.80e4
36.51
1.85e4
1.03e5
38.38
7.76e5
5.66e5
38.01
1.31e6
1.17e6
38.01
Table 1: Perplexity and average task accuracy (%) of dense Llama models across model sizes and weight bit widths. FP denotes the unquantized full-precision model.
Quantization
Perplexity
Accuracy (%)
Bits
Method
WT2
C4
PIQA
HellaSwag
MMLU
BoolQ
ARC-C
ARC-E
WinoGrande
Mean
W3
GPTQ
9.29
15.01
78.24
75.85
72.16
86.15
49.57
73.99
68.90
72.12
+ DRQ
9.20
14.90
78.78
76.00
73.07
85.87
52.99
76.05
67.56
72.90
Table 2: Perplexity and seven-task accuracy (%) for Qwen3-30B-A3B W3A16.
Figure 3: Six-method comparison on Llama-3.2-3B W3.
Figure 4: Reconstruction loss and PPL across calibration sizes for Llama-3.2-3B-Instruct with GPTQ W4. Loss ratios divide the reconstruction loss summed across layers after refinement by that before refinement, using each run’s own calibration activations.
Figure 5: Quantization time for Llama-3.1-8B W3.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
W,W(q)
Full-precision and quantized weights in Rm×d , after any scaling or rotation by the base PTQ method
q,Q
Integer codes representing quantized weights and the grid defined by fixed bit width, grouping, scales, and zero points
W0
Initial quantized weights produced by the base PTQ method, used to limit increases in reconstruction loss
X
Calibration input activation matrix in Rd×n , with one token per column
H
Activation second-moment matrix XX⊤/n
Xd,Xc(h)
Input activations in the search and checking subsets, with nd and nc(h) columns. DRQ computes H and L from the search subset
Appendix
Table 3: Summary of notation.
Bits
Method
PIQA
HellaSwag
MMLU
BoolQ
ARC-C
ARC-E
WinoGrande
Mean
FP
Unquantized
74.16
60.76
45.86
69.42
38.05
63.26
59.27
58.68
W2
GPTQ
51.52
26.47
24.19
42.54
24.66
25.63
48.22
34.75
+ DRQ
51.80
26.79
26.41
43.15
25.77
26.22
51.38
35.93
AWQ
51.74
25.99
23.66
47.65
25.68
23.53
49.49
35.39
+ DRQ
51.31
26.93
24.86
51.25
25.51
24.20
51.54
36.51
W3
GPTQ
67.90
52.69
28.32
62.26
32.68
53.16
56.51
50.50
Appendix
Table 4: Llama-3.2-1B-Instruct accuracy (%) on each downstream task.
Bits
Method
PIQA
HellaSwag
MMLU
BoolQ
ARC-C
ARC-E
WinoGrande
Mean
FP
Unquantized
75.68
70.47
60.45
77.74
45.56
67.76
67.72
66.48
W2
GPTQ
51.69
27.74
23.86
40.24
25.09
26.01
47.43
34.58
+ DRQ
50.60
29.11
24.83
40.43
25.00
25.34
49.88
35.03
AWQ
51.31
26.76
26.89
42.66
25.26
26.09
48.70
35.38
+ DRQ
51.36
28.44
26.89
60.67
24.74
26.30
50.28
38.38
W3
GPTQ
69.15
61.76
50.04
63.61
33.53
52.65
60.06
55.83
Appendix
Table 5: Llama-3.2-3B-Instruct accuracy (%) on each downstream task.
Bits
Method
PIQA
HellaSwag
MMLU
BoolQ
ARC-C
ARC-E
WinoGrande
Mean
FP
Unquantized
81.12
79.25
68.05
83.94
55.03
79.59
73.80
74.40
W2
GPTQ
49.95
27.67
24.11
43.03
26.37
27.53
49.80
35.49
+ DRQ
51.80
29.11
24.60
47.43
25.85
26.56
51.14
36.64
AWQ
51.03
26.60
26.89
61.93
24.40
25.00
48.78
37.80
+ DRQ
51.03
26.67
26.89
60.28
24.74
25.21
51.22
38.01
W3
GPTQ
76.93
75.40
57.01
81.53
45.48
71.25
71.59
68.46
Appendix
Table 6: Llama-3.1-8B-Instruct accuracy (%) on each downstream task.
Bits
Method
PIQA
HellaSwag
MMLU
BoolQ
ARC-C
ARC-E
WinoGrande
Mean
FP
Unquantized
83.73
84.63
82.29
88.01
63.23
83.67
78.93
80.64
W2
GPTQ
63.71
60.43
29.01
58.50
28.67
40.61
56.67
48.23
+ DRQ
67.08
64.81
35.14
68.56
34.22
51.64
60.46
54.56
AWQ
51.52
26.66
24.65
62.17
26.71
24.75
49.49
37.99
+ DRQ
51.31
26.61
24.65
62.17
26.54
25.04
49.72
38.01
W3
GPTQ
82.70
82.90
79.86
86.57
59.30
81.78
77.58
78.67
Appendix
Table 7: Llama-3.1-70B-Instruct accuracy (%) on each downstream task.
Perplexity
Accuracy (%)
Method
WT2
C4
PIQA
HellaSwag
MMLU
BoolQ
ARC-C
ARC-E
WinoGrande
Mean
RTN
17.97
26.96
69.42
62.41
45.21
72.39
37.12
56.78
62.51
57.98
+ DRQ
17.60
26.12
70.18
62.03
45.20
73.58
36.95
57.95
62.04
58.28
GPTQ
15.74
30.34
69.15
61.76
50.04
63.61
33.53
52.65
60.06
55.83
+ DRQ
14.63
22.11
71.33
65.25
49.25
73.33
38.57
59.60
64.33
60.24
AWQ
15.19
20.80
72.96
65.02
51.59
73.18
39.85
64.73
64.48
61.69
Appendix
Table 8: Perplexity and seven-task accuracy (%) for Llama-3.2-3B W3.
Perplexity
Accuracy (%)
Bits
Method
WT2
C4
PIQA
HellaSwag
MMLU
BoolQ
ARC-C
ARC-E
WinoGrande
Mean
W2
QuaRot + GPTQ
72.62
164.40
53.05
30.91
24.11
44.59
22.95
29.46
49.57
36.38
+ DRQ
67.53
159.03
53.32
31.06
23.33
46.67
24.23
29.76
49.41
36.82
W3
QuaRot + GPTQ
8.58
14.86
78.18
74.07
58.60
81.44
46.67
69.95
70.17
68.44
+ DRQ
8.24
14.23
78.35
75.41
60.18
82.63
48.81
71.25
71.43
69.72
Appendix
Table 9: QuaRot + GPTQ with DRQ on Llama-3.1-8B-Instruct.
Figure 6: Calibration token counts and integer-code changes in Qwen3-30B-A3B W3A16. (a) The 18,624 quantized linear layers grouped by the number of calibration tokens received, with layer counts shown beside the bars. (b) The fraction of linear layers with at least one changed integer code in each token-count group. The 0–2-token groups contain expert linear layers exclusively, while the ≥3 group includes both expert and attention linear layers.
Figure 7: Performance across different α and β combinations on Llama-3.2-3B-Instruct at W3.
Base PTQ method
Layer refinement (s)
Including integration (s)
RTN
556.74
788.04
GPTQ
227.28
233.24
AWQ
482.56
488.99
OmniQuant
520.38
792.13
GPTAQ
296.92
967.00
ParoQuant (earlier run)
220.79
794.04
Appendix
Table 10: DRQ refinement time for Llama-3.2-3B W3.
Figure 8: DRQ refinement time across model sizes and weight bit widths.
Method
Batch size
Input (tokens)
Prefill (ms)
Decoding (ms/token)
Throughput (tokens/s)
Peak memory (GiB)
GPTQ
1
128
33.61
16.52
60.04
7.86
+ DRQ
33.49
16.49
60.16
7.86
GPTQ
1
2,048
156.58
16.49
56.92
8.01
+ DRQ
157.86
16.59
56.58
8.01
GPTQ
4
128
49.07
17.04
231.35
7.89
+ DRQ
49.05
16.90
233.17
7.89
Appendix
Table 11: Inference cost across batch sizes and input lengths on Llama-3.1-8B-Instruct, W4A16. Output length is fixed at 128 tokens.