Post-training quantization is a powerful tool for compressing large language models. The most scalable methods quantize every layer in parallel, but quantization errors then compound through the residual stream, as no layer corrects for the errors of the layers before it. Sequential quantization accounts for this error compounding by re-calibrating each layer on the already-quantized outputs of its predecessors, yielding stronger results but at the cost of a serial schedule that becomes a bottleneck at scale. As a solution, we propose parallel quantization with activation denoising, which recovers much of the sequential benefit while keeping quantization fully parallel. Rather than re-calibrating layer-by-layer, we take a robustness perspective and model the upstream error as noise, regularizing to be robust to it through a preprocessing step followed by metric-weighted rounding. Applied at every layer, this regularization forms a depth-compounding smoothness penalty that dampens how strongly quantization errors amplify through the model. Unlike orthogonal rotations commonly used in quantization, which must preserve the model's function, we multiply the weights by a more general linear transformation. We find that the two are complementary and their effects compound. Empirically, our robustness regularization recovers a significant part of sequential quantization's benefit in a single parallel pass, at a fraction of its time. Overall, by treating compounding quantization errors as a robustness problem, we offer a principled foundation for more efficient and accurate LLM quantization at scale.
Figures & tables
Figure 1: A robustness perspective on the parallel-sequential quantization gap. Quantization errors enter the residual stream at every layer and accumulate with depth. Parallel quantization (left) is efficient but ignores this error; sequential quantization (right) corrects it, but only one layer after another. Instead of measuring and correcting this error layer by layer, we regularize every layer against noise (middle): we inject random noise in a single extra forward pass and use its moments to construct a denoising filter Fℓ that is absorbed into each layer’s weights before rounding. Every layer thereby becomes more robust to noise, at no additional inference cost. Overall, our findings suggest that accumulated quantization error can be treated as a robustness problem: regularizing all layers against noise in parallel offers an efficient alternative to correcting the error sequentially.
Quantization objective
Llama 1B
Llama 3B
Llama 8B
Llama 13B
Acc ↑
Avg ↓
Gap ↑
Acc ↑
Avg ↓
Gap ↑
Acc ↑
Avg ↓
Gap ↑
Acc ↑
Avg ↓
Gap ↑
FP16
54.8
11.33
–
62.6
9.17
–
67.1
7.57
–
66.5
5.73
–
3-bit GPTQ
Par.
symmetric correction
49.4
17.07
0%
58.5
12.11
0%
63.9
9.94
0%
64.2
6.40
0%
Denoising reg. ( ours )
49.3
16.51
31%
57.1
11.97
21%
64.9
9.81
30%
64.5
6.37
18%
Seq.
symmetric correction
49.8
16.53
31%
59.0
11.99
17%
63.5
9.81
29%
63.7
6.38
13%
asymmetric correction
50.3
15.29
100%
58.4
11.41
100%
63.8
9.51
100%
64.5
6.21
100%
Table 1: Parallel denoising closes a large part of the parallel–sequential gap for both scalar and vector quantization. 3 -bit GPTQ and 2 -bit QuIP# on Llama-3.2-1B/3B, Llama-3-8B, and Llama-2-13B. Acc: zero-shot accuracy; Avg: average perplexity; Gap: fraction of the perplexity gap closed, from 0% (parallel) to 100% (sequential asym. correction). On Llama-2-13B at 2 bits, the two methods nearly coincide, so we only indicate whether a method is below the sequential ( >100% ) or above the parallel ( <0% ). Per-dataset and learned-rotation results in Appendix A .
Figure 2: Projected multi-GPU quantization time (A100, QuIP#).
Figure 3: Random rotation and noise ablation (W3 Llama-3.2-1B). Avg. perplexity over 100 draws: parallel GPTQ and sequential GPTAQ resample the Hadamard rotation (stars: learned-rotation), while for our denoising we resample noise either for random or learned rotations (gray backgrounds). Notably, our parallel denoising significantly reduces the parallel–sequential gap and is insensitive to the random noise draw.
Figure 4: Denoising reduces weight norms, and two alternatives to full sequential quantization. Left (Frobenius-norm reduction on Llama-3.2-1B): ratio ∥WF∥F/∥W∥F of denoised to original weights before rounding, across depth and module type. The ratio is below 1 for all 112 linears, most strongly in the attention output projection. Middle (quantization-time trade-off): quantizing the first K blocks sequentially and the remaining L−K blocks in parallel, we find the sequential benefit is front-loaded: the first quarter of blocks closes 40 – 60% of the gap, but the last few percent need a prefix spanning most of the depth. Right (inference-time trade-off on Llama-3.2-3B): avg. perplexity vs. avg. effective bit-width, promoting a prefix of all/MLP-only/ vproj -only modules to 4 bits (green diamonds: finer group-size; red dashed: seq. GPTAQ). We find promoting only vproj to 4 bits reaches sequential quality while raising the avg. bit-width by only ≈0.03 bits.
Appendix figures & tables35 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Noise-model mechanism ablation (Llama-3.2-1B W3A16, Gaussian denoising, random Hadamard rotation, α=0.01 ). Average perplexity over the 5×5(ρ,δ) grid. Each title reports the best value and the % of the parallel–sequential gap it closes. Left : no injected noise, so the objective collapses to parallel GPTQ and ρ,δ have no effect. Middle : local injection only (noise enters the moments but does not propagate), closing 11% of the gap. Right : the full method with propagated noise, closing 28% . Local injection alone already improves over parallel GPTQ, while propagation accounts for the majority of the improvement.
Figure 6: Full ρ×λ sweep , Llama-3.2-1B W3A16 ( 9×9 grid over [0,1]2 ). Left : injected Gaussian noise (best avg. perplexity 16.31 at λ=0.4,ρ=0.3 ). Right : the real sequential upstream error (best 15.12 at λ=0.8,ρ=0.4 ). In both, λ is the input noise scale and ρ the target scale; the low-perplexity region is below the diagonal ( ρ<λ ).
Quantization objective
Llama 1B
Llama 3B
Llama 8B
Llama 13B
Acc ↑
Avg ↓
Gap ↑
Acc ↑
Avg ↓
Gap ↑
Acc ↑
Avg ↓
Gap ↑
Acc ↑
Avg ↓
Gap ↑
FP16
54.8
11.33
–
62.6
9.17
–
67.1
7.57
–
66.5
5.73
–
Random rot.
Parallel
symmetric correction
49.4
17.07
0%
58.5
12.11
0%
63.9
9.94
0%
64.2
6.40
0%
Denoising reg. ( ours )
49.3
16.51
31%
57.1
11.97
21%
64.9
9.81
30%
64.5
6.37
18%
+ GPTQ-W3 proxy
50.2
16.30
43%
59.1
11.88
34%
64.6
9.78
36%
64.2
6.33
36%
Seq.
symmetric correction
49.8
16.53
31%
59.0
11.99
17%
63.5
9.81
29%
63.7
6.38
13%
Appendix
Table 2: W3A16 GPTQ summary across Llama-3.2-1B/3B, Llama-3-8B, and Llama-2-13B. Same quantization objectives as Table 1 , here reported for the fixed random-Hadamard and the learned SpinQuant rotation side by side, including the GPTQ-W3 proxy noise.
Method
Bits
Grid
Setting
Acc ↑
Wiki ↓
C4 ↓
FW ↓
Avg ↓
Gap ↑
FP16
16
–
–
54.8
9.76
11.54
12.70
11.33
–
Random rot.
Parallel
GPTQ
3.00
25
damp=0.002
49.4
13.21
18.80
19.20
17.07
0%
Gauss. (ours)
3.00
25
ρ=0.55,δ=0.05
49.3
12.91
18.06
18.57
16.51
31%
W3 (ours)
3.00
25
ρ=0.65,δ=0.05
50.1
13.01
18.04
18.53
16.53
30%
W3+Gauss (ours)
3.00
25
ρ=0.25,δ=0.05
50.2
12.76
17.79
18.36
16.30
43%
Seq.
GPTQ
3.00
25
damp=0.05
49.8
12.99
18.03
18.56
16.53
31%
Appendix
Table 3: Llama-3.2-1B at W3A16 (GPTQ). The upper block uses the fixed random-Hadamard rotation, the lower block the SpinQuant-learned rotation; rows marked (ours) are our denoising objective. Gauss. denotes injected Gaussian noise, W3 corresponds to the additional GPTQ-W3 proxy noise, and W3+Gauss combines both.
Figure 7: Llama-3.2-1B W3A16 hyperparameter sweeps, fixed random-Hadamard rotation. Top : the one-dimensional baseline sweeps, avg. perplexity against the swept knob, with the selected optimum starred. Bottom : our denoising objective over its 5×5 grid of the two swept knobs ρ and δ (which sets the noise scale through λ=ρ+δ ); one heatmap per upstream-error proxy (Gaussian noise, the GPTQ-W3 proxy, and their combination), with the best (minimum-perplexity) cell boxed.
Figure 8: Llama-3.2-1B W3A16 hyperparameter sweeps, learned SpinQuant rotation. Panels as in Fig. 7 .
Method
Bits
Grid
Setting
Acc ↑
Wiki ↓
C4 ↓
FW ↓
Avg ↓
Gap ↑
FP16
16
–
–
62.6
7.82
9.29
10.41
9.17
–
Random rot.
Parallel
GPTQ
3.00
25
damp=0.001
58.5
9.61
13.14
13.60
12.11
0%
Gauss. (ours)
3.00
25
ρ=0.35,δ=0.0375
57.1
9.47
12.96
13.47
11.97
21%
W3 (ours)
3.00
25
ρ=0.35,δ=0.05
58.7
9.46
13.02
13.45
11.98
19%
W3+Gauss (ours)
3.00
25
ρ=0.35,δ=0.05
59.1
9.41
12.88
13.34
11.88
34%
Seq.
GPTQ
3.00
25
damp=0.002
59.0
9.51
13.03
13.45
11.99
17%
Appendix
Table 4: Llama-3.2-3B at W3A16 (GPTQ). Columns and rotation blocks as in Table 3 . Gauss. denotes injected Gaussian noise, W3 corresponds to the additional GPTQ-W3 proxy noise, and W3+Gauss combines both.
Figure 9: Llama-3.2-3B W3A16 hyperparameter sweeps, fixed random-Hadamard rotation. Top : the one-dimensional baseline sweeps, avg. perplexity against the swept knob, with the selected optimum starred. Bottom : our denoising objective over its 5×5 grid of the two swept knobs ρ and δ (which sets the noise scale through λ=ρ+δ ); one heatmap per upstream-error proxy (Gaussian noise, the GPTQ-W3 proxy, and their combination), with the best (minimum-perplexity) cell boxed.
Figure 10: Llama-3.2-3B W3A16 hyperparameter sweeps, learned SpinQuant rotation. Panels as in Fig. 9 .
Method
Bits
Grid
Setting
Acc ↑
Wiki ↓
C4 ↓
FW ↓
Avg ↓
Gap ↑
FP16
16
–
–
67.1
6.14
7.78
8.78
7.57
–
Random rot.
Parallel
GPTQ
3.00
25
damp=0.005
63.9
7.60
10.90
11.30
9.94
0%
Gauss. (ours)
3.00
25
ρ=0.55,δ=0.0375
64.9
7.50
10.75
11.17
9.81
30%
W3 (ours)
3.00
25
ρ=0.65,δ=0.05
64.1
7.51
10.83
11.22
9.85
19%
W3+Gauss (ours)
3.00
25
ρ=0.35,δ=0.025
64.6
7.51
10.70
11.14
9.78
36%
Seq.
GPTQ
3.00
25
damp=0.02
63.5
7.56
10.72
11.16
9.81
29%
Appendix
Table 5: Llama-3-8B at W3A16 (GPTQ). Columns and rotation blocks as in Table 3 . Gauss. denotes injected Gaussian noise, W3 corresponds to the additional GPTQ-W3 proxy noise, and W3+Gauss combines both.
Figure 11: Llama-3-8B W3A16 hyperparameter sweeps, fixed random-Hadamard rotation. Top : the one-dimensional baseline sweeps, avg. perplexity against the swept knob, with the selected optimum starred. Bottom : our denoising objective over its 5×5 grid of the two swept knobs ρ and δ (which sets the noise scale through λ=ρ+δ ); one heatmap per upstream-error proxy (Gaussian noise, the GPTQ-W3 proxy, and their combination), with the best (minimum-perplexity) cell boxed.
Figure 12: Llama-3-8B W3A16 hyperparameter sweeps, learned SpinQuant rotation. Panels as in Fig. 11 .
Method
Bits
Grid
Setting
Acc ↑
Wiki ↓
C4 ↓
FW ↓
Avg ↓
Gap ↑
FP16
16
–
–
66.5
4.88
5.80
6.49
5.73
–
Random rot.
Parallel
GPTQ
3.00
25
damp=0.005
64.2
5.37
6.66
7.17
6.40
0%
Gauss. (ours)
3.00
25
ρ=0.35,δ=0.0375
64.5
5.31
6.64
7.15
6.37
18%
W3 (ours)
3.00
25
ρ=0.55,δ=0.05
63.7
5.33
6.64
7.15
6.37
15%
W3+Gauss (ours)
3.00
25
ρ=0.35,δ=0.0375
64.2
5.29
6.60
7.11
6.33
36%
Seq.
GPTQ
3.00
25
damp=0.002
63.7
5.35
6.64
7.15
6.38
13%
Appendix
Table 6: Llama-2-13B at W3A16 (GPTQ). Columns and rotation blocks as in Table 3 . Gauss. denotes injected Gaussian noise, W3 corresponds to the additional GPTQ-W3 proxy noise, and W3+Gauss combines both.
Figure 13: Llama-2-13B W3A16 hyperparameter sweeps, fixed random-Hadamard rotation. Top : the one-dimensional baseline sweeps, avg. perplexity against the swept knob, with the selected optimum starred. Bottom : our denoising objective over its 5×5 grid of the two swept knobs ρ and δ (which sets the noise scale through λ=ρ+δ ); one heatmap per upstream-error proxy (Gaussian noise, the GPTQ-W3 proxy, and their combination), with the best (minimum-perplexity) cell boxed.
Figure 14: Llama-2-13B W3A16 hyperparameter sweeps, learned SpinQuant rotation. Panels as in Fig. 13 .
Method
Bits
Grid
Setting
Acc ↑
Wiki ↓
C4 ↓
FW ↓
Avg ↓
Gap ↑
FP16
16
–
–
54.8
9.76
11.54
12.70
11.33
–
Random rot.
Parallel
GPTQ
4.00
25
damp=0.02
53.1
10.48
13.30
14.11
12.63
0%
Gauss. (ours)
4.00
25
ρ=0.55,δ=0.05
53.7
10.41
13.29
14.15
12.61
7%
W4 (ours)
4.00
25
ρ=0.65,δ=0.025
52.8
10.45
13.28
14.10
12.61
8%
W4+Gauss (ours)
4.00
25
ρ=0.45,δ=0.05
54.0
10.41
13.28
14.13
12.61
9%
Seq.
GPTQ
4.00
25
damp=0.001
53.4
10.47
13.27
14.09
12.61
8%
Appendix
Table 7: Llama-3.2-1B at W4A16 (GPTQ). Columns and rotation blocks as in Table 3 . Gauss. denotes injected Gaussian noise, W4 corresponds to the additional GPTQ-W4 proxy noise, and W4+Gauss combines both. At 4 bits the parallel–sequential gap is small ( ≈0.26 perplexity).
Figure 15: Llama-3.2-1B W4A16 hyperparameter sweeps, fixed random-Hadamard rotation. Top : the one-dimensional baseline sweeps, avg. perplexity against the swept knob, with the selected optimum starred. Bottom : our denoising objective over its 5×5 grid of the two swept knobs ρ and δ (which sets the noise scale through λ=ρ+δ ); one heatmap per upstream-error proxy (Gaussian noise, the GPTQ-W4 proxy, and their combination), with the best (minimum-perplexity) cell boxed.
Figure 16: Llama-3.2-1B W4A16 hyperparameter sweeps, learned SpinQuant rotation. Panels as in Fig. 15 .
Method
Grid
Setting
Acc ↑
Wiki ↓
C4 ↓
FW ↓
Avg ↓
Gap ↑
FP16
–
–
54.8
9.76
11.54
12.70
11.33
–
Parallel
symmetric correction
25
damp=0.55
44.3
19.13
30.36
31.66
27.05
0%
Denoising reg. ( ours )
25
ρ=0.25,δ=0.0125
45.1
18.01
28.46
30.07
25.51
38%
Seq.
symmetric correction
25
damp=0.3
43.8
19.05
29.14
30.88
26.36
17%
asymmetric correction
25
β=0.3333
45.1
16.30
25.67
27.15
23.04
100%
Appendix
Table 8: QuIP# at W2A16 (2-bit), Llama-3.2-1B. Parallel: symmetric correction quantizes each layer from clean activations, Denoising reg. is our objective (Gaussian noise injection, ρ / δ grid). Sequential: symmetric correction uses the symmetric objective on the sequential input, i.e. the outputs of the already-quantized upstream layers; asymmetric correction additionally targets the clean full-precision output (GPTAQ-correction applied to QuIP#). Both then apply E8 rounding. Acc: zero-shot average accuracy (%). Gap: 0%= parallel symmetric, 100%= sequential asymmetric.
Figure 17: QuIP# W2A16 hyperparameter sweeps, Llama-3.2-1B. The three baseline sweeps (parallel symmetric, sequential symmetric, sequential asymmetric), avg. perplexity against the swept knob with the selected optimum starred, and our denoising objective over its 5×5 grid of ρ and δ (which sets the noise scale through λ=ρ+δ ), best cell boxed.
Method
Grid
Setting
Acc ↑
Wiki ↓
C4 ↓
FW ↓
Avg ↓
Gap ↑
FP16
–
–
62.6
7.82
9.29
10.41
9.17
–
Parallel
symmetric correction
25
damp=0.25
52.7
12.25
18.47
19.15
16.62
0%
Denoising reg. ( ours )
25
ρ=0.35,δ=0.025
53.9
11.61
17.21
17.90
15.57
86%
Seq.
symmetric correction
25
damp=0.25
52.6
12.28
18.44
19.19
16.64
−1%
asymmetric correction
25
β=0.3333
51.3
11.23
17.12
17.84
15.40
100%
Appendix
Table 9: QuIP# at W2A16 (2-bit), Llama-3.2-3B. Columns and blocks as in Table 8 .
Figure 18: QuIP# W2A16 hyperparameter sweeps, Llama-3.2-3B. Panels as in Figure 17 .
Method
Grid
Setting
Acc ↑
Wiki ↓
C4 ↓
FW ↓
Avg ↓
Gap ↑
FP16
–
–
67.1
6.14
7.78
8.78
7.57
–
Parallel
symmetric correction
25
damp=0.3
57.7
9.67
15.54
16.06
13.76
0%
Denoising reg. ( ours )
25
ρ=0.25,δ=0.0125
58.8
9.11
14.25
14.75
12.70
119%
Seq.
symmetric correction
25
damp=0.5
58.0
9.98
15.66
16.19
13.94
−20%
asymmetric correction
25
β=0.3333
56.2
9.05
14.55
15.01
12.87
100%
Appendix
Table 10: QuIP# at W2A16 (2-bit), Llama-3-8B. Columns and blocks as in Table 8 .
Figure 19: QuIP# W2A16 hyperparameter sweeps, Llama-3-8B. Panels as in Figure 17 .
Method
Grid
Setting
Acc ↑
Wiki ↓
C4 ↓
FW ↓
Avg ↓
Gap ↑
FP16
–
–
66.5
4.88
5.80
6.49
5.73
–
Parallel
symmetric correction
25
damp=0.6
61.0
5.90
7.81
8.12
7.28
–
Denoising reg. ( ours )
25
ρ=0.25,δ=0.0125
61.0
5.81
7.72
8.05
7.19
–
Seq.
symmetric correction
25
damp=0.2
60.6
5.98
8.10
8.37
7.48
–
asymmetric correction
25
β=0.375
59.3
5.81
7.89
8.16
7.28
–
Appendix
Table 11: QuIP# at W2A16 (2-bit), Llama-2-13B. Columns and blocks as in Table 8 . Here parallel-symmetric and sequential-asymmetric nearly coincide.
Figure 20: QuIP# W2A16 hyperparameter sweeps, Llama-2-13B. Panels as in Figure 17 .
Figure 21: Injected-noise scale γ sweep (Llama-3.2-1B W3A16, λ=0.4 , ρ=0.3 , random Hadamard rotation). Average perplexity against γ over ten log-spaced values in [0.01,0.2] .
Figure 22: Isotropic vs. per-channel Gaussian noise (Llama-3.2-1B, δ=0.025 , random Hadamard rotation, γ=0.05 ). Average perplexity against the target noise scale ρ ; grey dashed marks the parallel baseline ( 0% gap) and red dotted the sequential baseline ( 100% ). Left (W3 GPTQ, α=0.01 ): both isotropic and per-channel drop below the parallel baseline. Right (W2 QuIP#, α=0.3 ): only per-channel improves over the baseline; isotropic noise stays above it throughout.
Quantization objective
Llama 1B
Llama 3B
Llama 8B
Llama 13B
Acc ↑
Avg ↓
Gap ↑
Acc ↑
Avg ↓
Gap ↑
Acc ↑
Avg ↓
Gap ↑
Acc ↑
Avg ↓
Gap ↑
FP16
54.8
12.12
–
62.6
9.85
–
67.1
8.28
–
66.5
6.15
–
3-bit GPTQ
Par.
symmetric correction
49.4
19.00
0%
58.5
13.37
0%
63.9
11.10
0%
64.2
6.92
0%
Denoising reg. ( ours )
49.3
18.32
34%
57.1
13.22
21%
64.9
10.96
32%
64.5
6.90
12%
Seq.
symmetric correction
49.8
18.30
35%
59.0
13.24
18%
63.5
10.94
36%
63.7
6.89
13%
asymmetric correction
50.3
17.00
100%
58.4
12.65
100%
63.8
10.65
100%
64.5
6.72
100%
Appendix
Table 12: Held-out generalization ablation. Same objectives, models, and sweeps as Table 1 , every configuration is selected by WikiText-2 test perplexity, but the “Avg” column reports the average of the two held-out sets (C4 and FineWeb), which are unseen by the selection. The conclusions match Table 1 , showing that the qualitative improvements persist under held-out evaluation.
Figure 23: Effective bit-width vs. perplexity on Llama-3.2-1B/3B and Llama-3-8B (W3A16). Same plot as Figure 4 (right) for Llama-3.2-3B, extended to Llama-3.2-1B and Llama-3-8B.
Figure 24: Per-model sequential-prefix diagnostics (W3A16), one row per model. The first K blocks are quantized sequentially (GPTAQ, β=0.1 ) and the remaining ones in parallel (GPTQ), at fixed hyperparameters. Left: avg. perplexity vs. K ; the gray line interpolates linearly between K=0 and K=L , so a curve below it means the gains are front-loaded. Middle: per-block residual-stream SNR for K∈{0,L/2,L} . Right: per-block SNR gain of sequential over parallel quantization.
Figure 25: Projected quantization time vs. number of GPUs on Llama-2-13B (W2A16 QuIP#). Projection from measured single-GPU times. Parallel quantization and our parallel denoising distribute the per-block rounding across GPUs, whereas the sequential asymmetric correction cannot be parallelized.
Method
Rot.
Llama 1B
Llama 3B
Llama 8B
Llama 13B
Quant
Eval
Quant
Eval
Quant
Eval
Quant
Eval
Parallel GPTQ
×
2.5
1.0
5.6
1.9
11.6
3.4
16.6
5.2
✓
2.5
1.0
5.6
1.9
11.6
3.4
16.6
5.2
Gaussian (ours)
×
5.7
0.9
13.1
1.9
29.4
3.5
43.2
5.2
✓
5.8
1.0
13.1
1.9
29.4
3.5
43.3
5.2
W3 proxy (ours)
×
8.2
1.0
19.0
1.9
42.4
3.5
61.2
5.2
Appendix
Table 13: Per-row runtime on a single A100 (minutes, mean over each sweep), split into quantization (Quant) and perplexity evaluation (Eval), for Llama-3.2-1B/3B, Llama-3-8B, and Llama-2-13B at W3A16. Rot.: whether the rotation is a learned SpinQuant rotation ( × means random rotation).
Schedule
Llama 1B
Llama 3B
Llama 8B
Llama 13B
Quant
Eval
Quant
Eval
Quant
Eval
Quant
Eval
Parallel symmetric
23.6
0.5
64.7
1.1
163.3
2.3
138.9
1.4
Denoising reg. (ours)
27.3
0.5
73.9
1.1
183.2
2.3
145.1
1.4
Sequential symmetric
23.8
0.5
65.3
1.1
164.1
2.3
140.1
1.4
Sequential asymmetric
25.4
0.5
69.2
1.1
173.3
2.3
142.1
1.4
Appendix
Table 14: Per-schedule runtime for the QuIP# W2A16 sweeps (minutes, mean over each sweep). A100 for 1 B/ 3 B/ 8 B, H200 for 13 B. QuIP#’s rounding is more expensive than GPTQ’s ( table 13 ).
Model
Clean (s)
Noisy (s)
Overhead (s)
Noisy/Clean
Llama-3.2-1B
5.1
7.1
2.0
1.40
Llama-3.2-3B
12.3
16.6
4.3
1.35
Llama-3-8B
25.7
33.3
7.6
1.29
Llama-2-13B
42.9
53.5
10.6
1.25
Appendix
Table 15: Clean vs. noisy calibration forward-pass time on a single A100 (seconds, n=128 sequences, seqlen 2048 ). Clean is the full-precision forward pass every parallel method runs to build its moments. Noisy is the additional pass with injected per-channel Gaussian noise that our denoising objective needs. Overhead = Noisy − Clean is the extra cost of our second pass.
Post-training quantization compresses large language models (LLMs) by storing their weights at reduced precision, and each quantized weight introduces an error into the hidden states. Naively, these errors should accumulate with depth and corrupt next-token prediction; randomly initialized models accumulate these discrepancies rapidly, whereas quantized pretrained models accumulate much less hidden-state error and largely maintain downstream task performance, even though they were never trained with quantization noise. This raises the question we address: why does post-training quantization work? Comparing full-precision and quantized forward passes, we identify two mechanisms that characterize pretrained quantization robustness. First, the error a layer newly introduces tends to oppose the error it inherits from the layer's input. The two cancel partially such that the discrepancy between full-precision and quantized passes grows slowly. This counteracting residual interaction develops during pretraining. Our quantitative analysis identifies it as a major factor slowing hidden-error growth. Second, LM-head geometry preferentially preserves the scores and probabilities of high-ranked tokens, which typically represent the model's most confident predictions. Together, these mechanisms explain why quantization error that passes through numerous layers can still produce only small output changes, and we verify the findings across models and quantization settings.
Yuxiang Chen, Michael Beyer, Jun Zhu +1
Dept. of Comp. Sci. and Tech., Institute for AI, BNRist Center, THBI Lab, Tsinghua-Bosch Joint ML Center, Tsinghua University · College of AI, Tsinghua University · Bosch AI Research, Renningen, Germany
Post-training quantization (PTQ) is essential for efficient large language model inference, but reliably quantizing activations remains challenging when weights, activations, and KV caches are all quantized to 4-bit precision. A key difficulty lies in massive activations, whose extreme values dominate the activation range and amplify quantization errors. State-of-the-art methods mainly mitigate massive activations through transformation-based smoothing, such as orthogonal rotations and affine scaling, but overlook the cross-layer dynamics of the residual stream. In this paper, we show that massive activations emerge and disappear in a phase-wise pattern across network depth, triggering large residual changes. These changes cause newly injected layer-wise updates to dominate the 4-bit quantization scale and weaken historical residual information. To characterize this behavior, we introduce Jump Ratio and Historical Feature SNR. This suggests that static transformation-based smoothing cannot fully resolve dynamic quantization instability caused by cross-layer residual changes. Based on this analysis, we propose DynamicPTQ, a Dynamic Post-Training Quantization policy for phase-aware mixed-precision activation quantization. DynamicPTQ identifies quantization-sensitive layers from residual-stream dynamics and assigns 8-bit activation precision only to these layers, while keeping weights, KV caches, and other activations in 4-bit precision. It can be directly integrated with strong PTQ baselines such as QuaRot, SpinQuant, and FlatQuant. Experiments on LLaMA-2 and LLaMA-3 show that DynamicPTQ consistently improves perplexity and zero-shot QA performance under W4A4KV4 quantization, while achieving 1.05 to 1.07 times throughput improvement with modest memory overhead. These results demonstrate a practical path toward robust low-bit LLM inference.
Zimo Zhao, Maolin Wang, Bowen Yu +3
City University of Hong Kong · Hong Kong, China · Zhejiang University of Technology +1
Post-training quantization has emerged as a widely adopted technique for compressing and accelerating the inference of Large Language Models (LLMs). The primary challenges in LLMs quantization stem from activation outliers, which significantly degrade model performance especially at lower bit precision. While recent approaches attempt to mitigate outliers through linear transformations across feature dimensions, our analysis reveals that the transformed weights and activations still exhibit persistent outlier patterns with concentrated magnitude distributions. In this paper, we first model the mathematical relationship between quantization error and outliers, and then introduce a new metric Flatness to quantify the distribution of outliers. Based on this, we derive the theoretical optimal solution with respect to Flatness. Building on these insights, we propose Bidirectional Diagonal Quantization (BDQ), a novel post-training quantization framework that effectively disperses outlier patterns through optimized matrix transformations. BDQ strategically distributes outlier magnitudes across matrix dimensions via learned diagonal operations. Extensive experiments demonstrate that BDQ establishes a new quantization benchmark. It achieves less than 1% accuracy drop in W4A4 quantization on the LLaMA-3-8B model. In the more challenging W2A4KV16 experiment, compared to state-of-the-art approaches, BDQ reduces the performance gap by 39.1% on the DeepSeek-R1-Distill-LLaMA-70B model.
Xiusheng Huang, Zhe Li, Xuanwu Yin +5
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China · School of Artificial Intelligence, University of Chinese Academy of Sciences · Beijing Academy of Artificial Intelligence +2