Extreme low-bit LLM compression is most challenging when weights, activations, and KV caches are quantized together: their distributions differ, and quantization errors interact throughout the network. We introduce CanonQ, a unified quantization-aware training framework that addresses these challenges by separating source canonicalization from task-aware adaptation. Fixed rotations and energy normalization map heterogeneous tensor sources to canonical coordinates, enabling frozen Gaussian-reference codebooks to be reused across layers and models. Joint training then adapts the network to the coupled errors of weight, activation, and cache quantization within a common scalar/vector interface. We bound frozen-codebook transfer error and local task loss, and derive an exact normalization-aware straight-through Jacobian that links quantization distortion to gradient bias. The strongest gains arise under joint W2A4KV2 compression: across LLaMA3-1B/3B/8B, CanonQ-Omni achieves up to 14.28x lower WikiText-2 perplexity and up to 57.9% higher mean zero-shot accuracy than prior state-of-the-art and representative quantization baselines. The benefits extend to Qwen3-1.7B, code generation, and mathematical reasoning: on instruction-tuned MobileLLM-Pro-1B at W2A16KV16, CanonQ achieves relative improvements of 41.7% in HumanEval pass@1 and 39.1% in GSM8K exact match over the strongest evaluated quantization baseline.
Figures & tables
Figure 1: Schematic overview of CanonQ . Fixed source-specific transforms and normalization map tensor blocks to canonical coordinates. Frozen Gaussian-reference codebooks are reused across layers and models for matching (b,d) . Joint QAT adapts the model to coupled quantization errors at W2A4KV2 .
W2A8KV2
W2A4KV2
Method
PPL ↓
Acc. ↑
PPL ↓
Acc. ↑
FP Model
9.75
58.54
9.75
58.54
RTN (PTQ)
5.8e5
37.31
5.5e5
36.72
GPTQ (PTQ)
5.3e4
38.23
9.5e4
37.83
QuaRot (PTQ)
3.4e4
37.39
4.1e4
38.11
KIVI
1.1e5
34.85
3.2e5
35.16
Table 1: Joint-compression and cross-model results. Columns report PPL ↓ and eight-task accuracy ↑ . The suffix denotes dW ; canonical A/K/V paths use d=1 . CanonQ-Omni uses canonical KV quantization in the headline KV2 results. Full task scores appear in Tabs. 11 , 12 and 13 .
Benchmark
Metric
Tasks
FP base
ParetoQ
VQLLM-QAT-d4
CanonQ-d1
MBPP
pass@1
500
46.4 (232)
28.0 (140)
27.8 (139)
32.2 (161)
HumanEval
pass@1
164
60.4 (99)
43.9 (72)
33.5 (55)
62.2 (102)
GSM8K
EM
1319
53.8 (709)
34.3 (453)
32.7 (431)
47.8 (630)
Table 3: Code generation and mathematical reasoning on MobileLLM-Pro-Instruct-1B W2A16KV16 . Scores are percentages, with correct-task counts in parentheses. MBPP/HumanEval report greedy pass@1; GSM8K reports final-answer exact match (EM). Quantized settings share the teacher and evaluation step. GSM8K also guides learning-rate selection; see Sec. A.5 .
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Source
Index bits
Group g
Stored bits, r=1
Stored bits, r=2
Weights
2
128
2.125
2.250
Activations
4
128
4.125
4.250
Keys / values
2
64
2.250
2.500
Appendix
Table 5: Canonical W2A4KV2 storage before padding and unquantized modules. Each decoding scalar is 16 bits; r counts independent factors after folding. A rates apply only to materialized quantized activation buffers. Global tables are separate.
Figure 2: Paired scalar marginals before and after HDP, using identical normalized groups. Dynamic A/K/V sources approach the Gaussian reference; weights remain near it. These scalar diagnostics illustrate alignment with the shared design reference. Full statistics appear in Tab. 6 .
W1 to N(0,1)
KS to N(0,1)
Source
No rotation
HDP
No rotation
HDP
Weights W
0.015±0.008
0.017±0.012
0.006±0.003
0.010±0.006
Activations A
0.101±0.055
0.010±0.004
0.039±0.019
0.005±0.002
Keys K
0.060±0.017
0.026±0.013
0.035±0.011
0.014±0.006
Values V
0.030±0.014
0.018±0.008
0.016±0.006
0.009±0.003
Appendix
Table 6: Paired scalar-marginal distances to N(0,1) for a trained LLaMA3-1B W2A4KV2 checkpoint. Values are layer means ± standard deviations, not uncertainty across training seeds. Both conditions use the same groups and normalization.
Format
Method / cache path
PPL ↓
Avg. ZS ↑
FP
Full precision
9.75
58.54
W2A4
HDP+LM PTQ
7043.43
38.83
W2A4
CanonQ-Omni-d1 QAT
13.84
54.71
W2A4KV2
HDP+LM PTQ , uniform KV
8868.91
36.76
W2A4KV2
CanonQ-Omni-d1 QAT, canonical KV
15.42
53.56
Appendix
Table 7: Zero-update canonical PTQ and trained CanonQ-Omni-d1 on LLaMA3-1B . The W2A4 pair shares the W/A codec; the KV2 pair also changes the cache codec.
Format
Cache path
PPL ↓
Avg. ZS ↑
W2A8KV2
uniform
22.82
47.72
canonical
15.70
53.45
W2A4KV2
uniform
21.65
48.30
canonical
15.42
53.56
Appendix
Table 8: Cache-path ablation for LLaMA3-1B CanonQ-Omni-d1 . W/A paths remain canonical; the comparison changes the complete cache path.
Method
ARC-e
ARC-c
BoolQ
PIQA
SIQA
HellaS.
OBQA
Wino.
Avg. ZS ↑
PPL ↓
FP16 Base
64.57
42.85
65.35
75.05
44.91
64.31
50.00
61.25
58.54
9.75
W2A16
RTN (PTQ)
27.19
24.69
54.03
50.00
35.89
25.43
23.05
50.94
36.40
1.2e6
GPTQ (PTQ)
26.95
25.31
61.66
49.12
36.67
25.47
29.69
50.78
38.21
1.2e4
QuaRot (PTQ)
28.12
24.69
62.35
50.34
36.23
26.72
28.12
49.45
38.25
3.2e3
ParetoQ
64.60
39.69
61.05
72.75
42.79
56.74
47.85
58.75
55.53
12.66
Appendix
Table 9: LLaMA3-1B weight-only quantization. All task scores are zero-shot accuracies (%); PPL is measured on WikiText-2 .
Method
ARC-e
ARC-c
BoolQ
PIQA
SIQA
HellaS.
OBQA
Wino.
Avg. ZS ↑
PPL ↓
FP base
64.57
42.85
65.35
75.05
44.91
64.31
50.00
61.25
58.54
9.75
W2A8
RTN (PTQ)
26.91
25.16
51.47
50.29
36.43
25.72
23.05
50.08
36.14
1.2e6
GPTQ (PTQ)
26.68
25.39
62.08
48.39
37.11
25.56
29.69
50.39
38.16
1.3e4
QuaRot (PTQ)
28.71
24.38
62.44
50.15
36.08
26.74
28.91
49.92
38.42
3.2e3
ParetoQ
58.99
35.73
61.56
68.16
41.43
45.72
42.58
54.53
51.09
18.22
Appendix
Table 10: LLaMA3-1B with quantized weights and activations, retaining the floating-point cache. The three panels vary weight and activation precision.
Method
ARC-e
ARC-c
BoolQ
PIQA
SIQA
HellaS.
OBQA
Wino.
Avg. ZS ↑
PPL ↓
FP16 Base
64.57
42.85
65.35
75.05
44.91
64.31
50.00
61.25
58.54
9.75
W2A8KV2
RTN (PTQ)
26.88
22.50
56.16
51.07
37.26
25.30
27.54
51.80
37.31
5.8e5
GPTQ (PTQ)
27.42
26.41
61.24
49.76
37.70
25.72
25.59
52.03
38.23
5.3e4
QuaRot (PTQ)
25.86
24.06
61.90
51.95
36.43
24.81
25.00
49.14
37.39
3.4e4
KIVI
27.73
23.36
37.59
52.34
38.67
26.75
23.83
48.52
34.85
1.1e5
Appendix
Table 11: LLaMA3-1B joint W/A/KV quantization. CanonQ-Omni uses canonical KV quantization; other QAT methods use uniform KV.
Method
ARC-e
ARC-c
BoolQ
PIQA
SIQA
HellaS.
OBQA
Wino.
Avg. ZS ↑
PPL ↓
FP base
72.99
49.84
75.12
78.12
48.29
74.05
53.52
68.91
65.11
7.73
W2A16
RTN (PTQ)
27.11
25.94
47.30
50.34
36.91
25.45
25.78
49.30
36.02
3.6e5
GPTQ (PTQ)
27.34
23.83
61.54
51.12
35.35
26.10
29.49
48.83
37.95
7.3e3
QuaRot (PTQ)
28.01
25.70
56.43
50.24
37.45
25.61
26.17
50.08
37.46
3.3e2
ParetoQ
73.97
50.00
68.42
75.83
46.78
68.30
55.47
63.91
62.84
9.29
Appendix
Table 12: LLaMA3-3B : weight-only, weight–activation, and joint W/A/KV quantization. The same FP reference applies to all three panels. Accuracy is in percent; lower PPL is better.
Qwen3-1.7B, W2A4
Method
ARC-e
ARC-c
BoolQ
PIQA
SIQA
HellaS.
OBQA
Wino.
Avg. ZS ↑
PPL ↓
FP base
68.79
41.48
79.12
71.39
44.73
59.51
38.67
61.09
58.10
16.33
ParetoQ
41.01
25.25
60.76
59.43
41.48
28.67
30.66
51.56
42.35
43.56
CanonQ-Omni-d1
58.11
34.12
62.38
67.13
43.04
46.60
43.95
52.34
50.96
17.80
LLaMA3-8B, W2A4KV2
FP base
80.95
57.88
83.34
81.14
49.34
79.50
55.56
73.43
70.14
6.14
Appendix
Table 13: Cross-family and scale transfer: Qwen3-1.7B W2A4 and LLaMA3-8B W2A4KV2 .
Figure 3: Quality under joint compression on LLaMA3-1B . (a) CanonQ-Omni-d1 retains low PPL as activation and cache precision decrease; the vertical axis is logarithmic. (b) Its W2A4 improvement over ParetoQ spans all eight tasks, with uneven gain magnitudes. Points summarize separately trained configurations; task scores appear in Tab. 10 .
PPL ↓
× FP ↓
Method
2K
4K
8K
2K
4K
8K
FP reference
9.71
9.21
8.74
1.00
1.00
1.00
ParetoQ
50.42
46.67
45.63
5.19
5.07
5.22
CanonQ-W
43.16
40.24
39.50
4.44
4.37
4.52
CanonQ-Omni
16.86
15.55
14.86
1.74
1.69
1.70
Appendix
Table 14: LLaMA3-1B W2A4KV2 : all quantized methods receive 8K continued QAT for 40K updates. WikiText-2 PPL and its ratio to the same-length FP reference.
We introduce ShamAN-Q, a sub-1-bit post-training quantization method that extends NanoQuant by replacing each its diagonal reconstruction geometry with a tractable dense curvature metric, using a general paradigm popularized by the Shampoo optimizer. For each linear weight, ShamAN-Q fits a Kronecker product to the empirical Fisher information matrix of a small calibration set by Kullback--Leibler minimization, forming a Mahalanobis reconstruction loss from the result. The continuous ADMM updates from NanoQuant become solutions to Sylvester equations, while its discrete projection and deployment format remain unchanged. Because the curvature is local to a given set of weights, ShamAN-Q re-measures the input curvature statistic for each layer immediately before layer factorization, periodically refreshing all statistics on the partially quantized model. ShamAN-Q also redistributes the uniform rank from NanoQuant across layers at the same total number of bits. On Qwen3-Base, ShamAN-Q lowers WikiText-2 perplexity at ≈1 bpw from 27.56 to 22.96 (0.6B), 19.21 to 16.72 (1.7B), and 14.29 to 13.80 (4B) while matching or improving zero-shot accuracy on the Eleuther LM Evaluation Harness. On 0.6B, ShamAN-Q at ≈0.8 bpw matches the published perplexity of NanoQuant at ≈1.0 bpw.
Post-training quantization (PTQ) is a widely adopted technique for compressing large language models (LLMs) without retraining. Existing second-order PTQ methods, including GPTQ, construct quantization objectives exclusively from input activation statistics, effectively assuming that all output channels contribute equally to the layer-wise reconstruction objective. We propose KronQ, a PTQ framework that challenges this assumption by introducing the gradient covariance into the quantization pipeline. Under the Kronecker-factored Hessian approximation, the quantization loss depends jointly on both the activation and gradient covariances, and KronQ exploits this at two complementary levels. (1) KronQ introduces bidirectional incoherence processing, extending the existing input-side random rotation to the output dimension using the gradient covariance, reducing weight magnitude variance across both input and output dimensions. (2) KronQ derives a new sensitivity metric for inter-layer mixed-precision allocation, driven by the gradient and activation Hessian traces. Notably, in the case of 2-bit weight-only quantization on LLaMA-3-70B, while GPTQ and GPTAQ diverge or produce degenerate quantizations (>2000 perplexity on WikiText-2), KronQ achieves 7.93 perplexity.
Donghyun Lee, Yuhang Li, Ruokai Yin +1
University of Southern California · Yale University
Large language models (LLMs) have driven major progress in NLP, yet their substantial memory and compute demands still hinder practical deployment. Binarization can compress weights to 1 bit, fundamentally lowering compute and bandwidth cost. However, existing methods cannot address activation heavy tails and thus must keep activations in high precision, preventing true end-to-end acceleration. To overcome this limitation, we propose BWLA (Binarized Weights and Low-bit Activations), the first post-training quantization framework that preserves high accuracy while achieving 1-bit weight quantization together with low-bit activations (e.g., 6 bits). The Orthogonal-Kronecker Transformation (OKT) learns an orthogonal mapping via EM minimization, converting unimodal weights into symmetric bimodal forms while suppressing activation tails and incoherence. The Proximal SVD Projection (PSP) then performs lightweight low-rank refinement through proximal SVD projection, further enhancing quantizability with minimal overhead. On Qwen3-32B, BWLA reaches a Wikitext2 perplexity of 11.92 under 6-bit activations (vs. 38 from SOTA), improves five zero-shot tasks by more than 70%, and delivers 3.26 times inference speedup, demonstrating strong potential for real-world LLM compression and acceleration.