We introduce ShamAN-Q, a sub-1-bit post-training quantization method that extends NanoQuant by replacing each its diagonal reconstruction geometry with a tractable dense curvature metric, using a general paradigm popularized by the Shampoo optimizer. For each linear weight, ShamAN-Q fits a Kronecker product to the empirical Fisher information matrix of a small calibration set by Kullback--Leibler minimization, forming a Mahalanobis reconstruction loss from the result. The continuous ADMM updates from NanoQuant become solutions to Sylvester equations, while its discrete projection and deployment format remain unchanged. Because the curvature is local to a given set of weights, ShamAN-Q re-measures the input curvature statistic for each layer immediately before layer factorization, periodically refreshing all statistics on the partially quantized model. ShamAN-Q also redistributes the uniform rank from NanoQuant across layers at the same total number of bits. On Qwen3-Base, ShamAN-Q lowers WikiText-2 perplexity at ≈1 bpw from 27.56 to 22.96 (0.6B), 19.21 to 16.72 (1.7B), and 14.29 to 13.80 (4B) while matching or improving zero-shot accuracy on the Eleuther LM Evaluation Harness. On 0.6B, ShamAN-Q at ≈0.8 bpw matches the published perplexity of NanoQuant at ≈1.0 bpw.
Figures & tables
Figure 1: WikiText-2 perplexity (PPL) (left, lower is better) and zero-shot mean on reasoning tasks in Eleuther LM Evaluation Harness (right, higher is better) vs. Qwen3-Base size at a 1.0 bpw budget. ShamAN-Q matches or improves on NanoQuant in both PPL and zero-shot mean at every size. All points are from Table 1 ; only reported results are plotted.
Figure 2: ShamAN-Q end to end. Phase 1 runs the calibration set through the frozen teacher, fits and shrinks the curvature statistics (Aγ,Gγ) of every linear layer, and allocates ranks at bit parity. Phase 2 processes decoder blocks in order. The factorized prefix and the untouched full-precision suffix are frozen while block b is tuned ( TuneFP ), factorized layer by layer by Kron–LB–ADMM under the transported statistics (A,G) (Figure 3 ), and refined by latent STE. The statistics of the unfactorized layers are refreshed on the partially quantized model after every Δ blocks. Phase 3 distills only the scale vectors against the teacher, with the sign matrices frozen. Flames mark trained parameters, snowflakes frozen ones, and blue blocks denote differences between ShamAN-Q and NanoQuant.
Figure 3: ShamAN-Q pipeline for one layer. Calibration fits and shrinks the curvature statistics, yielding the transported metric (A,G) and diagonal maps (Din,Dout) . Preconditioning maps W to W . Kron–LB–ADMM returns the continuous weight factors (U,V) and scale vectors (s1,s2) . The transported statistics alter only the continuous U/V optimization component; Euclidean consensus leaves SVID, dual updates, scale extraction, and sign extraction unchanged from NanoQuant.
PPL ↓
0-shot mean ↑
bpw
Method
0.6B
1.7B
4B
8B
14B
0.6B
1.7B
4B
8B
14B
≈ 1.0
NanoQuant (paper)
27.56
19.21
14.29
12.47
10.92
–
–
–
0.4894
–
NanoQuant (repr.)
29.21
18.76
14.86
12.38
11.49
0.387
0.427
0.455
0.474
0.498
ShamAN-Q
22.96
16.72
13.80
11.82
10.91
0.427
0.448
0.464
0.497
0.506
≈ 0.8
NanoQuant (paper)
33.79
25.31
19.33
14.83
12.88
–
–
–
–
–
NanoQuant (repr.)
42.22
22.37
17.99
15.63
–
0.387
0.414
0.431
0.443
–
Table 1: WikiText-2 perplexity and zero-shot mean across a suite of reasoning tasks for < 1 bpw Qwen3-Base models at the same total bits as NanoQuant ( Chong et al., 2026 ) , with both reported and reproduced values. The upper block uses a 1.0 bpw budget, the middle block a 0.8 bpw budget, and the lower block a 0.55 bpw budget. Unreported results are marked with –.
Table 5
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Component
NanoQuant diagonal geometry
ShamAN-Q dense geometry
Curvature state
Θ(m+n)
Θ(m2+n2) statistics and eigensystems
Calibration fit
Θ(N(m+n))
Θ(N(m2+n2)+m3+n3) per pass, three passes
Fresh input statistic
–
Θ(Nn2) per shared-input group
Refresh
–
one fit pass over unfactorized layers every Δ blocks
Rank probe
–
three K=50 ADMM solves per layer, once
Setup
Θ(mn)
Θ(m3+n3+m2n+mn2)
Appendix
Table 5: Per-layer costs for an m×n weight, factor rank r , and N calibration tokens.
Large Language Models (LLMs) are widely used across many domains, but their scale makes deployment challenging. Post-Training Quantization (PTQ) reduces memory footprint without retraining by leveraging a small calibration set. Recent Hessian-based PTQ methods compensate quantization error via cross-channel dependencies, but such approaches degrade at low bit-widths due to noisy curvature estimates from limited calibration data. We propose DASH-Q, a robust PTQ framework using diagonal Hessian approximation and iterative weighted least squares. By discarding noise-prone dependencies, DASH-Q filters sampling noise while prioritizing the preservation of salient feature power. We outperform other PTQ baselines in ultra low-bit regime, improving zero-shot accuracy by 7.01% on average and up to 14.01% over the strongest baselines across five baseline LLM models, while showing robust and stable performance with very small calibration data.
Post-training quantization (PTQ) has become an important technique for reducing the inference cost of Large Language Models (LLMs). While recent mixed-precision methods improve ultra-low bit quantization by preserving critical subspaces in high precision, they typically construct these subspaces relying solely on activation statistics. This ignores the fundamental nature of linear operations, where the output perturbation is jointly driven by both activation and weight quantization noise. In this paper, we propose CoQuant, a joint weight-activation subspace projection method. By theoretically modeling the expected output error, CoQuant formulates a closed-form weighted PCA solution that balances activation and weight covariances to select the optimal high-precision subspace. Extensive experiments on Llama-3.2 and Qwen2.5 models show that CoQuant consistently outperforms strong PTQ baselines in both WikiText perplexity and zero-shot common-sense reasoning accuracy. These results demonstrate that joint weight-activation subspace modeling provides a principled and effective direction for low-bit LLM quantization. The source code is available at https://github.com/Zachary5895/CoQuant.
Zhe Ding, Su Pan, Duowei Pan
School of Internet of Things, Nanjing University of Posts and Telecommunications, Nanjing, China · Amazon AGI, Seattle, WA, USA
Post-training quantization (PTQ) is a widely adopted technique for compressing large language models (LLMs) without retraining. Existing second-order PTQ methods, including GPTQ, construct quantization objectives exclusively from input activation statistics, effectively assuming that all output channels contribute equally to the layer-wise reconstruction objective. We propose KronQ, a PTQ framework that challenges this assumption by introducing the gradient covariance into the quantization pipeline. Under the Kronecker-factored Hessian approximation, the quantization loss depends jointly on both the activation and gradient covariances, and KronQ exploits this at two complementary levels. (1) KronQ introduces bidirectional incoherence processing, extending the existing input-side random rotation to the output dimension using the gradient covariance, reducing weight magnitude variance across both input and output dimensions. (2) KronQ derives a new sensitivity metric for inter-layer mixed-precision allocation, driven by the gradient and activation Hessian traces. Notably, in the case of 2-bit weight-only quantization on LLaMA-3-70B, while GPTQ and GPTAQ diverge or produce degenerate quantizations (>2000 perplexity on WikiText-2), KronQ achieves 7.93 perplexity.
Donghyun Lee, Yuhang Li, Ruokai Yin +1
University of Southern California · Yale University