Trellis-coded quantization enables high-dimensional compression of large language model (LLM) weights at ultra-low bit widths without the exponentially large codebooks required by conventional vector quantization. Practical deployment, however, presents two challenges: reconstructing compressed weights at sufficient parallel throughput to avoid making dequantization an inference bottleneck, and maintaining quantization accuracy without costly incoherence transformations. We address these challenges with two complementary techniques. First, we introduce an ultra-low-complexity trellis dequantizer that uses a structured, hardware-efficient state-to-value mapping while preserving diverse reconstruction choices for trellis search. Second, we reformulate discrete trellis path optimization with a curvature-aware objective that reflects model sensitivity directly in the original coordinate space. Together, these techniques enable high-quality ultra-low-bit trellis quantization with inexpensive, highly parallel runtime reconstruction and without relying on Hadamard-based incoherence processing.
Figures & tables
Figure 1: Overview of XOR-Trellis. A structured state-to-value mapping provides diverse FP4 reconstruction choices, while Hessian-aware Viterbi search weights errors using the diagonal factor D . The resulting approximately 2-bit trellis path is decoded with the same lightweight mapping.
PPL ↓
0-Shot Raw Accuracy (%) ↑
Model
Method
Rotation
WT2
C4
ArcC
ArcE
BoolQ
HSwag
PiQA
Avg.
Llama-3.1-8B
BF16
–
5.84
8.43
51.28
81.61
82.05
60.06
80.14
71.03
QTIP-3INST
Hadamard
8.48
11.93
39.85
73.15
78.47
51.99
75.63
63.82
XOR-Trellis
Hadamard
8.10
11.45
39.80
73.15
75.26
53.61
76.88
63.74
XOR-Trellis
None
8.69
12.23
37.97
68.69
76.06
51.13
75.24
61.82
XOR-Trellis + D-weighted
None
8.06
11.40
38.48
72.39
79.63
53.30
76.82
64.12
Table 1: Perplexity (PPL) is evaluated on WikiText-2 (WT2) and C4 with context lengths of 4096 for Llama-3.1-8B, 2048 for Llama-3.1-8B-Instruct, and 8192 for Llama-3.1-70B-Instruct. Zero-shot raw accuracy ( acc , %) is reported on five downstream tasks.
PPL ↓
0-Shot Accuracy (%) ↑
Method
Rotation
Hessian
WT2
C4
ArcC
ArcE
BoolQ
HSwag
PiQA
Avg.
BF16
–
–
7.22
10.44
51.71
81.82
84.10
59.05
79.92
71.32
YAQA-3INST
Hadamard
YAQA-2S
9.57
13.75
44.71
77.57
83.46
52.35
76.17
66.85
YAQA-3INST
None
YAQA-2S
15511.81
8986.10
21.93
25.59
37.83
25.81
53.48
32.93
YAQA-3INST + D-weighted
None
YAQA-2S
6500.82
3745.74
20.90
26.47
37.83
26.14
52.39
32.75
XOR-Trellis
Hadamard
YAQA-2S
9.00
12.90
46.16
78.70
83.82
53.85
77.20
67.95
Table 2: Weight-only PTQ results on Llama-3.1-8B-Instruct using the two-sided YAQA Hessian. Perplexity (PPL) is evaluated on WikiText-2 (WT2) and C4 with a context length of 2048, and zero-shot raw accuracy ( acc , %) is reported on five downstream tasks.
Method
Rotation
WT2 ↓
ArcC
ArcE
BoolQ
HSwag
PiQA
Avg.
BF16
–
5.12
43.43
76.35
77.74
57.11
78.07
66.54
QTIP-3INST
Hadamard
6.95
34.56
69.53
71.93
48.15
74.48
59.73
LLVQ (shape-gain, 2-bit gain)
None
7.27
–
–
–
–
–
–
LLVQ (shape-gain, 2-bit gain)
Input
6.90
–
–
–
–
–
–
LLVQ (shape-gain, 2-bit gain)
Input + Output
6.83
35.50
69.80
73.00
49.70
75.20
60.64
LLVQ (shape-gain, 0-bit gain)
Input + Output
6.48
–
–
–
–
–
–
Table 3: Comparison with LLVQ on Llama-2-7B. Perplexity (PPL) is evaluated on WikiText-2 (WT2) with a context length of 4096, and zero-shot raw accuracy ( acc , %) is reported where available. All quantized results are reported without layerwise or end-to-end finetuning.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
PPL ↓
0-Shot Raw Accuracy (%) ↑
Model
Method
Rotation
WT2
C4
ArcC
ArcE
BoolQ
HSwag
PiQA
Avg.
Llama-3.2-1B-Instruct
BF16
–
13.16
18.58
35.67
68.81
69.39
45.11
74.05
58.61
QTIP-3INST
Hadamard
24.28
30.44
26.19
55.56
62.23
36.32
67.08
49.48
XOR-Trellis
Hadamard
22.89
28.95
25.77
56.14
60.67
37.02
66.43
49.21
XOR-Trellis
None
30.04
36.35
23.89
48.44
63.21
34.74
63.44
46.74
XOR-Trellis + D-weighted
None
23.04
28.86
26.28
57.11
63.00
38.09
67.57
50.41
Appendix
Table 4: Weight-only PTQ results on Llama-3.2-1B-Instruct and Llama-3.2-3B-Instruct. Perplexity (PPL) is evaluated on WikiText-2 (WT2) and C4 with a context length of 2048, and zero-shot raw accuracy ( acc , %) is reported on five downstream tasks. Lower PPL and higher accuracy indicate better performance.
PPL ↓
0-Shot Raw Accuracy (%) ↑
Model
Method
Rotation
WT2
C4
ArcC
ArcE
BoolQ
HSwag
PiQA
Avg.
Llama-2-7B
BF16
–
5.12
6.63
43.43
76.35
77.74
57.11
78.07
66.54
QTIP-3INST
Hadamard
6.95
9.08
34.56
69.53
71.93
48.15
74.48
59.73
XOR-Trellis
Hadamard
6.65
8.75
36.52
71.46
70.34
49.83
73.72
60.37
XOR-Trellis
None
7.34
10.07
25.17
53.87
68.41
35.87
63.38
49.34
XOR-Trellis + D-weighted
None
6.41
8.52
37.12
71.17
75.05
51.16
74.27
61.75
Appendix
Table 5: Weight-only PTQ results on Llama-2-7B and Llama-2-13B. Perplexity (PPL) is evaluated on WikiText-2 (WT2) and C4 with a context length of 4096, and zero-shot raw accuracy ( acc , %) is reported on five downstream tasks. Lower PPL and higher accuracy indicate better performance.
PPL ↓
0-Shot Accuracy (%) ↑
Method
Rotation
WT2
C4
ArcC
ArcE
BoolQ
HSwag
PiQA
Avg.
BF16
–
9.00
12.48
55.46
83.54
86.54
57.04
76.55
71.83
QTIP-3INST
Hadamard
11.35
14.97
48.63
78.49
85.90
49.93
74.65
67.52
XOR-Trellis
Hadamard
10.29
14.23
47.70
78.28
84.01
51.30
74.65
67.19
XOR-Trellis
None
11.57
15.34
40.61
70.12
82.54
50.33
71.76
63.07
XOR-Trellis + D-weighted
None
10.34
14.38
46.93
78.07
83.70
51.09
74.59
66.88
Appendix
Table 6: Weight-only PTQ results on Qwen3-8B. Perplexity (PPL) is evaluated on WikiText-2 (WT2) and C4 with a context length of 4096, and zero-shot raw accuracy ( acc , %) is reported on five downstream tasks. Lower PPL and higher accuracy indicate better performance.
PPL ↓
0-Shot Accuracy (%) ↑
Method
Rotation
WT2
C4
ArcC
ArcE
BoolQ
HSwag
PiQA
Avg.
BF16
–
4.91
7.42
50.17
80.81
83.73
61.25
80.52
71.30
QTIP-3INST
Hadamard
6.04
8.77
42.75
75.67
82.72
54.62
78.24
66.80
XOR-Trellis
Hadamard
5.68
8.41
43.60
76.81
83.06
55.99
79.22
67.74
XOR-Trellis
None
6.71
9.78
38.74
73.11
80.64
52.86
76.61
64.39
XOR-Trellis + D-weighted
None
5.71
8.41
44.88
76.73
82.63
56.17
78.62
67.81
Appendix
Table 7: Weight-only PTQ results on Mistral-7B. Perplexity (PPL) is evaluated on WikiText-2 (WT2) and C4 with a context length of 4096, and zero-shot raw accuracy ( acc , %) is reported on five downstream tasks. Lower PPL and higher accuracy indicate better performance.
Matrix-level outcome
ρ
95% CI
no-RHT diagonal- D gain GD
0.545
[0.454, 0.631]
legacy RHT gain GRHT
0.629
[0.537, 0.718]
absolute legacy error reduction, EnoRHT−ERHT
0.606
[0.535, 0.676]
no-RHT diagD advantage over legacy RHT, log(ERHT,legacy/EnoRHT,diagD)
−0.291
[ −0.414 , −0.172 ]
Appendix
Table 8: Spearman correlations with no-RHT robust anisotropy logAD . Intervals use complete-layer bootstrap samples. The absolute error delta and direct no-RHT-diagD comparison are post-hoc diagnostics.
Figure 2: Local full-Hessian error effects versus no-RHT robust diagonal anisotropy. (a) RHT reduces legacy error by a larger factor for more anisotropic matrices ( ρ=0.629 ). (b) The corresponding absolute relative-error reduction follows the same trend ( ρ=0.606 ). Positive values in (a) and (b) favor RHT. (c) Positive values favor no-RHT diagonal- D over legacy RHT. Its negative association ( ρ=−0.291 ) shows that diagonal anisotropy alone does not determine which complete method wins.
Metric
r=00
r=01
r=10
r=11
Pooled
Signed correlation
0.000000
0.000000
0.000000
0.000000
0.000000
Absolute-value correlation
0.000000
0.000000
0.000000
0.000000
0.000000
Appendix
Table 9: Exhaustive transition-correlation results for XOR-Trellis. Values are computed over all 216 states.
Property
Value
Count per payload code
4096
Mean
0
Standard deviation
2.9262
Mean absolute value
2.2500
P(q=0)
12.5%
P(∣q∣≥4)
25.0%
Appendix
Table 10: Additional exhaustive properties of the XOR-Trellis generator.
Figure 3: Gate-level implementation of the XOR-Trellis generator. Four masked parity functions are computed from the 14-bit history using XOR2 gates, producing the two-bit palette selector and two permutation bits. Together with the two transition-slot bits, these signals index a fixed 4×4 palette lookup to produce one of the 16 FP4 payload codes. The parity logic requires 14 XOR2 gates before the lookup, resulting in a compact combinational implementation.
Deploying large language models (LLMs) is challenging due to their massive parameters and high computational costs. Ultra low-bit quantization can significantly reduce storage and accelerate inference, but extreme compression (i.e., mean bit-width <= 2) often leads to severe performance degradation. To address this, we propose Squeeze10-LLM, effectively "squeezing" 16-bit LLMs' weights by 10 times. Specifically, Squeeze10-LLM is a staged mixed-precision post-training quantization (PTQ) framework and achieves an average of 1.6 bits per weight by quantizing 80% of the weights to 1 bit and 20% to 4 bits. We introduce Squeeze10LLM with two key innovations: Post-Binarization Activation Robustness (PBAR) and Full Information Activation Supervision (FIAS). PBAR is a refined weight significance metric that accounts for the impact of quantization on activations, improving accuracy in low-bit settings. FIAS is a strategy that preserves full activation information during quantization to mitigate cumulative error propagation across layers. Experiments on LLaMA and LLaMA2 show that Squeeze10-LLM achieves state-of-the-art performance for sub-2bit weight-only quantization, improving average accuracy from 43% to 56% on six zero-shot classification tasks--a significant boost over existing PTQ methods.
Qingcheng Zhu, Yangyang Ren, Linlin Yang +6
Beihang University · Communication University of China · Beijing Jiaotong University
Large language models (LLMs) have driven major progress in NLP, yet their substantial memory and compute demands still hinder practical deployment. Binarization can compress weights to 1 bit, fundamentally lowering compute and bandwidth cost. However, existing methods cannot address activation heavy tails and thus must keep activations in high precision, preventing true end-to-end acceleration. To overcome this limitation, we propose BWLA (Binarized Weights and Low-bit Activations), the first post-training quantization framework that preserves high accuracy while achieving 1-bit weight quantization together with low-bit activations (e.g., 6 bits). The Orthogonal-Kronecker Transformation (OKT) learns an orthogonal mapping via EM minimization, converting unimodal weights into symmetric bimodal forms while suppressing activation tails and incoherence. The Proximal SVD Projection (PSP) then performs lightweight low-rank refinement through proximal SVD projection, further enhancing quantizability with minimal overhead. On Qwen3-32B, BWLA reaches a Wikitext2 perplexity of 11.92 under 6-bit activations (vs. 38 from SOTA), improves five zero-shot tasks by more than 70%, and delivers 3.26 times inference speedup, demonstrating strong potential for real-world LLM compression and acceleration.
Scalar quantization of large language models (LLMs) is fundamentally limited by information-theoretic bounds. While vector quantization (VQ) overcomes these limits by encoding blocks of parameters jointly, practical implementations must avoid the need for expensive lookup mechanisms or other explicit codebook storage. Lattice approaches address this through highly structured and dense packing. This paper explores the Leech lattice, which, with its optimal sphere packing and kissing configurations at 24 dimensions, is the highest dimensional lattice known with such optimal properties. To make the Leech lattice usable for LLM quantization, we extend an existing search algorithm based on the extended Golay code construction, to i) support indexing, enabling conversion to and from bitstrings without materializing the codebook, ii) allow angular search over union of Leech lattice shells, iii) propose fully-parallelisable dequantization kernel. Lastly, we provide a geometric reinterpretation of combining shape--gain quantization with GPTQ-style Hessian corrections: the standard scale-correction step of shape--gain acts as a retraction onto a product of spheres, yielding a Spherical GPTQ primarily acting on directions. We find that low-angular-distortion LLVQ reduces sensitivity to Hadamard/rotation preprocessing, and enables a strong Hadamard-free PTQ in practice. LLVQ delivers state-of-the-art LLM quantization performance, outperforming recent methods such as Quip#, QTIP, and PVQ. The results highlight the effectiveness of high-dimensional lattices for scalable, theoretically grounded model compression.
Tycho F. A. van der Ouderaa, Mart van Baalen, Paul Whatmough +1