Tetra: Serving Leech-Lattice Quantized LLMs at 2.7 Bits per Parameter
Organizations: Scub, Bordeaux, France
Abstract
Leech-lattice quantization gives good quality at two bits per weight, but its codebooks hold more than 10^14 points, too many for a lookup table. Our earlier kernel expanded the codes at load time and read 4.804 bits per weight from GPU memory for 2 bits of code. We present Tetra, a new codebook on the same lattice. A 24-weight block still takes 48 bits, most of which index a 64-state trellis of the Golay code and one shared 16 KiB table. The kernel decodes a block with six table loads and two small lookups inside the matrix-vector product, and reads 2.148 bits per weight. For full models, we retrain one scale per matrix row, store the matrices that lose the most as 4-bit integers, and pay for them with 4-bit embedding tables. Our Qwen3-4B, 8B and 14B files hold 2.73, 2.70 and 2.73 bits per parameter over the whole model. They score 63.37, 69.58 and 75.66 on the full MMLU test set, 4.76, 4.21 and 2.46 points below 4-bit AWQ at 5.3 to 6.0 bits per parameter. They generate 113.8, 95.0 and 57.2 tokens per second in our engine. On GSM8K, through the served kernel, they lose 9.63, 4.62 and 3.26 points to FP16. At 4B our file scores 23.6 points above llama.cpp's IQ2_XXS (2.48 bits per parameter). Every number we measured for a table or figure comes from one NVIDIA L40S GPU. We preregistered the main experiments.
Figures & tables
| Qwen3-4B | Qwen3-8B | Qwen3-14B | |
| projections, Tetra int4 | 168 84 | 199 53 | 181 99 |
| int4 projections | , , down 12–23 | , down 10–26 | , , down 10–28 |
| 4-bit embedding tables | 1 (tied head) | 2 | 2 |
| file size (bytes) | 1,418,224,685 | 2,815,098,745 | 5,087,000,541 |
| bits per parameter, whole model | 2.7320 | 2.6953 | 2.7305 |
| FP16 | AWQ w4g128 | Planes14 | Tetra | no weights | |
| matrices timed | 252 | 252 | 252 | 216 | 252 |
| median ms | 10.973 | 3.261 | 4.997 | 3.424 | 2.289 |
| GB read per pass | 7.27 | 1.90 | 2.18 | 0.95 | 0.07 |
| b/weight in VRAM | 16.000 | 4.179 | 4.804 | 2.148 | 0.159 |
| GB/s, fastest round | 662 | 583 | 437 | 278 | 32 |
| format and engine | b/param | MMLU | decode tok/s | weights GB | |
|---|---|---|---|---|---|
| 4B | FP16, vLLM | 16.00 | 70.14 | 83.1 | 8.04 |
| f16 dense path, our engine a | 16.00 | 43.0 | 8.04 | ||
| AWQ w4g128, vLLM | 5.30 | 68.14 | 200.5 | 2.67 | |
| IQ2_XXS, llama.cpp | 2.48 | 39.78 | 312.9 | 1.25 | |
| Tetra , our engine | 2.73 | 63.37 | 113.8 | 1.38 | |
| LLVQ, cited b | 2 b/w c | 62.8 |
| FP16 | AWQ | Tetra | FP16 minus Tetra | AWQ minus Tetra | |
|---|---|---|---|---|---|
| 4B | 92.12 | 89.01 | 82.49 | 9.63 | 6.52 |
| 8B | 93.25 | 92.95 | 88.63 | 4.62 | 4.32 |
| 14B | 95.30 | 95.38 | 92.04 | 3.26 | 3.34 |
| Tetra | against | weights GB | MMLU | GSM8K |
| 8B | AWQ 4B | 2.76 vs 2.67 | ||
| 14B | AWQ 8B | 5.04 vs 6.10 | ||
| 8B | FP16 4B | 2.76 vs 8.04 | ||
| 14B | FP16 8B | 5.04 vs 16.38 | ||
| 14B | FP16 4B | 5.04 vs 8.04 | ||
| AWQ 8B against AWQ 4B | 6.10 vs 2.67 | |||
| step | b/param | MMLU | gain | 95 % interval, McNemar | |
|---|---|---|---|---|---|
| 4B | Tetra int4 v_proj | 2.7475 | 57.95 | ||
| trained row scales | 2.7475 | 61.11 | , | ||
| int4 , down; 4-bit tables | 2.7320 | 63.37 | , | ||
| 8B | Tetra int4 v_proj | 3.0683 | 64.87 | ||
| trained row scales | 3.0683 | 68.16 | , | ||
| int4 down; 4-bit tables | 2.6953 | 69.58 | , |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| arm | matrices | median ms | GB read | b/weight | GB/s |
|---|---|---|---|---|---|
| FP16, our control | 252 | 10.973 | 7.27 | 16.000 | 662 |
| FP16, cuBLAS | 252 | 10.828 | 7.27 | 16.000 | 672 |
| Slot32 (ours, earlier) | 252 | 5.732 | 2.50 | 5.510 | 437 |
| Planes14 (ours, earlier) | 252 | 4.997 | 2.18 | 4.804 | 437 |
| Planes12x (ours, earlier) | 252 | 5.366 | 1.97 | 4.342 | 368 |
| Golay70 (ours, earlier) | 252 | 8.064 | 1.63 | 3.589 | 202 |
| quantity | predicted | measured | inside? |
|---|---|---|---|
| Tetra time at tile 64, tile sweep | 3.77 ms | 3.423 ms | no, 0.08 ms under |
| gain from tile 128 to 64 | no, 1.1 over | ||
| row-scale training, 4B | yes | ||
| row-scale training, 8B | yes | ||
| row-scale training, 14B | no, 0.03 under | ||
| sealed composition, 4B | no, 0.06 over |