Serving a large language model (LLM) across a fleet of deployments requires several weight-precision operating points. Multi-precision formats serve them all from one stream whose prefixes are valid lower-precision codes, instead of storing multiple copies. We present StagQ, a multi-precision weight format whose main stream is a 2-bit group-wise affine base followed by a configurable number of 1-bit refinement planes on a dyadic step schedule. Every supported precision is a readable prefix, decoded by an affine map derived from metadata shared across all precisions, with no per-weight lookup. A sparse side record, filled both before and after the grid is fitted, holds out the few weights the grid serves worst. We report two configurations of the encoder. At two bits the cheaper one leads the strongest multi-precision baseline on Llama-3.1-8B, Phi-4, and OLMo-2-7B by 3.1 to 7.0 MMLU points, at a slightly lower logical rate. At three bits it leads on Llama-3.1-8B, leads on Phi-4 at a higher rate, and ties on OLMo-2-7B. At four bits it ties on all three, at a higher rate. In a batch-one matrix-vector product on an NVIDIA A100 GPU, timed on synthetic weights, our kernel is faster than the two baseline kernels in most shape-precision cases.
Figures & tables
Figure 1: The StagQ main stream. Each added plane halves the spacing, and every precision is again a uniform grid. The weight w=2.15 is encoded as q=2,b3=1,b4=0,b5=1,b6=0 .
Figure 2: Measured position δu of the conditional centroid inside each code cell, at five precisions on a common vertical scale. Blue: Llama-3.1-8B at g=16 ; green: g=32 . Further checkpoints are in Appendix D .
Figure 3: The StagQ encoder and the fields its rate charges. Dashed lines in each stage’s color map the stage to the field it fills. A p M read stops after plane M .
Method
Bit
Rate
MMLU
ARC-C
ARC-E
HS
PIQA
WG
CSR Avg
bfloat16
16
16.00
65.23
53.67
81.06
78.91
81.28
73.72
73.73
MatGPTQ
2
2.25
24.65
24.74
24.45
26.29
51.41
49.17
35.21
Any-Precision
2
2.01
23.82
24.57
34.09
27.90
57.24
50.36
38.83
AnyBCQ
2
2.38
35.33
37.29
62.67
62.76
73.94
57.93
58.92
StagQ-XXS
2
2.35
38.38
38.05
60.94
63.63
73.23
65.27
60.23
StagQ-XS
2
2.64
48.67
41.98
70.12
70.05
75.52
67.32
65.00
Table 1: Accuracy on Llama-3.1-8B, in AnyBCQ’s column order (HS is HellaSwag and WG WinoGrande). Bit is the read precision M and Rate the logical bpw . Bold marks the best quantized entry per column within a precision.
Wiki ↓
MMLU ↑
tok/s ↑
Bit
bf16
MatQ
AP
AB
XXS
bf16
MatQ
AP
AB
XXS
fp16
AP
AB
XXS
Llama-3.1-8B
2
6.24
> 10 6
4674.38
19.03
12.77
65.23
24.65
23.82
35.33
38.38
46.16
73.91
74.15
73.68
3
6.24
11.07
8.58
8.09
7.24
65.23
44.84
55.69
58.13
61.89
46.16
70.10
70.16
70.88
4
6.24
7.48
6.70
6.84
6.57
65.23
59.29
63.55
63.17
64.11
46.16
66.32
66.36
68.31
Phi-4
Table 2: MatGPTQ (MatQ), Any-Precision LLM (AP), AnyBCQ (AB) and StagQ-XXS (XXS). Tok/s is the mean of six timings in one run. OLMo-2-7B timings omit every per-channel multiply, favoring us.
side record
Wiki ↓
MMLU ↑
Bit
( bpw )
2
3
4
5
6
2
3
4
5
6
neither guard
0
28.47
7.60
6.66
6.43
6.38
26.90
58.03
62.62
64.29
64.74
post-grid only
0.0056
23.18
7.47
6.58
6.36
6.31
28.51
58.63
63.11
64.47
64.88
pre-grid only
0.0010
24.77
7.57
6.69
6.48
6.43
25.00
59.74
63.84
64.40
64.78
both
0.0065
20.86
7.44
6.60
6.41
6.36
25.09
60.30
64.26
64.64
64.77
Table 3: Guard ablation, Llama-3.1-8B, g=32 . The side-record column is the rate the guards add. Every row stores the payload bits and 0.25bpw of per-group pairs.
Wiki ↓
Bit
StagQ-XXS
dedicated
Δ MMLU
2
12.77
11.99
+3.62
3
7.24
7.18
−0.03
4
6.57
6.44
+0.72
Table 4: Single-precision control on Llama-3.1-8B. Δ MMLU is the dedicated fit minus StagQ-XXS.
Bit
Any-Precision
AnyBCQ
StagQ
2
3.53×
3.93×
4.13×
3
2.76×
3.01×
3.39×
4
2.34×
2.43×
3.00×
Table 5: Median speedup over cuBLAS, tcuBLAS/t∗ , on synthetic weights. The StagQ kernel here applies no activation scale. Its layout reserves guard space only on the three attention shapes. Appendix F gives the per-shape latencies.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Pre-grid guard
Model
Configuration
Feed-forward g
Attention g
ratio criterion
gain criterion
Llama-3.1-8B
StagQ-XS
16
16
feed-forward, T=4
attention, τ=10−8
StagQ-XXS
32
v 16, q,k,o 32
feed-forward, T=4
attention, τ=10−9
Phi-4
StagQ-XS
16
16
attention, down, T=4
—
StagQ-XXS
32
32 (fused)
down, T=4
attention, τ=10−8
OLMo-2-7B
StagQ-XS
16
16
all, T=8
—
Appendix
Table 6: Encoder settings per model and configuration. Each pre-grid guard column names the matrices under that criterion and its threshold. A dash marks a criterion that no matrix uses. The two criteria are exclusive per matrix. On Phi-4 the fused gate/up projection carries only the post-grid guard.
Figure 4: The measurement of Figure 2 on the three further dense checkpoints (rows: Qwen3.6-27B, OLMo-2-7B, Qwen3-8B) at five precisions (columns), on a common vertical scale. Blue: g=16 ; green: g=32 . R2 of the proportional fit (blue / green) is marked from p3 on.
bpw
Wiki ↓
MMLU ↑
g=32 throughout
2.26
20.86
25.09
+v at g=16
2.26
16.68
29.46
→ StagQ-XXS
2.35
12.77
38.38
g=16 throughout
2.51
10.82
43.71
→ StagQ-XS
2.64
9.68
48.67
Appendix
Table 7: Llama-3.1-8B at two bits, with rates in logical bpw . The +v row changes only the value projection’s group size. The rows marked → change several settings at once. The g=32 row is the both-guards row of Table 3 . AnyBCQ scores 19.03 and 35.33 at 2.38bpw .
cuBLAS
Any-Precision
AnyBCQ
StagQ
Shape
16-bit
2-bit
3-bit
4-bit
2-bit
3-bit
4-bit
2-bit
3-bit
4-bit
8B qkv/o 4096×4096
29.4
11.5
14.3
16.6
12.2
15.3
17.9
13.2
14.7
14.6
8B down 14336×4096
97.8
29.2
38.6
47.1
25.1
32.5
40.2
24.1
29.0
32.6
8B gate/up 4096×14336
97.9
27.3
34.7
41.9
24.9
32.5
40.7
23.7
28.9
32.8
14B qkv/o 5120×5120
50.1
15.5
18.7
22.5
14.3
18.2
22.2
16.3
18.4
20.0
14B down 17920×5120
141.3
40.8
48.4
54.4
32.1
37.7
49.4
30.8
36.4
41.6
Appendix
Table 8: Batch-one GEMV latency ( μ s) on one A100 (40 GB) for the nine shapes (input × output) of AnyBCQ’s kernel table, from Llama-3-8B, Phi-4 and Llama-3-70B in row order. Bold marks the fastest quantized kernel per shape and width.
School of Computer Science and Technology, Soochow University, Suzhou, China · Hithink Research, Hangzhou, China · Electronic Engineering, Tsinghua University, Beijing, China