Group-wise post-training quantizers for large language models round weights onto a grid that is not refit to the resulting integer codes. We show that this leaves accuracy on the table: the best grid depends on the codes, input correlations couple the errors of different groups, and useful code changes often involve many codes at once. We propose JARQ , a plug-in refinement that starts from any group-wise quantizer and alternates a joint least-squares fit of all group scales with bounded Babai proposals that move many codes of a group together on the current grid. The problem is a bilinear box-constrained mixed-integer least-squares problem; the solver is backpropagation-free, does not increase the layer-wise objective under exact scale solves, and keeps the host's bit width, groups, zero points, and inference cost. Across Llama-2, Llama-3, and Qwen models with RTN, GPTQ, OmniQuant, and AWQ hosts, JARQ lowers perplexity in 90 of 96 comparisons, cuts three-bit RTN perplexity by up to 36%, raises mean multiple-choice accuracy in 23 of 24 configurations, and improves QEP, QuaRot, and OJBKQ outputs, at under a minute per 7B block.
Figures & tables
Figure 1: JARQ on one group of two weights with correlated inputs. Ellipses are error level sets around w ( ⋆ ); dots are grid points. (a) The host codes are optimal on the host grid. (b) Refitting the scale moves the grid. (c) On the new grid, one Babai step moves both codes (Proposition 3.1 ). (d) The objective stays at or below the host level.
J(S,Q)=∥Y−RW∥F2.
(11)
Algorithm 1 JARQ layer refinement
Llama-2-7B
Llama-2-13B
Llama-3-8B
Qwen2.5-7B
Qwen3-4B
Qwen3-8B
Method
WT2
C4
WT2
C4
WT2
C4
WT2
C4
WT2
C4
WT2
C4
W3A16
FP16
5.472
6.973
4.884
6.468
6.136
8.881
6.848
10.442
13.638
16.626
9.715
13.290
RTN
6.663
8.405
5.521
7.179
12.048
16.444
11.954
16.054
22.447
27.013
13.456
17.582
+ Ours
6.204 − 6.88
8.217 − 2.24
5.270 − 4.54
7.158 − 0.29
8.219 − 31.8
12.667 − 23.0
7.625 − 36.2
11.916 − 25.8
16.194 − 27.9
21.625 − 19.9
10.712 − 20.4
15.103 − 14.1
GPTQ
6.083
7.990
5.248
7.034
7.618
13.242
7.665
11.677
15.153
18.623
10.749
14.720
Table 1: Perplexity before (Base) and after ( + Ours) refinement with K=3 (lower is better). Small numbers give the relative change of + Ours over Base in percent (green: lower, red: higher).
Model
Host
W/A
ARC-E
ARC-C
MMLU
Wino.
BoolQ
Avg.
Llama-3-8B
FP16
W16A16
77.53
54.10
65.45
74.03
82.26
70.67
RTN
W3A16
61.07/ 67.85
40.53/ 44.45
47.17/ 53.11
67.48/ 71.03
69.11/ 75.60
57.07/ 62.41 + 5.34
W4A16
77.31/ 77.78
51.62 / 51.11
62.63/ 63.57
72.38/ 72.38
79.69/ 81.07
68.73/ 69.18 + 0.45
GPTQ
W3A16
63.59/ 69.32
39.16/ 43.00
52.81/ 57.47
71.35 / 71.19
69.88/ 70.34
59.36/ 62.26 + 2.90
W4A16
77.48/ 78.75
53.67 / 53.33
63.71 / 63.45
74.11/ 74.43
81.44/ 82.51
70.08/ 70.49 + 0.41
OmniQuant
W3A16
63.80/ 68.52
40.96/ 43.00
53.45 / 52.55
68.75/ 70.72
75.08 / 71.41
60.41/ 61.24 + 0.83
Table 2: Multiple-choice accuracy (%; higher is better). Entries are Base/ + Ours; Avg. is the five-task mean, followed by the change of + Ours in points (green: higher, red: lower). Bold marks the higher value; ties are not bolded.
Figure 2: Generative accuracy at W4A16 on GSM8K (a–c) and MATH-500 (d–f) for three hosts: Base (gray), + Ours (blue), and FP16 (dashed). Labels give the change in points; titles give the mean over hosts. Axes are zoomed per panel. W3A16 GSM8K scores are in Table 10 .
Model
W/A
Base
+ Ours
Qwen2.5-7B
W3A16
7.67
7.58
W4A16
7.02
7.00
Qwen3-8B
W3A16
11.00
10.52
W4A16
9.94
9.88
Table 3: Compatibility results for (a) QEP ( Arai and Ichikawa, 2025 ) , (b) QuaRot ( Ashkboos et al., 2024 ) on Llama-2-7B, and (c) OJBKQ ( Wang et al., 2026 ) (WT2 perplexity, with C4 also in (b); lower is better). Settings are in Appendix A.4 .
Measurement
RTN
GPTQ
Indep. scale fits: gap (%)
20.98
25.29
3-pass coord. descent: gap (%)
1.31
0.01
3-pass coord. descent: time ratio
1.29
1.32
Babai gain after 1-code search (%)
40.14
13.53
Median codes changed
30
6
Joint over Fixed, held-out (%)
12.32
5.47
Table 4: Ablations. (a) Medians over 16 layers at W3A16 (Appendix C ). (b) Full-model WT2/C4 perplexity on Llama-3-8B (L3) and Qwen3-4B (Q3); bold: lowest (Appendix C.5 ).
Figure 3: Per-block error change (%) from refinement with the AWQ host at W3A16 on held-out WT2, for the seven linear modules of Llama-2-7B (top) and Qwen3-8B (bottom); below zero is better. Local error uses fixed full-precision inputs, propagated error each model’s own inputs. Local error falls in all 476 block–module pairs, propagated error in 468. Other settings: Appendix D .
Figure 4: (a, b) WT2 perplexity of RTN at W3A16 versus group size. (c) Perplexity change from refinement on the 32B models (lower is better).
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
K
Refinement time (s)
Total time (s)
Overhead
GPU memory (GiB)
CPU memory (GiB)
0
0.00±0.00
128.37±0.83
—
2.07
6.63
3
48.43±0.29
176.80±1.12
+37.73%
2.48
5.95
5
57.61±0.41
185.98±1.20
+44.88%
2.53
5.86
7
67.15±0.74
195.52±1.50
+52.31%
2.58
5.79
10
80.65±0.55
209.02±1.37
+62.83%
2.64
5.69
Appendix
Table 5: Time and memory for the first Qwen3-8B block at W3A16 with the OmniQuant host. Times are mean ± sample standard deviation over three repetitions; total time includes the host. Memory is the time-weighted mean over both stages.
Figure 11
One block
Two blocks
Model
Host
Host
+ Ours
Δ
Host
+ Ours
Δ
Llama-2-7B
RTN
4.9
61.5
56.6
7.8
119.8
112.0
GPTQ
33.0
86.6
53.6
64.0
170.1
106.1
AWQ
28.4
81.5
53.1
49.3
154.0
104.7
OmniQuant
69.7
126.6
56.9
135.6
244.9
109.3
Qwen3-4B
RTN
2.8
40.3
37.5
4.8
79.4
74.6
Appendix
Table 6: Quantization time (s) of each host alone and with refinement ( + Ours) at W3A16 over the first block and the first two blocks; Δ is the time added by refinement. Medians of the completed repetitions (one or two per entry).
Figure 7: Perplexity gap to FP16, PPL−PPLFP16 , on the 32B models (lower is better). Gray is Base ( K=0 ); blue is + Ours ( K=3 ). Axes start at zero, with a tighter range for C4 in (a). Labels give PPL+Ours−PPLBase . Absolute values are in Tables 7 and 8 .
WT2
C4
Host
W/A
Base
+ Ours
Base
+ Ours
FP16
W16A16
5.0176
8.9503
OmniQuant
W3A16
5.7804
5.6818
9.5809
9.5765
W4A16
5.1951
5.2085
9.0920
9.0431
AWQ
W3A16
5.8960
5.7181
9.5722
9.5719
W4A16
5.2222
5.1824
9.0894
9.0833
Appendix
Table 7: Qwen2.5-32B perplexity (lower is better). Base uses K=0 and + Ours uses K=3 . Bold marks the better value in each Base/ + Ours pair (compared before rounding).
WT2
C4
W/A
Base
+ Ours
Base
+ Ours
W3A16
8.2562
8.0641
11.7171
11.5977
W4A16
7.7106
7.7025
10.9813
10.9500
Appendix
Table 8: Qwen3-32B perplexity with the OmniQuant host (lower is better). Base uses K=0 and + Ours uses K=3 . Bold marks the better value in each Base/ + Ours pair (compared before rounding).
GSM8K
MATH-500
Calibration
Base
+ Ours
Base
+ Ours
WT2
59.67
66.19
29.4
28.2
C4
61.49
67.70
29.8
30.2
Pile
69.98
72.02
42.4
45.6
Appendix
Table 9: GSM8K and MATH-500 accuracy (%; higher is better) of Qwen3-4B with GPTQ at W3A16 under three calibration corpora. Bold marks the better value in each Base/ + Ours pair.
Model
Host
W/A
Base
+ Ours
Llama-3-8B
FP16
W16A16
47.92
GPTQ
W3A16
—
—
W4A16
42.53
42.61
OmniQuant
W3A16
20.17
22.29
W4A16
42.68
43.67
AWQ
W3A16
21.30
21.30
Appendix
Table 10: GSM8K accuracy (%; higher is better). Gray marks FP16 references and blue shading marks + Ours. Bold marks the better value in each Base/ + Ours pair; ties are not bolded. — denotes a configuration not reported.
Model
Host
W/A
Host
Codes only
Scales only
Full JARQ
Llama-3-8B
RTN
W3A16
12.048/16.444
8.689/14.171
8.580/13.311
8.219 / 12.667
W4A16
6.727/9.636
6.586/9.605
6.490/9.583
6.481 / 9.532
GPTQ
W3A16
7.618 /13.242
7.714/12.318
7.784/12.628
7.649/ 12.011
W4A16
6.423/ 9.386
6.419/9.416
6.417/9.411
6.416 /9.408
Qwen3-4B
RTN
W3A16
22.447/27.013
21.205/27.078
17.968/22.316
16.194 / 21.625
W4A16
15.968/18.291
14.629/ 17.924
14.598/17.965
14.555 /17.930
Appendix
Table 11: Full-model component ablation (WT2/C4 perplexity; lower is better). Bold marks the lowest value per setting and corpus.
Host
Local NMSE
Linear output SSE
Linear combined
Final block NMSE
Final block combined
RTN
1,818/1,904
1,888/1,904
1,888/1,904
8/8
8/8
GPTQ
1,739/1,904
1,447/1,904
1,443/1,904
7/8
6/8
AWQ
1,721/1,904
1,648/1,904
1,644/1,904
6/8
8/8
OmniQuant
1,794/1,904
1,456/1,904
1,456/1,904
2/8
3/8
Appendix
Table 12: Pairs in which refinement lowers the error. The first three columns count all modules and blocks of both models, bit widths, and corpora; the last two count the final block only.
Figure 8: Local linear NMSE ( 14 ) for Llama-2-7B. Rows are the seven projections; the left and right halves are WT2 and C4. At the final block, refinement lowers it in 104/112 module pairs (RTN 26/28, GPTQ 24/28, AWQ 26/28, OmniQuant 28/28).
Figure 10: Propagated output SSE ( 15 ) for Llama-2-7B, laid out as in Figure 9 . At the final block, refinement lowers it in 75/112 module pairs (RTN 27/28, GPTQ 16/28, AWQ 21/28, OmniQuant 11/28).
Figure 12: Per-module combined error ( 16 ) for Llama-2-7B, laid out as in Figure 9 . At the final block, refinement lowers it in 75/112 module pairs (RTN 27/28, GPTQ 16/28, AWQ 21/28, OmniQuant 11/28).
Model
W/A
Data
g=64
g=128
g=256
Llama-2-7B
W3A16
WT2
6.4009/ 6.0186
6.6630/ 6.2043
7.1019/ 6.2951
C4
8.1095/ 7.8924
8.4050/ 8.2166
8.9828/ 8.4944
W4A16
WT2
5.6780/ 5.5938
5.7261/ 5.6256
5.7495/ 5.6474
C4
7.2114/ 7.1347
7.2458/ 7.1857
7.2983/ 7.2414
Qwen3-8B
W3A16
WT2
12.3856/ 10.4292
13.4564/ 10.7123
15.7918/ 11.1354
C4
16.0018/ 14.5767
17.5825/ 15.1033
20.3924/ 15.8302
Appendix
Table 13: Group-size ablation with the RTN host, ν=0.6 , and K=3 . Entries are Base/ + Ours perplexity (lower is better). Bold marks the better value in each pair (compared before rounding).
Model
W/A
g=128 (min)
g=64 (min)
Increase
Llama-2-7B
W3A16
32.13
87.96
+173.78%
W4A16
32.06
87.04
+171.50%
Qwen3-8B
W3A16
38.54
95.36
+147.43%
W4A16
39.24
96.00
+144.64%
Appendix
Table 14: Full-model quantization and refinement time for the RTN host with ν=0.6 and K=3 . Increase is 100(T64/T128−1) .
Llama-2-7B
Qwen3-8B
ν
W3A16
W4A16
W3A16
W4A16
Base ( K=0 )
6.6630/8.4050
5.7261/7.2458
13.4564/17.5825
10.1369/13.6431
0.0
6.2254/8.2287
5.6203 /7.1873
10.5207 / 15.0312
9.9718/13.5872
0.2
6.1879/8.2263
5.6242/ 7.1837
10.6558/15.1019
9.9430/13.5646
0.4
6.1944/8.2391
5.6295/7.1875
10.7394/15.1621
9.9146 / 13.5575
0.6
6.2043/8.2166
5.6256/7.1857
10.7123/15.1033
9.9980/13.6271
Appendix
Table 15: Regularization-strength ablation with the RTN host, g=128 , and K=3 . Entries are WT2/C4 perplexity (lower is better). Bold marks the best of the six strengths for each model, bit width, and corpus (compared before rounding).
Changed coordinate
New codes qT
Refitted error ϕ(q)
First
(0,2)
4/101
(2,2)
1
(3,2)
25/104
Second
(1,0)
9/100
(1,1)
1
(1,3)
9/409
Appendix
Table 16: All single-code neighbors of (1,2)T in ( 28 ), with an exact scale refit for each candidate. The current state’s error is 1/104 .
Figure 14: Error after refitting the scale for each code pair in ( 28 ) (darker is lower). Black: current codes (1,2) ; dashed: single-code moves; green: the Babai step to (2,3) .
Post-training quantization (PTQ) is widely used to reduce the memory and computational cost of large language models. Existing PTQ methods typically obtain an initial quantized model through heuristic rules or greedy optimization, and once quantization is completed the resulting integer assignments are usually treated as final. This observation motivates a complementary optimization stage within PTQ that keeps quantized weights improvable after an executable quantized model has been produced, while preserving the quantized format. We introduce ReQuant, a backpropagation-free fixed-grid refinement procedure for this stage. Agnostic to the PTQ initializer, ReQuant takes an existing quantized model as a feasible starting point and iteratively revisits its discrete weight assignments on the fixed quantization grid. Accepted updates strictly reduce the mean squared reconstruction error and remain on the original grid. In this way, ReQuant turns the initially fixed PTQ output into an iteratively optimizable discrete solution and serves as a plug-and-play post-processing stage for existing PTQ pipelines. Experiments across diverse model families, bit-widths, and downstream tasks show that ReQuant consistently improves quantized models from heterogeneous PTQ initializers, with especially large gains on simple initializers and lower bit-widths. Notably, ReQuant can refine a simple round-to-nearest initialization across multiple sweeps until it approaches or surpasses GPTAQ under the same quantization format. These results establish ReQuant as a practical complementary stage for further improving existing PTQ pipelines.
Yongge Ma, Guoan Wang, Feiyu Wang +5
School of Computer Science, Peking University · School of Software and Microelectronics, Peking University · School of Physics, Peking University +2
Serving large language models (LLMs) under diverse deployment constraints requires flexible trade-offs between accuracy, memory footprint, and throughput. However, conventional quantization methods typically require a separate checkpoint for each target bit-width. We introduce Recurrent Residual Quantization (RRQ), a post-training quantization (PTQ) framework that represents weights as a low-bit quantized base together with a sequence of quantized residual corrections, enabling multiple effective precisions from a single checkpoint. Starting from a 2-bit model obtained via post-training quantization (PTQ) or round-to-nearest (RTN), RRQ progressively adds lightweight 2-bit residuals generated via RTN to construct 4-, 6-, and 8-bit representations. The method is calibration-free and avoids joint multi-bit optimization. In our Qwen3-8B setup, the full all-RTN 2-/4-/6-/8-bit package is constructed in 1,293 seconds, 3.3 times faster than the measured MatGPTQ construction. Experiments on six recent LLMs show competitive accuracy at 6 and 8 bits, with model-dependent behavior at 4 bits. The code will be made publicly available upon publication.
Low-bit quantization shrinks language models but treats precision as a single global hyper-parameter: every weight uses the same bit-width. We introduce Variable Bit-width Quantization (VBQ), a training-time method in which each contiguous group of 64 weights learns its own resolution from {1,2,4,8} bits via a Gumbel-Softmax relaxation, trained jointly by an alternating optimization that gives the precision logits a clean, task-aligned signal. VBQ discovers a consistent, strongly heterogeneous allocation within individual projection types, not merely across layers, impossible to express with per-layer methods: 69% of groups collapse to 1 bit, the LM head averages 1.09 bits, while the first MLP block keeps ~2.5 bits. This pattern is stable enough to freeze into a fixed recipe and reuse without further search. The recipe yields a "bigger-but-smaller" regime: a 131M model at 1.82 mean bits reaches perplexity 4.2 on TinyStories, beating a 55M FP16 model (PPL 4.4) at 3.8x less storage, and lets a 1.46B model on FineWeb-Edu match a 593M FP16 control at ~3.7x less storage with 2.5x more parameters. As quality-per-byte, VBQ is 3.9-8.4x more efficient than FP16. The recipe maps directly to packed low-bit storage, so it also accelerates inference: with custom fused dequantize-and-multiply kernels, memory-bandwidth-bound autoregressive decode is faster at equal output, and the speedup grows with scale (parity at 131M, 1.9x at 1.0B, 4.7x at 9B on Apple silicon). A distributional analysis (KL divergence and argmax-flip rate) reveals a striking mechanism: deeper layers progressively self-heal the quantization error injected by early layers. The win is a from-scratch, train-time phenomenon; scaling the search economically beyond 1.5B parameters remains open. VBQ reframes precision as a learnable, non-uniform resource and shows that spending a fixed bit budget unevenly beats spending it uniformly.