We introduce ApexQuant, a calibration-free quantization method that recursively re-quantizes the residual error, serving as a refinement layer on top of existing quantizers. We establish that a fresh random rotation returns each residual to the uniform distribution on the hypersphere, which characterizes the rate of progressive error decay across successive passes. This result lets us determine, before any weight is read, how many passes a layer needs for a target weight-space error. Every prefix is itself a valid lower-rate model, so one artifact serves several precisions. We instantiate ApexQuant with three interchangeable stages, scalar, E8 and trellis, and validate it on four open-weight LLMs and on Earth-observation and medical domains where in-distribution data is often unattainable as imagery arrives under restrictive licences or due to patient material under privacy constraints. Progressive re-isotropization comes within a few percent of full precision at four bits and gives the best two-bit arm we measure, in a completely data-free setting.
Figures & tables
Figure 1: ApexQuant: the progressive quantizer end to end. (A) The three arrows into the index store are the error each pass contributes, 1:ρ:ρ2 , so later passes cost the same bits and correct geometrically less. (B) Dashed: the prediction ρP , a quadrature over fd as established in proposition 2 , with b=2 bits in blue, and b=4 bits in orange. Markers: difference measured on eight Clay v1.5 layers, agreeing to at least three decimals at every pass. Red: b = 2 bits with the rotation held fixed , the counterfactual to Proposition 1 . (C) A prefix of the passes is bit-for-bit a valid lower-rate model, so one artifact serves every tier from one weight stream and a mixed batch reads each pass once. Storing one artifact per tier instead shares nothing, which is the six blocks against ten: measured on a Mistral block stack at tiers (3,3,2,2) , nesting stores 6.09 against 10.16 bits per weight and decodes 1.56× faster ( section 5.2 , appendix H ).
Method
Data used
Quantizer
Bit-widths
GPTQ, AWQ, OmniQuant
calibration set
scalar
one
AQLM
calibration set
vector
one
QuaRot
calibration set
scalar
one
QuIP#, QTIP
calibration set (Hessian)
vector
one
HQQ
none
scalar
one
SINQ
none
scalar
one
Table 1: The methods closest to ours, on the three axes of this section. The data column is what a method must read besides the weights, and the last is how many bit-widths one stored model serves. ApexQuant reads none, with the single optional exception of appendix I .
Figure 2: The rate loss of table 2 , read off the rate–distortion bound. Each stage sits above the Gaussian bound D(R)=2−2R by a horizontal distance measured in bits, and strengthening the stage, from rounding one coordinate at a time to choosing 256 together, closes that distance from 0.45 to 0.10 bits. That distance is proposition 3 ’s Δ1 , and P passes pay exactly P of them: successive refinement is costless for a Gaussian source, but only for a stage that achieves D(R) exactly. Measured rather than argued in appendix D .
Stage
Dim.
ρ2 (ours)
ρ2 (publ.)
Rate loss (bits/pass)
Encoding
Scalar Lloyd–Max
1
0.1171
0.118
0.45
nearest of 2b values
E8 lattice
8
0.0906
0.089
0.27
nearest lattice point
Trellis ( L=14 )
256
0.0720
0.069
0.10
Viterbi search
Rate–distortion bound
∞
—
0.0625
0
—
Table 2: The three stages evaluated at two bits per weight. ρ2 is the single-pass error of equation 3 , the published column reports the values given by QTIP and QuIP# for a Gaussian source; and rate loss denotes the additional bits per pass relative to the final row, computed as 21log2(ρ2/ρRD2) from our measured column. The trellis row uses L=14 , the widest our kernel accelerates ( appendix C ).
4 -bit class
2 -bit class
Method
b/w
TL
Phi
Mist.
Qwen
b/w
TL
Phi
Mist.
Qwen
RTN-absmax
4.03
4.419
0.130
0.247
0.148
2.03
8.056
6.742
8.939
14.532
QuaRot-RTN
4.03
0.116
0.059
0.054
0.072
2.03
8.144
2.552
8.543
4.217
SINQ, g=64
≤4.53
0.035
0.028
0.019
0.039
≤2.53
4.537
1.431
2.107
6.152
SINQ, g=512
≤4.09
0.069
0.050
0.040
0.096
≤2.09
10.230
4.476
8.887
9.985
HQQ, g=64
4.50
0.046
0.033
0.025
0.044
2.50
8.115
1.912
6.361
4.091
Table 3: Language models, KL(pfp16∥pquant) in nats per token on WikiText-2 raw at sequence length 2048 , lower is better. Bits per weight is measured storage including side information, which puts the group-size- 64 arms of HQQ and SINQ at 4.50 and 4.53 against our 4.06 ; both also appear at g=512 , matched to ours. SINQ’s overhead is shape-dependent, so its entries are bounds. In the 4 -bit class the scalar stage is one pass at four bits and the vector stages two passes at two, the only 4 -bit form here keeping a valid 2 -bit prefix. Every variant here is data-free; and appendix D includes the extended results with the calibrated counterpart A-SINQ as well as perplexity, tail percentiles, the lp codebook and a group-size sweep. Model abbreviations: TL = TinyLlama-1.1B, Phi=Phi-1.5, Mist.=Mistral-7B, and Qwen = Qwen2.5-7B.
task metric ( ↑ )
KL ( ↓ )
Model
Task
full
P=1
P=2
P=3
P=1
P=2
P=3
Clay v1.5
EuroSAT
96.30
96.00
96.15
96.35
0.224
0.010
0.002
UNI
PatchCamelyon
87.81
85.34
86.68
87.56
0.641
0.033
0.003
Table 4: Results beyond language modeling, evaluated at the three precision tiers ( 2.06 , 4.13 and 6.19 bits). KL divergence is against the same model at full precision. The Clay probe carries a standard error of 0.43 points, so only KL divergence separates its tiers. The KL protocol, the sampling procedure and each model’s quantized fraction are available in appendix E .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Why the rotation has to be fresh? A stage leaves a residual that is structured in the frame it was just quantized in, so the marginal the codebook was designed for is destroyed by the very act of using it. Drawing a fresh rotation restores that marginal exactly ( proposition 1 ), which is what lets the same codebook be optimal at pass two as at pass one, so nothing is refitted ( corollary 1 ) and the error is ρP ( proposition 2 ). Reusing the previous rotation does not: the residual is already adapted to it, the coordinates it produces are not fd , and the geometric decay stalls, its measured per-pass ratios climbing 0.418 , 0.597 , 0.734 towards one ( fig. 1 B). Freshness, not orthogonality, is the active ingredient. The coordinate laws are drawn schematically.
Figure 4: The three data-free stages as one dataflow. Every lane takes the same input, the rotated residual normalized onto the sphere, and every lane emits the same thing, b bits per weight plus one fp16 scale for the group, so a stage can be swapped for another without changing the format, the rotation or anything downstream. The lanes differ in exactly one place, the middle box: how many coordinates are quantized jointly . That is the only axis along which they differ and it is what moves ρ2 , from 0.1171 when coordinates are rounded one at a time to 0.0720 when 256 are chosen together by a Viterbi search. The constants are those of table 2 and are computed from the marginal before any weight is read; appendix C reports what the trellis lane costs to run.
L
states
reference
kernel
speedup
ρ2
8
256
16.30
532.06
32.7×
0.090871
10
1,024
16.22
304.78
18.8×
0.080997
12
4,096
13.38
587.35
43.9×
0.075239
14
16,384
2.90
197.61
68.1×
0.071999
16
65,536
0.72
1.83
2.5×
0.070240
Appendix
Table 5: Viterbi encoder throughput in millions of weights per second, d=512 , b=2 , one RTX 4090. At L=16 the path-cost array exceeds the register budget, so the reference implementation runs and the speedup collapses. Bold marks the shipped configuration rather than a column best: L=16 reaches a lower ρ2 but cannot use the register-resident kernel. L=14 is the width table 2 reports and the encoder ships; the tiered artifacts of table 9 predate the kernel and use L=12 , which is why that table carries both. ρ2 is quoted to six places here and rounded to four in table 2 .
TinyLlama-1.1B
Phi-1.5
Method
bits/w
PPL
KL
KL 99.9
PPL
KL
KL 99.9
4 -bit class
QuaRot-RTN
4.03
8.75
0.116
2.51
22.75
0.059
0.66
SINQ, g=64
≤4.53
8.09
0.035
0.79
22.30
0.028
0.37
SINQ, g=512
4.09
8.41
0.069
1.60
22.46
0.050
0.60
HQQ, g=64
4.50
8.18
0.046
1.47
22.35
0.033
0.41
Appendix
Table 6: Extended results for the two models near 1 B, adding perplexity and the 99.9 th-percentile KL to table 3 ’s mean KL. Bits per weight is measured per arm; SINQ’s bit-width is an upper bound because its dual-scale overhead depends on the layer shape.
Mistral-7B-v0.1
Qwen2.5-7B
Method
bits/w
PPL
KL
KL 99.9
PPL
KL
KL 99.9
4 -bit class
QuaRot-RTN
4.03
5.65
0.054
1.99
7.36
0.072
1.43
SINQ, g=64
4.51
5.43
0.019
0.53
7.16
0.039
0.89
SINQ, g=512
4.07
5.52
0.040
1.23
7.51
0.096
2.26
HQQ, g=64
4.50
5.46
0.025
0.79
7.19
0.044
1.04
Appendix
Table 7: The same results at 7 B, scored by the protocol described below.
Method
TinyLlama
Phi-1.5
Mistral-7B
Qwen2.5-7B
Reference
Full precision
55.7
68.5
76.7
75.2
4 -bit class
HQQ, g=64
54.9
68.1
75.9
75.0
ApexQuant, scalar
54.9
68.4
75.1
74.1
ApexQuant, trellis
55.2
68.1
76.0
74.4
Appendix
Table 8: Classification accuracy (mean of ARC-Easy, WinoGrande, HellaSwag), scored by length-normalized log-likelihood; chance is 33.3 . Binomial standard error per cell is 0.58 – 0.69 points. At four bits all gaps are within noise of full precision. Bold at two bits marks the best data-free arm. The last row reports the trellis arm’s above-chance margin as a percentage of full precision’s.
Stage
P=1 ( 2.06 bit)
P=2 ( 4.125 bit)
P=3 ( 6.19 bit)
Scalar Lloyd–Max
6.472(0.131)
0.4157(0.0324)
0.1154(0.0194)
E8 P lattice
3.648(0.120)
0.3850(0.0533)
0.0502(0.0070)
Trellis, 1MAD ( L=12 )
3.996(0.168)
0.2626(0.0300)
0.0440(0.0075)
Trellis, 1MAD ( L=14 )
3.068(0.134)
0.2312(0.0279)
0.0581(0.0151)
Appendix
Table 9: Gemma-3-12B, KL(pbf16∥pquant) in nats per token on held-out WikiText-2, lower is better, reported as mean (standard error) over 32 windows of 2048 tokens. Every arm is data-free and differs only in the stage of table 2 ; g=256 . The unquantized reference is scored once with its tower offloaded and its log-probabilities cached; all four quantized rows are then evaluated against that single cache, ensuring an identical reference across rows ( experiments/kl_two_phase.py ). The P=1 and P=2 columns are prefixes of the same P=3 artifact, verified to be bit-identical to encoding at that depth directly ( experiments/prefix_identity_gemma.py ). The first three rows read the published artifacts; the L=14 row is quantized here with all other settings matched, because table 2 reports L=14 while the published artifact uses L=12 . No cell is bolded, as the stages are separated at 2 bits but not at 6 ( experiments/kl_separability.py ).
LLaMA2-7B
LLaMA3-8B
Method
Data
Elast.
≈4 bit
≈2 bit
≈4 bit
≈2 bit
Full precision
—
—
5.503
5.503
6.198
6.198
ApexQuant trellis (ours)
none
yes
5.613
71.80
6.550
52.51
ApexQuant E8 P (ours)
none
yes
5.723
218.7
6.735
165.7
ApexQuant scalar (ours)
none
yes
5.750
2046
6.817
3134
MoBiQuant
128 seq., 20 ep.
yes
5.82
10.91
7.31
58.12
Appendix
Table 10: WikiText-2 perplexity at sequence length 2048 ; lower is better. Our rows use the harness of table 3 ; all others are from Wang et al. (2026) (Tables 1 and 2). Bit-widths for published elastic rows are averages over routed tokens excluding router parameters; ours are stored bits per weight ( 4.03 scalar, 4.06 vector, 2.03 throughout). Bold marks the lowest quantized entry per column, excluding full precision. ∗ Reproduced using the released quantizer and evaluation script ( experiments/anyprecision_llama2_7b.txt ); the 6.05 and 2e3 in Wang et al. (2026) do not match the 5.624 and 35.21 we obtain.
Figure 5: What “elastic” means in bytes. The passes of one weight row are stored back to back, so the first P′ of them are a contiguous prefix: reading it yields exactly the artifact a P′ -pass run would have produced, checked by byte comparison on three published tiered models, and raising a deployed model’s rate appends indices rather than rebuilding it. The reads shown yield 2.03 , 4.06 and 6.09 bits per weight at P′=1,2,3 , and the one stream that serves them all costs its deepest tier alone. Fixed-width elasticity has to keep one independent copy per tier: on the Mistral 32 -block stack at tiers (3,3,2,2) that is a 6.09 and a 4.06 copy, 10.16 bits per weight against nesting’s 6.09 , which is where the 1.67× storage factor comes from. That factor holds unconditionally, while the 1.56× latency factor needs a batch deep enough for the extra passes to stop being free ( appendix H ).
Scenario
fast (ms)
max (ms)
A
Naive sequential, four batch-1 decodes
28.11
28.19
B
Premium-only batch, M=2 , three passes
19.38
13.46
C
Homogeneous premium batch, M=4
19.20
14.02
D
Tiered batch, M=4 , mixed (ours)
19.47
13.95
E
Separate artifacts, three-pass plus a distinct two-pass
31.81
21.76
Appendix
Table 11: Decode latency for one batch, Mistral 32 -block stack, tiers (3,3,2,2) , graph-captured under two autotune scopes. M is the batch size. Row D is ours.
Figure 6: The calibration pass, the single exception to reading only weights. Layers differ in how much a given quantization error hurts the model, so the allocation measures that directly rather than inferring it: each candidate layer is quantized alone with the rest left at full precision and the displacement of the model’s output is recorded, 97 forward passes on Clay rather than one in total. The ranking is then spent as pass counts, and nothing is refitted. What it buys is a strict Pareto improvement in embedding error, 0.0406 at 4.99 bits per weight against 0.0534 at 6.19 for uniform: lower error and less storage. What it does not buy is the downstream task, where the allocated arm sits 0.15 points below uniform on the EuroSAT probe, about a third of one standard error, for 0.8 more stored bits; on language models the sign flips between models. The gain concentrates on a small slice of layers rather than spreading evenly, which is the structure an unweighted weight-space objective cannot see. Which slice we do not claim: this run keeps the patch-embedding linears dense throughout, so it is not the input interface the panel labels ( appendix I ).
ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (RTN) schemes when quantizing weights near the centers of quantization intervals. Starting from a pretrained LLM, ReRound trains a conditional diffusion model to produce continuous reconstructions of low-bit weights for the LLM. These reconstructed weights act as a guidance signal to disambiguate the rounding direction of weights located close to interval midpoints. To integrate this reconstruction-guided rounding with conventional RTN, ReRound introduces a tolerance metric measuring how far the quantized weight (not the final quantized integer) is away from the midpoint: quantized weights within a tolerance region around midpoints are quantized using diffusion-based reconstructions, whereas weights closer to quantization boundaries are quantized with RTN. By sweeping the tolerance parameter, ReRound generates multiple candidate quantized integer weight matrices and selects the de-quantized weight matrix candidate whose leading singular values most closely match those of the original full-precision weights. This selected candidate determines the tolerance parameter ReRound uses. ReRound is particularly effective for smaller LLMs. Across a range of such models, it consistently outperforms standard RTN for 3-bit and 4-bit weight quantization. ReRound achieves superior accuracy compared to an extensive set of calibration-free methods, remains competitive with calibration-dependent approaches, and operates entirely offline, introducing no additional overhead during low-bit inference. The ReRound strategy represents a new approach for low-bit quantization. The method applies to AI models beyond LLMs. This paper focuses on its applications to small LLMs.
Aggressive weight quantization to 2-bit precision offers substantial throughput and memory gains for large language model (LLM) inference, but typically incurs severe accuracy degradation. These gains are particularly relevant for edge and on-device deployment, where memory capacity and bandwidth are primary constraints. In this work, we extend Recover-LoRA -- a lightweight, data-free accuracy recovery method originally developed for general model weight corruption -- to the setting of ultra-low-bit quantization. We propose a selective mixed-precision strategy in which only gate and up projection layers of the MLP are quantized to 2-bit (W2), while all other linear layers remain at higher precision, yielding a mixed-precision GateUp configuration. We demonstrate via roofline analysis across three model families (4B--20B) and two hardware platforms that a W4/W2-GateUp deployment (4-bit base with 2-bit gate/up) delivers 7.5--23.3% TPS improvement over uniform W4 depending on model and context length, while confining quantization error to a predictable subset of layers. We then apply Recover-LoRA -- training low-rank adapters on the quantized layers via logit distillation with synthetic data -- to recover accuracy lost from 2-bit quantization of the gate and up layers. In a case study on Qwen3-4B, Recover-LoRA achieves 80--95% accuracy recovery on 9 of 12 benchmarks, using only 10k synthetic training samples and no labeled data. We further demonstrate that synthetic data performs comparably to curated labeled data for distillation-based recovery, and that recovery generalizes to out-of-distribution evaluation tasks. Our results present Recover-LoRA as a practical post-quantization accuracy recovery tool for aggressive weight compression in deployment settings.
Serving large language models (LLMs) under diverse deployment constraints requires flexible trade-offs between accuracy, memory footprint, and throughput. However, conventional quantization methods typically require a separate checkpoint for each target bit-width. We introduce Recurrent Residual Quantization (RRQ), a post-training quantization (PTQ) framework that represents weights as a low-bit quantized base together with a sequence of quantized residual corrections, enabling multiple effective precisions from a single checkpoint. Starting from a 2-bit model obtained via post-training quantization (PTQ) or round-to-nearest (RTN), RRQ progressively adds lightweight 2-bit residuals generated via RTN to construct 4-, 6-, and 8-bit representations. The method is calibration-free and avoids joint multi-bit optimization. In our Qwen3-8B setup, the full all-RTN 2-/4-/6-/8-bit package is constructed in 1,293 seconds, 3.3 times faster than the measured MatGPTQ construction. Experiments on six recent LLMs show competitive accuracy at 6 and 8 bits, with model-dependent behavior at 4 bits. The code will be made publicly available upon publication.