As large language model inference shifts toward lower precision, post-training quantization (PTQ) becomes increasingly brittle, making quantization-aware training (QAT) essential for preserving model quality. However, QAT has a structural mismatch: gradient updates are applied to latent full-precision weights, while the loss and gradients are computed on lossy reconstructions of those weights. This mismatch can lead to suboptimal training trajectories and a higher loss floor. Second-order PTQ methods address a similar problem by minimizing loss-aware reconstruction error, but applying such expensive reconstruction repeatedly during QAT as the weights evolve is impractical. We introduce QUASAR, a QAT method that brings lightweight, loss-aware reconstruction into the training loop. At each training step, QUASAR reconstructs the latent weights by searching over a small set of clipping ranges and fitting dequantization parameters through saliency-weighted least squares. We use an exponential moving average of squared gradients as the per-parameter saliency signal. Our theoretical analysis shows that optimizing QUASAR's reconstruction objective tightens both the convergence and final-loss bounds of QAT. We evaluate QUASAR across four model families and across INT4, INT3, INT2, and NVFP4 quantization formats. QUASAR consistently achieves lower training and evaluation loss than competitive QAT methods and outperforms QAT and PTQ baselines on downstream benchmarks. At INT2, QUASAR improves average accuracy over the best QAT baseline by 13.3 points with quantization-aware distillation and by 10.9 points with QAT on mathematical reasoning data. Notably, after distillation on only about 600M tokens, QUASAR's INT4 Gemma-4 E4B checkpoint outperforms the corresponding QAT checkpoint released by Google, with 66% lower KL divergence and 1.8 points higher average accuracy.
Figures & tables
Figure 1: The structural mismatch of gradients in standard QAT versus QUASAR. Under the STE, both methods evaluate the gradient at the reconstructed weights r and apply it to the latent weights w . QUASAR minimizes the loss-aware reconstruction error, making the gradient at r a more faithful proxy for the gradient at w .
Figure 2: The training loops for standard QAT and QUASAR differ only in the reconstruction w↦q↦r . Standard QAT determines the quantization grid from each group’s extreme weights; QUASAR searches over code assignments and fits the dequantizer to minimize loss-aware reconstruction error. The right panels show the resulting grids and error distributions; empirical measurements appear in Appendix A .
Figure 3: QAD of Qwen3-4B-Thinking-2507 (left) and Llama-3.1-8B-Instruct (right) at INT4, INT3, and INT2 on Open-PerfectBlend. We show training loss over the full run, the final 1,000 steps, and held-out evaluation loss. All losses are forward KL to the full-precision model.
Qwen3-4B-Thinking-2507
Llama-3.1-8B-Instruct
Method
KL ↓
Top-1 ↑
HMMT ’26
AIME ’25
MATH-500
MMLU-Pro
SuperGPQA
LCB v6
LongBench-v2
RULER
Avg.
KL ↓
Top-1 ↑
MATH-500
MMLU-Pro
SuperGPQA
LCB v6
LongBench-v2
RULER
Avg.
Teacher (BF16)
–
–
43.1
77.9
98.2
72.7
48.1
74.6
44.5
87.9
68.4
–
–
49.6
42.3
22.4
17.0
30.0
89.6
41.8
2-bit (INT2)
RTN
11.166
0.5
0.0
0.0
1.9
0.0
10.0
0.0
0.0
0.0
1.5
10.810
0.5
0.0
0.0
10.2
0.0
0.0
0.0
1.7
GPTQ
1.140
67.7
0.0
0.0
2.6
3.2
8.4
0.0
0.0
0.3
1.8
2.380
51.8
2.9
1.8
8.7
0.0
0.0
0.1
2.3
AWQ
3.016
37.0
0.0
0.0
2.4
0.0
5.2
0.0
0.0
0.0
1.0
6.769
12.9
1.6
0.0
10.7
0.0
0.0
0.0
2.0
Table 1: Reasoning, coding, and long-context results for Qwen3-4B-Thinking-2507 and Llama-3.1-8B-Instruct at INT2, INT3, and INT4. We report KL and top-1 agreement between quantized checkpoints and the BF16 original, as well as benchmark accuracy.
2-bit (INT2)
3-bit (INT3)
4-bit (INT4)
Method
PPL ↓
MATH-500
GSM8K
AIME’24
AIME’25
HMMT’25
Avg.
PPL ↓
MATH-500
GSM8K
AIME’24
AIME’25
HMMT’25
Avg.
PPL ↓
MATH-500
GSM8K
AIME’24
AIME’25
HMMT’25
Avg.
Base (no SFT)
–
41.0
66.4
9.6
3.3
0.8
24.2
–
41.0
66.4
9.6
3.3
0.8
24.2
–
41.0
66.4
9.6
3.3
0.8
24.2
FP-SFT
1.59
83.8
91.3
22.9
22.5
10.0
46.1
1.59
83.8
91.3
22.9
22.5
10.0
46.1
1.59
83.8
91.3
22.9
22.5
10.0
46.1
RTN
3.5×105
0.9
0.3
0.0
0.0
0.0
0.2
2.24
12.8
7.6
0.0
1.7
0.0
4.4
1.66
78.2
90.0
15.0
13.3
7.5
40.8
GPTQ
3.12
2.1
1.1
0.0
0.0
0.0
0.6
1.68
71.0
85.0
11.2
11.2
4.2
36.5
1.61
81.7
90.6
19.1
20.5
8.5
44.1
AWQ
130
1.5
1.2
0.0
0.0
0.0
0.5
1.83
57.3
65.1
10.0
8.3
1.7
28.5
1.63
80.3
90.5
15.8
17.5
8.3
42.5
Table 2: Comparison of QAT and PTQ methods for adapting Qwen3-4B-Base on OpenMathReasoning at INT2, INT3, and INT4. We report held-out perplexity and accuracy on five math benchmarks.
Fidelity
Accuracy (%) ↑
Hard reasoning and coding
Long horizon
Model
Method
Format
bpw
KL ↓
Top-1 ↑
HMMT’26
AIME’25
SuperGPQA
OJBench
MRCR
MRCR ≥ 64k
MRCR 8-needle
LongProc
LongBench-v2
RULER
Avg.
BF16
BF16
16
–
–
44.6
68.4
48.8
22.4
25.8
19.2
16.2
48.9
38.4
80.5
41.3
RTN
NVFP4 W4A4
4.50
.053
93.0
37.4
64.6
45.6
14.7
22.4
15.2
14.1
39.3
36.0
76.6
36.6
GPTQ
NVFP4 W4A4
4.50
.039
94.0
40.3
64.0
46.2
16.8
23.6
14.6
15.2
39.5
36.0
78.5
37.5
Standard QAT
NVFP4 W4A4
4.50
.041
94.0
39.0
61.2
45.7
19.4
21.6
14.1
12.9
36.5
37.4
77.9
36.6
Table 3: Comparison of QUASAR with QAT and PTQ baselines across NVFP4, INT4, and Q4_0 (llama.cpp GGUF format). We report fidelity to BF16 and accuracy on ten reasoning, coding, and long-horizon benchmarks.
Figure 4: Ablation Study: INT2 QAD for Qwen3-4B-Thinking-2507 with different choices for QUASAR’s components. Training and evaluation (held-out) loss both use forward KL.
Appendix figures & tables40 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Per-weight loss-aware reconstruction error for 14 Qwen3-4B weight groups at INT3. Color shows ∣h(r−w)∣ under (a) Standard QAT with min–max reconstruction and (b) QUASAR. Each row contains 128 weights ordered by saliency h .
Figure 6: Scale search for Qwen3-4B weight groups at INT3. (a) One extreme weight sets the min–max range. (b) QUASAR clips the outlier and covers the bulk. (c) Distribution of selected range factors f⋆ across 28.4M groups; 99.6% use a range narrower than min–max.
Figure 7: Distribution of selected clipping-range factors f⋆ across weight groups after QAD at INT4, INT3, and INT2 for Qwen3-4B-Thinking-2507 and Llama-3.1-8B-Instruct. Lower bit widths favor narrower ranges.
Figure 8: Saliency h in one Qwen3-4B layer. Weight groups span 128 weights along the input dimension. The large variation within each group lets QUASAR prioritize the most salient weights.
Component
Standard QAT
QUASAR
Teacher forward
0.373
0.373
Student compute
0.366
0.366
Weight reconstruction
0.294
0.335
Loss computation
0.065
0.065
Backward
1.805
1.805
Optimizer update
0.021
0.022
Appendix
Table 4: Wall-clock time of one training step, per component and in seconds: quantization-aware distillation of Qwen3-4B-Thinking-2507 at INT3 on 8x H100 GPUs. QUASAR differs from Standard QAT only in the weight reconstruction row.
Format
Codes
Group
Stored parameters
Fit
Statement
INT sym. (GPTQ layout)
INT4
32–128
scale
Eq. ( 7 )
Alg. 3
INT affine (AWQ layout)
INT4
32–128
scale, int. zero-point
Eq. ( 8 )
Alg. 4
NVFP4
E2M1
16
E4M3 scale, FP32 tensor scale
Eq. ( 7 ), E4M3-rounded
Alg. 5
MXFP4
E2M1
32
E8M0 scale
search only
Alg. 6
Appendix
Table 5: QUASAR variants for deployment formats with fast inference kernels. The formats differ in their code grids, group sizes, stored parameters, and dequantization constraints.
Figure 9: Comparison of saliency weights ht and ht2 . We report held-out loss under AdamW at INT2, INT3, and INT4 (top) and under SGD at three INT2 learning rates (bottom).
Figure 10: Reconstruction error tracks final KL. End-of-training loss-aware reconstruction error versus final held-out KL to the full-precision teacher, for Standard QAT, Denoising QAT, and QUASAR at INT4, INT3, and INT2.
Figure 11: QUASAR reduces reconstruction error across projection types. Median reduction in loss-aware reconstruction error relative to Standard QAT across Qwen3-4B-Thinking-2507 modules at INT4, INT3, and INT2, measured at initialization (left) and after training (right). Error is measured against the full-precision weights.
Figure 12: Reconstruction error drives final loss. Five QUASAR INT2 runs differ only in the lower bound imposed on the scale-search factor f . Restricting the search raises both the final reconstruction error ST and held-out KL after 1,000 training steps. The star marks the unrestricted search.
Figure 13: Reconstruction error predicts final loss before training. Across Standard QAT, Denoising QAT, and QUASAR at INT4, INT3, and INT2, reconstruction error at initialization predicts held-out KL to the full-precision model after 4,096 steps ( R2=0.98 ). Results use Qwen3-4B-Thinking-2507.
Figure 14: Reconstruction error controls gradient mismatch. Across INT4, INT3, and INT2, the squared gradient mismatch rises with the loss-aware reconstruction error S , supporting Assumption 3. Circles, triangles, and squares show fixed clipping factors for INT4, INT3, and INT2, respectively; stars show QUASAR’s per-group scale-search solutions.
Figure 15: QUASAR’s saliency tracks curvature. Each point represents one weight matrix from a trained INT2 Qwen3-4B-Thinking checkpoint and compares its mean saliency h (Adam’s second moment) with a 32-sample Hutchinson estimate of its mean Hessian diagonal. The log-scale correlation is r=0.81 ( Dong et al., 2020 ; Yao et al., 2020 ) .
Figure 16: Training trajectories support a PL relation. Each point is one step of an INT2 Qwen3-4B-Thinking run, with loss and gradient measured at the reconstruction r . For all three methods, the squared gradient norm scales with the gap above the trajectory minimum. QUASAR’s estimated μ^=1.72 is about 2.5× the baseline values. In Theorem 1 (ii), a larger μ gives faster geometric convergence and smaller floor terms, which scale with 1/μ or 1/μ2 . This provides an optimization-side view of the healing curves (Figure 3 ): QUASAR’s reconstruction adds less loss, so more of the remaining gap produces useful gradient, enabling faster healing and a lower loss floor.
Figure 17: Reconstruction error dominates the measured bound. We evaluate each term of Theorem 1 (i) on an INT2 QUASAR run with plain SGD and learning rate 10−3 . We measure λmax by power iteration (approximating L ), σ2 from per-batch gradients, and C using Figure 14 . The resulting bound is within 3.1× of the observed average squared gradient norm, and reconstruction error is its largest term.
Figure 18: QUASAR retains its advantage under plain SGD. Training (top) and held-out (bottom) loss for Standard QAT and QUASAR at INT2 across three constant learning rates. QUASAR reaches lower loss floors at every stable learning rate.
Qwen3-4B-Thinking-2507
Llama-3.1-8B-Instruct
Method
KL ↓
Top-1 ↑
GSM8K
MMLU
ARC-C
ARC-E
HellaSwag
WinoGrande
TruthfulQA
IFEval
Avg.
KL ↓
Top-1 ↑
GSM8K
MMLU
ARC-C
ARC-E
HellaSwag
WinoGrande
TruthfulQA
IFEval
Avg.
Teacher (FP16)
–
–
87.0
68.7
53.1
76.7
65.7
66.1
57.5
54.3
66.1
–
–
70.1
68.3
55.6
79.9
79.5
73.6
54.5
73.9
69.4
2-bit (INT2)
RTN
11.166
0.5
0.0
25.4
25.6
25.9
26.0
52.0
46.8
8.3
26.3
10.810
0.5
0.0
25.5
26.2
26.1
26.1
51.8
47.6
11.8
26.9
GPTQ
1.140
67.7
0.2
23.4
25.1
30.0
30.1
49.3
52.9
8.9
27.5
2.380
51.8
0.0
24.9
23.3
28.9
30.0
48.0
49.0
9.1
26.7
AWQ
3.016
37.0
0.0
23.4
23.2
31.6
29.1
51.1
52.8
9.6
27.6
6.769
12.9
0.0
23.6
23.3
25.9
27.2
48.2
48.8
7.9
25.6
Appendix
Table 6: QUASAR gives the highest fidelity in all six model–bit-width settings. Held-out forward KL, top-1 agreement, and accuracy on eight benchmarks for QAD of Qwen3-4B-Thinking-2507 and Llama-3.1-8B-Instruct.
Figure 19: QUASAR gives the lowest held-out KL across both models and all bit widths. Forward KL to the full-precision model after QAD of Qwen3-4B-Thinking-2507 (left) and Llama-3.1-8B-Instruct (right). Lower is better.
Figure 20: QUASAR attains the highest INT2 average accuracy on both models. Average accuracy over the eight benchmarks of Table 6 versus bit width for Qwen3-4B-Thinking-2507 (left) and Llama-3.1-8B-Instruct (right). Gray circles mark the full-precision models.
Figure 21: QUASAR gives the best fidelity at each quantized model size. Average accuracy over eight benchmarks (left) and held-out KL to the full-precision model (right) versus model size. Filled markers denote Qwen3-4B-Thinking-2507, open markers denote Llama-3.1-8B-Instruct, and gray circles mark the full-precision models.
Figure 22: QUASAR preserves INT2 accuracy broadly across tasks. Per-task accuracy normalized by the full-precision model for Qwen3-4B-Thinking-2507 (left) and Llama-3.1-8B-Instruct (right) on the eight benchmarks of Table 6 .
Figure 23: Lower held-out KL correlates with higher downstream accuracy. Average accuracy over eight benchmarks versus final forward KL for the QAT and PTQ methods of Table 6 , across both models and all three bit widths. Dashed lines mark full-precision accuracy.
Figure 24: QUASAR reaches the highest final top-1 agreement in all six settings. Top-1 agreement with the full-precision model during QAD of Qwen3-4B-Thinking-2507 (top) and Llama-3.1-8B-Instruct (bottom) at INT4, INT3, and INT2.
Figure 25: QUASAR is robust to learning rate. Final held-out KL after 1,000 steps of INT2 QAD of Qwen3-4B-Thinking-2507 across a 50× learning-rate range. QUASAR attains the lowest KL at every tested rate; Standard QAT diverges at high rates, while BitDistiller degrades at low rates.
Figure 26: QUASAR becomes the final leader early in training. Each strip identifies the method with the lowest held-out KL among six methods for Qwen3-4B-Thinking-2507. QUASAR leads from step 128 at INT3 and INT2, and from step 1,408 at INT4, through the end of training.
Figure 27: The judge prefers QUASAR to every INT2 quantized baseline. Pairwise win, tie, and loss rates on responses to 128 WildChat prompts for Qwen3-4B-Thinking-2507 and Llama-3.1-8B-Instruct. Llama-3.3-70B-Instruct serves as the judge.
vs. QUASAR
Qwen3-4B-Thinking-2507
Llama-3.1-8B-Instruct
RTN
+1.00[+1.00,+1.00]
+0.98[+0.96,+1.00]
GPTQ
+1.00[+1.00,+1.00]
+0.97[+0.94,+0.99]
AWQ
+1.00[+1.00,+1.00]
+0.99[+0.97,+1.00]
Standard QAT
+0.32[+0.21,+0.42]
+0.97[+0.94,+0.99]
LSQ
+0.27[+0.15,+0.38]
+0.37[+0.25,+0.49]
Denoising QAT
+0.31[+0.21,+0.41]
+0.98[+0.95,+1.00]
Appendix
Table 7: QUASAR wins all 14 comparisons with quantized baselines. Mean per-prompt preference in [−1,1] , with 95% bootstrap confidence intervals over 128 prompts. Positive values favor QUASAR. The final row compares QUASAR with the full-precision model.
Figure 28: QUASAR matches the full-precision models’ response stability at INT2. Mean word-repetition ratio over 128 WildChat prompts for Qwen3-4B-Thinking-2507 (left) and Llama-3.1-8B-Instruct (right), with greedy decoding. Dashed lines mark the full-precision models; counts show responses that close the reasoning block for Qwen or stop before the token cap for Llama.
Prompt: Write a story about a man being stranded on an island.
FP16
The salt stung Leo’s eyes before he even registered the crash. One moment, he was shouting into the cockpit of his battered sedan, the next, the world …
2-bit (INT2)
RTN
max o mil MD o. over md edge md mdish ai } ats } } further display ∘ scs <igrated found ”arql eventually o, same, if ( cs1) < < ∘ <ql <, \ql o < <! ( < …
The first thing Elias noticed wasn’t the ocean, but the absence of it. Not the roar of waves, but the silence of his own thoughts, a hollow space wh …
Appendix
Figure 29: The beginning of each method’s response to one held-out prompt, for Qwen3-4B-Thinking-2507 at INT2, INT3, and INT4 with greedy decoding.
Variant
Standard QAT
Adam fit
Init. only
Uniform fit
Fisher fit
QUASAR (ours)
Scale search
none
none
at init.
every step
every step
every step
Saliency
none
Adam 2nd moment
uniform
uniform
Fisher
Adam 2nd moment
KL ↓
0.175
0.164
0.172
0.114
0.113
0.112
Δ vs. Standard QAT
–
−6.1%
−1.6%
−34.7%
−35.5%
−35.9%
Appendix
Table 8: Ablation Study: Effect of QUASAR’s scale search and weighted dequantization fit on Qwen3-4B-Thinking-2507 with INT2 QAD. We report final held-out KL and its change from Standard QAT.
Figure 30: QUASAR reaches the lowest held-out perplexity at every bit width. Training perplexity over the full run (left) and final 1,000 steps (center), and held-out perplexity (right), for QAT of Qwen3-4B-Base on OpenMathReasoning. FP-SFT is included as a full-precision reference.
Figure 31: Answer outcomes for the INT2 checkpoints. Share of generated samples that produce a correct answer, an incorrect answer, or no final answer on MATH-500 and GSM8K. The base model and FP-SFT are included as references.
Figure 32: Variation across repeated INT2 samples. Ratio of pass@8 to avg@1 on MATH-500 for Qwen3-4B-Base checkpoints, using eight samples per problem at temperature 0.6. A ratio of 1 means that the same problems are solved on every attempt; higher values indicate less consistent successes. The base model and FP-SFT are included as references.
Figure 33: Divergence from FP-SFT along reasoning traces. Forward KL to FP-SFT by trace decile for selected quantized checkpoints, averaged over 1,000 MATH-500 traces generated by FP-SFT. Every checkpoint is evaluated on the same token sequences.
Figure 34: QUASAR reaches lower held-out loss for both NVFP4 models. Training forward KL over the full run (left) and final 1,000 steps (center), and held-out forward KL (right), for NVFP4 QAD of Qwen3-8B (top) and Qwen3.5-9B (bottom).
Downstream task accuracy (%) ↑
Method
Avg.
GSM8K
MMLU
ARC-C
ARC-E
Hella.
WinoG.
TQA
IFEval
GPQA-D
Qwen3-8B
BF16
68.6
86.7
72.9
56.7
80.9
75.0
67.7
54.4
84.8
38.4
Standard QAT
67.1
83.9
70.8
55.4
80.3
73.6
67.1
53.9
80.6
37.9
[][] QUASAR (ours)
67.6
87.0
71.1
55.7
80.4
73.6
66.8
54.2
81.2
38.4
Qwen3.5-9B
Appendix
Table 9: QUASAR gives the best average accuracy among the quantized checkpoints. Accuracy on nine downstream benchmarks for Qwen3-8B and Qwen3.5-9B. The materialized NVFP4 checkpoints are evaluated through vLLM’s native W4A4 path.
Figure 35: QUASAR better preserves long-horizon behavior in NVFP4. Expanded Qwen3.5-9B results from Table 3 : (a) OpenAI-MRCR score by context length; (b) share of long generations that terminate within the 32k-token output budget. Higher is better.
Method
MATH-500
LiveCodeBench v6
OJBench (C++)
OJBench (Python)
Qwen3-8B
BF16
0.5
13.6
1.3
2.2
RTN
1.4
40.8
20.3
26.7
GPTQ
0.5
24.3
11.6
10.3
Standard QAT
0.5
21.6
4.7
5.6
[][] QUASAR (ours)
0.5
18.4
3.0
2.2
Appendix
Table 10: QUASAR has the lowest or tied-lowest non-termination rate among the quantized checkpoints. Percentage of generations that reach the 32,768-token output budget without stopping, using the same generations as Table 3 . Lower is better.
Model
Format
bpw
Fidelity
Accuracy (%) ↑
Qwen3.5-4B
KL ↓
Top-1 ↑
GSM8K-P
MMLU-P
IFEval
MATH
AIME25
GPQA-D
Avg.
BF16
BF16
16
–
–
94.0
78.9
88.0
84.4
79.2
77.3
83.6
RTN
NVFP4 W4A16
4.50
.042
93.8
94.9
77.5
86.3
83.8
67.1
74.1
80.6
GPTQ
NVFP4 W4A16
4.50
.023
95.7
95.1
75.2
84.5
83.2
52.9
69.9
76.8
GPTQ
INT4 g128
7.49
.049
93.1
94.0
77.5
87.1
83.0
65.0
71.2
79.6
AWQ
INT4 g32
4.50
.042
93.9
93.9
77.6
87.6
84.2
72.5
74.6
81.7
Appendix
Table 11: Comparison with released 4-bit checkpoints across four model families. We report fidelity to BF16 and accuracy on each model’s benchmark suite. GGUF KL values are comparable only within each block. ‡ denotes mixed NVFP4, FP8, and BF16 precision.
Model
Row
Artifact
Qwen3.5-4B
BF16
Qwen/Qwen3.5-4B ( Qwen Team, 2026a )
GPTQ (INT4 g128)
RedHatAI/Qwen3.5-4B-quantized.w4a16 ( Red Hat AI, 2026b )
Post-training quantization (PTQ) converts a trained full-precision model into low-bit weights without task-level retraining, while quantization-aware training (QAT) incorporates quantization into the training loop. Although PTQ is efficient and often accurate at moderate bitwidths, it can fail sharply at aggressive bitwidths; QAT is more expensive but can often recover the lost accuracy. We propose a unified geometric framework that explains both PTQ failure and QAT recovery. We model full-precision training as following a low-loss \emph{river} inside a wider \emph{valley}: a normal neighborhood of the river forms a nearly flat \emph{basin}, while leaving this basin incurs a sharp loss increase. When the quantization grid is comparable to the basin width, local PTQ objectives, including rounding and Hessian-based second-order reconstruction, can select a high-loss deployed quantized point outside the basin even when nearby low-loss quantized points exist. In this regime, straight-through-estimator-based QAT has a useful bias: it evaluates gradients at the deployed quantized weights while updating latent full-precision weights, causing the gradient to sense the valley wall and acquire an inward component that steers subsequent quantized iterates back into the basin. We formalize this mechanism through a local landscape model, construct a geometric PTQ failure mode, and prove finite-time QAT recovery under local quantizer-compatibility assumptions. Experiments across vision and language models under multiple neural-network quantization schemes corroborate the predicted basin-crossing failure of PTQ and the corresponding recovery mechanism of QAT.
Hanyang Li, Jianhao Ma, Ying Cui
Department of IEOR, University of California, Berkeley · Department of Statistics and Data Science, University of Pennsylvania
The rapid development of LLMs incurs prohibitive memory footprints and intensive computational demands. Quantization-Aware Training (QAT) techniques have emerged as a promising solution to address these challenges by explicitly simulating quantization effects during model training, yielding low-bit models that achieve accuracy comparable to their full-precision counterparts. In this work, we provide a target-centric survey of QAT, aimed at clarifying both its theoretical foundations and its evolving implementation landscape. We systematically review existing QAT methods through a target-centric taxonomy and synthesize cross-target differences in error characteristics, numerical formats, and strategy transferability. We further summarize QAT evaluation paradigms and discuss challenges in optimization and deployment, outlining potential directions for future research.
Jiamin Song, Mengjie Zhao, Zijing Wang +6
1Northeastern University, China · 2Shandong University, China · 3CIS, LMU Munich; MCML, Germany
Quantization-aware training (QAT) is widely deployed but typically relies on the Straight-Through Estimator (STE), which passes gradients through non-differentiable quantizers by fiat. This often makes training brittle near bin boundaries and weakly aligned with the actual behavior of the low-precision model. We introduce JacQuant, a QAT framework that learns a lightweight surrogate of the model's local sensitivity to parameter changes and uses it to stabilize and accelerate training within standard variance-reduced optimizers. The surrogate is inexpensive (diagonal or block-diagonal), data-driven, and compatible with common weight and activation quantizers. On code-preserving training phases, we prove convergence for non-convex objectives and obtain linear rates under a PL condition, and we relate the learned sensitivity to end-to-end output fidelity via a simple calibration argument. Across LLM benchmarks at ≤2 bits, JacQuant consistently reaches higher accuracy than STE-based QAT, and the runtime analyses on various models show that the added cost remains negligible under practical group sizes. The method is drop-in and requires no changes to the forward quantizers; our empirical claims are scoped to ultra-low-bit LLM QAT.