QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction
Organizations: Cornell University · Together AI
Abstract
As large language model inference shifts toward lower precision, post-training quantization (PTQ) becomes increasingly brittle, making quantization-aware training (QAT) essential for preserving model quality. However, QAT has a structural mismatch: gradient updates are applied to latent full-precision weights, while the loss and gradients are computed on lossy reconstructions of those weights. This mismatch can lead to suboptimal training trajectories and a higher loss floor. Second-order PTQ methods address a similar problem by minimizing loss-aware reconstruction error, but applying such expensive reconstruction repeatedly during QAT as the weights evolve is impractical. We introduce QUASAR, a QAT method that brings lightweight, loss-aware reconstruction into the training loop. At each training step, QUASAR reconstructs the latent weights by searching over a small set of clipping ranges and fitting dequantization parameters through saliency-weighted least squares. We use an exponential moving average of squared gradients as the per-parameter saliency signal. Our theoretical analysis shows that optimizing QUASAR's reconstruction objective tightens both the convergence and final-loss bounds of QAT. We evaluate QUASAR across four model families and across INT4, INT3, INT2, and NVFP4 quantization formats. QUASAR consistently achieves lower training and evaluation loss than competitive QAT methods and outperforms QAT and PTQ baselines on downstream benchmarks. At INT2, QUASAR improves average accuracy over the best QAT baseline by 13.3 points with quantization-aware distillation and by 10.9 points with QAT on mathematical reasoning data. Notably, after distillation on only about 600M tokens, QUASAR's INT4 Gemma-4 E4B checkpoint outperforms the corresponding QAT checkpoint released by Google, with 66% lower KL divergence and 1.8 points higher average accuracy.
Figures & tables
| Qwen3-4B-Thinking-2507 | Llama-3.1-8B-Instruct | |||||||||||||||||||
| Method | KL | Top-1 | HMMT ’26 | AIME ’25 | MATH-500 | MMLU-Pro | SuperGPQA | LCB v6 | LongBench-v2 | RULER | Avg. | KL | Top-1 | MATH-500 | MMLU-Pro | SuperGPQA | LCB v6 | LongBench-v2 | RULER | Avg. |
| Teacher (BF16) | – | – | 43.1 | 77.9 | 98.2 | 72.7 | 48.1 | 74.6 | 44.5 | 87.9 | 68.4 | – | – | 49.6 | 42.3 | 22.4 | 17.0 | 30.0 | 89.6 | 41.8 |
| 2-bit (INT2) | ||||||||||||||||||||
| RTN | 11.166 | 0.5 | 0.0 | 0.0 | 1.9 | 0.0 | 10.0 | 0.0 | 0.0 | 0.0 | 1.5 | 10.810 | 0.5 | 0.0 | 0.0 | 10.2 | 0.0 | 0.0 | 0.0 | 1.7 |
| GPTQ | 1.140 | 67.7 | 0.0 | 0.0 | 2.6 | 3.2 | 8.4 | 0.0 | 0.0 | 0.3 | 1.8 | 2.380 | 51.8 | 2.9 | 1.8 | 8.7 | 0.0 | 0.0 | 0.1 | 2.3 |
| AWQ | 3.016 | 37.0 | 0.0 | 0.0 | 2.4 | 0.0 | 5.2 | 0.0 | 0.0 | 0.0 | 1.0 | 6.769 | 12.9 | 1.6 | 0.0 | 10.7 | 0.0 | 0.0 | 0.0 | 2.0 |
| 2-bit (INT2) | 3-bit (INT3) | 4-bit (INT4) | |||||||||||||||||||
| Method | PPL | MATH-500 | GSM8K | AIME’24 | AIME’25 | HMMT’25 | Avg. | PPL | MATH-500 | GSM8K | AIME’24 | AIME’25 | HMMT’25 | Avg. | PPL | MATH-500 | GSM8K | AIME’24 | AIME’25 | HMMT’25 | Avg. |
| Base (no SFT) | – | 41.0 | 66.4 | 9.6 | 3.3 | 0.8 | 24.2 | – | 41.0 | 66.4 | 9.6 | 3.3 | 0.8 | 24.2 | – | 41.0 | 66.4 | 9.6 | 3.3 | 0.8 | 24.2 |
| FP-SFT | 1.59 | 83.8 | 91.3 | 22.9 | 22.5 | 10.0 | 46.1 | 1.59 | 83.8 | 91.3 | 22.9 | 22.5 | 10.0 | 46.1 | 1.59 | 83.8 | 91.3 | 22.9 | 22.5 | 10.0 | 46.1 |
| RTN | 0.9 | 0.3 | 0.0 | 0.0 | 0.0 | 0.2 | 2.24 | 12.8 | 7.6 | 0.0 | 1.7 | 0.0 | 4.4 | 1.66 | 78.2 | 90.0 | 15.0 | 13.3 | 7.5 | 40.8 | |
| GPTQ | 3.12 | 2.1 | 1.1 | 0.0 | 0.0 | 0.0 | 0.6 | 1.68 | 71.0 | 85.0 | 11.2 | 11.2 | 4.2 | 36.5 | 1.61 | 81.7 | 90.6 | 19.1 | 20.5 | 8.5 | 44.1 |
| AWQ | 130 | 1.5 | 1.2 | 0.0 | 0.0 | 0.0 | 0.5 | 1.83 | 57.3 | 65.1 | 10.0 | 8.3 | 1.7 | 28.5 | 1.63 | 80.3 | 90.5 | 15.8 | 17.5 | 8.3 | 42.5 |
| Fidelity | Accuracy (%) | |||||||||||||||
| Hard reasoning and coding | Long horizon | |||||||||||||||
| Model | Method | Format | bpw | KL | Top-1 | HMMT’26 | AIME’25 | SuperGPQA | OJBench | MRCR | MRCR 64k | MRCR 8-needle | LongProc | LongBench-v2 | RULER | Avg. |
| BF16 | BF16 | 16 | – | – | 44.6 | 68.4 | 48.8 | 22.4 | 25.8 | 19.2 | 16.2 | 48.9 | 38.4 | 80.5 | 41.3 | |
| RTN | NVFP4 W4A4 | 4.50 | .053 | 93.0 | 37.4 | 64.6 | 45.6 | 14.7 | 22.4 | 15.2 | 14.1 | 39.3 | 36.0 | 76.6 | 36.6 | |
| GPTQ | NVFP4 W4A4 | 4.50 | .039 | 94.0 | 40.3 | 64.0 | 46.2 | 16.8 | 23.6 | 14.6 | 15.2 | 39.5 | 36.0 | 78.5 | 37.5 | |
| Standard QAT | NVFP4 W4A4 | 4.50 | .041 | 94.0 | 39.0 | 61.2 | 45.7 | 19.4 | 21.6 | 14.1 | 12.9 | 36.5 | 37.4 | 77.9 | 36.6 | |
Appendix figures & tables40 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Standard QAT | QUASAR |
|---|---|---|
| Teacher forward | 0.373 | 0.373 |
| Student compute | 0.366 | 0.366 |
| Weight reconstruction | 0.294 | 0.335 |
| Loss computation | 0.065 | 0.065 |
| Backward | 1.805 | 1.805 |
| Optimizer update | 0.021 | 0.022 |
| Format | Codes | Group | Stored parameters | Fit | Statement |
|---|---|---|---|---|---|
| INT sym. (GPTQ layout) | INT4 | 32–128 | scale | Eq. ( 7 ) | Alg. 3 |
| INT affine (AWQ layout) | INT4 | 32–128 | scale, int. zero-point | Eq. ( 8 ) | Alg. 4 |
| NVFP4 | E2M1 | 16 | E4M3 scale, FP32 tensor scale | Eq. ( 7 ), E4M3-rounded | Alg. 5 |
| MXFP4 | E2M1 | 32 | E8M0 scale | search only | Alg. 6 |
| Qwen3-4B-Thinking-2507 | Llama-3.1-8B-Instruct | |||||||||||||||||||||
| Method | KL | Top-1 | GSM8K | MMLU | ARC-C | ARC-E | HellaSwag | WinoGrande | TruthfulQA | IFEval | Avg. | KL | Top-1 | GSM8K | MMLU | ARC-C | ARC-E | HellaSwag | WinoGrande | TruthfulQA | IFEval | Avg. |
| Teacher (FP16) | – | – | 87.0 | 68.7 | 53.1 | 76.7 | 65.7 | 66.1 | 57.5 | 54.3 | 66.1 | – | – | 70.1 | 68.3 | 55.6 | 79.9 | 79.5 | 73.6 | 54.5 | 73.9 | 69.4 |
| 2-bit (INT2) | ||||||||||||||||||||||
| RTN | 11.166 | 0.5 | 0.0 | 25.4 | 25.6 | 25.9 | 26.0 | 52.0 | 46.8 | 8.3 | 26.3 | 10.810 | 0.5 | 0.0 | 25.5 | 26.2 | 26.1 | 26.1 | 51.8 | 47.6 | 11.8 | 26.9 |
| GPTQ | 1.140 | 67.7 | 0.2 | 23.4 | 25.1 | 30.0 | 30.1 | 49.3 | 52.9 | 8.9 | 27.5 | 2.380 | 51.8 | 0.0 | 24.9 | 23.3 | 28.9 | 30.0 | 48.0 | 49.0 | 9.1 | 26.7 |
| AWQ | 3.016 | 37.0 | 0.0 | 23.4 | 23.2 | 31.6 | 29.1 | 51.1 | 52.8 | 9.6 | 27.6 | 6.769 | 12.9 | 0.0 | 23.6 | 23.3 | 25.9 | 27.2 | 48.2 | 48.8 | 7.9 | 25.6 |
| vs. QUASAR | Qwen3-4B-Thinking-2507 | Llama-3.1-8B-Instruct |
|---|---|---|
| RTN | ||
| GPTQ | ||
| AWQ | ||
| Standard QAT | ||
| LSQ | ||
| Denoising QAT |
| Prompt: Write a story about a man being stranded on an island. | |
|---|---|
| FP16 | The salt stung Leo’s eyes before he even registered the crash. One moment, he was shouting into the cockpit of his battered sedan, the next, the world … |
| 2-bit (INT2) | |
| RTN | max o mil MD o. over md edge md mdish ai } ats } } further display scs <igrated found ”arql eventually o, same, if ( cs1) < < <ql <, \ql o < <! ( < … |
| GPTQ | *A [R] :Help ”l” (or: ”a” in 1000+). 0 * F : 0 (To the 0) : 0 : 0 : 0 : 0 : 0 : 0 : 0 : 0 : 0 : 10000: 0 : 10000: 0 : 0 : 0 : 0 : 0 : 0 : 0 : 0 : 0 : … |
| AWQ | go {r=0 {r=1, 100 100 100 100 REEN .1 100 100 10 100. 100 10 10 1 10 1 1 1 0 1 100 1 1 1 1 1 ( 1 1 2 1 02 1 1 1 1 1 1 1. 1 1 1 1 1 1 1 1 1 1 1 … |
| Standard QAT | The first thing Elias noticed wasn’t the ocean, but the absence of it. Not the roar of waves, but the silence of his own thoughts, a hollow space wh … |
| Variant | Standard QAT | Adam fit | Init. only | Uniform fit | Fisher fit | QUASAR (ours) |
| Scale search | none | none | at init. | every step | every step | every step |
| Saliency | none | Adam 2nd moment | uniform | uniform | Fisher | Adam 2nd moment |
| KL | 0.175 | 0.164 | 0.172 | 0.114 | 0.113 | 0.112 |
| vs. Standard QAT | – |
| Downstream task accuracy (%) | ||||||||||
| Method | Avg. | GSM8K | MMLU | ARC-C | ARC-E | Hella. | WinoG. | TQA | IFEval | GPQA-D |
| Qwen3-8B | ||||||||||
| BF16 | 68.6 | 86.7 | 72.9 | 56.7 | 80.9 | 75.0 | 67.7 | 54.4 | 84.8 | 38.4 |
| Standard QAT | 67.1 | 83.9 | 70.8 | 55.4 | 80.3 | 73.6 | 67.1 | 53.9 | 80.6 | 37.9 |
| [][] QUASAR (ours) | 67.6 | 87.0 | 71.1 | 55.7 | 80.4 | 73.6 | 66.8 | 54.2 | 81.2 | 38.4 |
| Qwen3.5-9B | ||||||||||
| Method | MATH-500 | LiveCodeBench v6 | OJBench (C++) | OJBench (Python) |
|---|---|---|---|---|
| Qwen3-8B | ||||
| BF16 | 0.5 | 13.6 | 1.3 | 2.2 |
| RTN | 1.4 | 40.8 | 20.3 | 26.7 |
| GPTQ | 0.5 | 24.3 | 11.6 | 10.3 |
| Standard QAT | 0.5 | 21.6 | 4.7 | 5.6 |
| [][] QUASAR (ours) | 0.5 | 18.4 | 3.0 | 2.2 |
| Model | Format | bpw | Fidelity | Accuracy (%) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Qwen3.5-4B | KL | Top-1 | GSM8K-P | MMLU-P | IFEval | MATH | AIME25 | GPQA-D | Avg. | ||||
| BF16 | BF16 | 16 | – | – | 94.0 | 78.9 | 88.0 | 84.4 | 79.2 | 77.3 | 83.6 | ||
| RTN | NVFP4 W4A16 | 4.50 | .042 | 93.8 | 94.9 | 77.5 | 86.3 | 83.8 | 67.1 | 74.1 | 80.6 | ||
| GPTQ | NVFP4 W4A16 | 4.50 | .023 | 95.7 | 95.1 | 75.2 | 84.5 | 83.2 | 52.9 | 69.9 | 76.8 | ||
| GPTQ | INT4 g128 | 7.49 | .049 | 93.1 | 94.0 | 77.5 | 87.1 | 83.0 | 65.0 | 71.2 | 79.6 | ||
| AWQ | INT4 g32 | 4.50 | .042 | 93.9 | 93.9 | 77.6 | 87.6 | 84.2 | 72.5 | 74.6 | 81.7 | ||
| Model | Row | Artifact |
| Qwen3.5-4B | BF16 | Qwen/Qwen3.5-4B ( Qwen Team, 2026a ) |
| GPTQ (INT4 g128) | RedHatAI/Qwen3.5-4B-quantized.w4a16 ( Red Hat AI, 2026b ) | |
| AWQ (INT4 g32) | cyankiwi/Qwen3.5-4B-AWQ-4bit ( cyankiwi, 2026 ) | |
| ModelOpt PTQ | cosmicproc/Qwen3.5-4B-NVFP4 ( cosmicproc, 2026 ) | |
| imatrix (GGUF) | bartowski/Qwen_Qwen3.5-4B-GGUF ( bartowski, 2026 ) | |
| KD-QAT g64 (GGUF) | YoozLabs/Qwen3.5-4B-qat-GGUF ( YoozLabs, 2026 ) |