OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models
Organizations: The Hong Kong Polytechnic University · Sun Yat-sen University · PolyU-Daya Bay Technology and Innovation Research Institute
Abstract
Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits. Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated answers, whereas the deployed quantized model condi- tions on prefixes generated by itself. Quantization errors can therefore move the model into states that are absent from offline recovery data. We introduce OnlineQAT, a two-stage framework that first obtains a usable low-bit initialization through block-wise QAT and then performs on-policy distillation (OPD) on student-generated responses. At each visited pre- fix, a frozen full-precision teacher provides a sampled reverse-KL training signal. On Qwen3-1.7B, OnlineQAT obtains the best average among the compared quantized methods: 57.28 at W3A16 and 32.52 at W2A16, im- proving over ReasoningQAT by 2.90 and 0.44 points, respectively. The results suggest that student-visited states provide a useful recovery signal beyond fixed-completion training, particularly at three bits.
Figures & tables
| Method | GSM8K | MMLU-R | MATH | IFEval | LCB | GPQA-D | AVG |
|---|---|---|---|---|---|---|---|
| BF16 (Think) | 89.9 | 73.9 | 93.4 | 72.5 | 33.2 | 40.1 | 67.17 |
| W3-Stg1 | 84.2 | 65.3 | 64.8 | 57.3 | 22.3 | 18.2 | 52.02 |
| W3-KD | 85.4 | 66.0 | 67.4 | 58.3 | 26.7 | 18.2 | 53.67 |
| W3-ReasoningQAT | 86.0 | 65.8 | 71.6 | 60.3 | 24.4 | 18.2 | 54.38 |
| W3-OnlineQAT | 86.3 | 68.0 | 73.2 | 59.5 | 27.4 | 29.3 | 57.28 |
| W2-Stg1 | 24.40 | 17.20 | 21.80 | 33.20 | 0.50 | 11.60 | 18.12 |