cs.CLOct 7, 2026

OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models

Authors: Wenjun Wang, Heng Li, Yanggan Gu, Hongxia Yang

Organizations: The Hong Kong Polytechnic University · Sun Yat-sen University · PolyU-Daya Bay Technology and Innovation Research Institute

Abstract

Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits. Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated answers, whereas the deployed quantized model condi- tions on prefixes generated by itself. Quantization errors can therefore move the model into states that are absent from offline recovery data. We introduce OnlineQAT, a two-stage framework that first obtains a usable low-bit initialization through block-wise QAT and then performs on-policy distillation (OPD) on student-generated responses. At each visited pre- fix, a frozen full-precision teacher provides a sampled reverse-KL training signal. On Qwen3-1.7B, OnlineQAT obtains the best average among the compared quantized methods: 57.28 at W3A16 and 32.52 at W2A16, im- proving over ReasoningQAT by 2.90 and 0.44 points, respectively. The results suggest that student-visited states provide a useful recovery signal beyond fixed-completion training, particularly at three bits.

Figures & tables

Explore similar work

CardsList
  1. Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning

    Sep 22, 2026Yuanteng Chen, Zhilei Liu, Peisong Wang +7Transport

  2. LC-QAT: Data-Efficient 2-Bit QAT for LLMs via Linear-Constrained Vector Quantization

    Date pendingHaoyu Wang, Xingyu Yu, Haiyan Zhao +2Large Language Model QuantizationQuantization-Aware Training

  3. Statistically-Lossless Quantization of Large Language Models

    May 4, 2026Michael Helcig, Eldar Kurtic, Dan AlistarhLarge Language Model QuantizationQuantized