Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive post-training quantization (PTQ) can degrade model quality. We present QATFactory, an open-source framework for deployment-aligned quantization-aware distillation (QAD) and reinforcement learning (QARL). QATFactory simulates deployment-time quantization while performing matrix multiplications in BF16, allowing models to adapt to quantization noise without requiring training hardware that natively supports the target format; for example, it supports NVFP4 training on H100 GPUs, which lack FP4 Tensor Cores. The framework supports NVFP4, MXFP4, and llama.cpp's Q4_K format; dense and mixture-of-experts models; and both full-parameter and LoRA-based training. It exports checkpoints directly to vLLM and llama.cpp without an additional lossy conversion step or added inference overhead. With QATFactory, we conduct extensive experiments on models ranging from 8B to 230B parameters and evaluate exported checkpoints in production inference engines. Across models and formats, QAD consistently improves deployed-model quality over strong PTQ baselines. On Qwen3.5-9B, QAD achieves average benchmark accuracies of 68.9% under NVFP4 and 66.0% under MXFP4, outperforming the best PTQ results of 65.4% and 56.4%, respectively. Through our experiments, we found that although both FP4 formats quantize weights and activations at deployment, the best training strategy is format-dependent: NVFP4 generally performs better when only weights are quantized during training, whereas MXFP4 benefits from quantizing both weights and activations. At a fixed training token budget, training on fewer 32K sequences improves average accuracy by 1.9 points over training on more 4K sequences.
Figures & tables
Figure 1: Comparison of post-training quantization (PTQ) and quantization-aware training (QAT) in QATFactory . PTQ directly quantizes a pretrained model and can lead to accuracy degradation. QATFactory instead trains the model with quantization in the loop using self-distillation or reinforcement learning, then exports deployment-ready low-precision checkpoints. On Qwen3.5-9B, QATFactory produces higher-quality NVFP4 models than RTN and GPTQ on both LiveCodeBench and BigCodeBench.
Figure 2: Overview of QATFactory. (a) Quantization-aware distillation (QAD): a frozen high-precision teacher supervises a quantization-aware student initialized from the same pretrained checkpoint. (b) Quantization-aware reinforcement learning (QARL): a low-precision inference engine generates rollouts, reward functions or verifiers evaluate them, and updated weights are synchronized back to the rollout engine. Right: the shared internal structure of a quantization-aware layer. Latent weights and activations undergo simulated quantization in BF16, and parameter updates flow through a straight-through estimator (STE).
NVFP4
MXFP4
Q4_K
Precision
W4A4
W4A4
W4A16 (Weight-only)
Element format
E2M1 (FP4)
E2M1 (FP4)
Unsigned INT4
Block structure
Block size 16
Block size 32
32 (in 256 super-block)
Scale hierarchy
Block E4M3, Tensor FP32
Block E8M0
Block 6-bit, Super-block FP16
Symmetry
Symmetric
Symmetric
Asymmetric
Effective bits / weight
4.50 bits
4.25 bits
4.50 bits
Table 1: Quantization formats supported by QATFactory. Effective bits per weight include element codes and scaling metadata. W4A4 denotes 4-bit weights and activations; W4A16 denotes 4-bit weights with high-precision activations.
Figure 3: Training loss, training top-1 agreement, evaluation KL on held-out data, and evaluation top-1 agreement for Qwen3.5 9B and DeepSeek R1 Distill Llama 8B. We compare weight-only (W4A16) and weight–activation (W4A4) training under NVFP4 and MXFP4.
Figure 4: Training and held-out evaluation loss for Qwen3 8B and Qwen3.5 9B under Q4_K quantization, and for Qwen3-30B-A3B and MiniMax M2.7 under NVFP4 quantization.
Category
Full-param QAD
LoRA r=16
Δ
Teacher weights (BF16)
16.68
0.00
− 16.68
Student weights, frozen (BF16)
3.79
16.68
+ 12.89
Student weights, trainable (FP32)
25.77
0.16
− 25.61
Gradients (FP32)
25.77
0.16
− 25.61
Optimizer states (FP32)
51.55
0.32
− 51.22
Persistent subtotal
123.56
17.32
− 106.24
Table 4: Training memory breakdown (GiB) measured for Qwen3.5 9B with an 8K sequence length, comparing full-parameter QAD with LoRA QAD ( r=16 ) under the same hardware configuration. Δ is the LoRA value minus the full-parameter QAD value. LoRA removes the separate teacher copy and confines trainable parameters, gradients, and optimizer states to the adapters. In this configuration, LoRA reduces GPU memory usage by roughly 2.9× .
Figure 5: Training loss, training top-1 agreement, held-out evaluation loss, and held-out top-1 agreement for LoRA QAD at ranks 4, 16, 64, and 256, compared with full-parameter QAD on Qwen3.5 9B.
Figure 6: LoRA QAD of Muse-Glimmer 30B at ranks 32 and 64: training loss, training top-1 agreement, forward KL divergence to the teacher, and top-5 Jaccard overlap. Faint lines show per-step values, and solid lines show their exponential moving averages.
Figure 7: Optimization and held-out fidelity for token-matched Qwen3.5-9B NVFP4 QAD. Rows denote the Math and Code source domains. Columns report the training objective, forward KL from the BF16 teacher, top-1 agreement, and top-10 Jaccard overlap with the teacher. The x-axis reports training progress as a fraction of each run’s final supervised-token budget; thin training traces are raw logs and thick traces are EMA-smoothed.
Figure 8: Training loss, rollout reward, and completion length over 1,024 optimizer steps for BF16 GRPO and NVFP4 QARL on Qwen3-8B-Base. NVFP4 rollouts use real quantized weights in vLLM, while the training forward pass uses simulated quantization.
The emergence of fine-grained numerical formats like NVFP4 presents new opportunities for efficient Large Language Model (LLM) inference. However, it is difficult to adapt existing Post-Training Quantization (PTQ) strategies to these formats: rotation-based methods compromise fine-grained block isolation; smoothing techniques struggle with significant 4-bit quantization errors; and mixed-precision approaches often conflict with hardware constraints on unified-precision computation. To address these challenges, we propose ARCQuant, a framework that boosts NVFP4 performance via Augmented Residual Channels. Distinct from methods that compromise block isolation or hardware uniformity, ARCQuant maintains a strictly unified NVFP4 format by augmenting the activation matrix with quantized residual channels. This design integrates the error compensation process directly into the matrix reduction dimension, enabling the use of standard, highly optimized GEMM kernels with minimal overhead. Theoretical analysis confirms that the worst-case error bound of our dual-stage NVFP4 quantization is comparable to that of standard 8-bit formats such as MXFP8. Extensive experiments on LLaMA and Qwen models demonstrate that ARCQuant achieves state-of-the-art accuracy, comparable to full-precision baselines in perplexity and downstream tasks. Furthermore, deployment on RTX 5090 and RTX PRO 6000 GPUs confirms practical benefits, achieving up to 3x speedup over FP16. Our code is available at https://github.com/actypedef/ARCQuant.
Haoqian Meng, Yilun Luo, Yafei Zhao +3
School of Computer Science and Technology, Tianjin University
Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, while compact weights reduce memory traffic during generation. We propose "disaggregated quantization" (DQ), which specializes computation formats, weights and storage placement to both of these phases. On Qwen 3 and Gemma 3, removing activation quantization specifically on decode improves accuracy on decode-heavy tasks without increasing inference cost. Training separate compute-native prefill weights accelerates prompt processing relative to weight-only inference while matching or exceeding its accuracy at 2-3-bit decode on both decode-heavy and prefill-heavy tasks. With released Qwen3.8-27B GGUF decoders, training an NVFP4 prefiller improves 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 on MMMU-Pro without modifying the decode checkpoint. To accommodate the additional checkpoint on a single device, offloaded disaggregated prefill (ODP) streams its weights from SSD, amortizing loading over prompt length. On the same 27B model, ODP delivers a 1.78x time-to-first-token speedup over the weight-only baseline at 8K prompt length in llama.cpp. We evaluate accuracy under disaggregated serving in vLLM and further validate shared-weight format disaggregation through post-training quantization on models up to 2.8T parameters.
Andrei Panferov, Maximilian Kleinegger, Sweta Priyadarshi +2
Deploying large language models (LLMs) in resource-constrained environments is hindered by heavy computational and memory requirements. We present LBLLM, a lightweight binarization framework that achieves effective W(1+1)A4 quantization through a novel three-stage quantization strategy. The framework proceeds as follows: (1) initialize a high-quality quantized model via PTQ; (2) quantize binarized weights, group-wise bitmaps, and quantization parameters through layer-wise distillation while keeping activations in full precision; and (3) training learnable activation quantization factors to dynamically quantize activations to 4 bits. This decoupled design mitigates interference between weight and activation quantization, yielding greater training stability and better inference accuracy. LBLLM, trained only using 0.016B tokens with a single GPU, surpasses existing state-of-the-art binarization methods on W2A4 quantization settings across tasks of language modeling, commonsense QA, and language understanding. These results demonstrate that extreme low-bit quantization of LLMs can be both practical and highly effective without introducing any extra high-precision channels or rotational matrices commonly used in recent PTQ-based works, offering a promising path toward efficient LLM deployment in resource-limited situations.
Siqing Song, Chuang Wang, Yong Lang +2
MAIS, Institute of Automation, Chinese Academy of Sciences, China · School of Artificial Intelligence, University of Chinese Academy of Sciences, China · Central Media Technology Institute, Huawei