cs.AISep 27, 2026

JustQuant: You Don't Need Smoothing, SVD, or Rotation for 4-Bit Activation Quantization

Authors: Kaicheng Yang, Kaisen Yang, Chunyu Liu, Xianglong Yan, Haotong Qin, Junyi Wu, Tianao Zhang, Xun Zhang, +3 more

Organizations: Shanghai Jiao Tong University · Tsinghua University · The Hong Kong Polytechnic University

Abstract

Recent generative models have become increasingly powerful, but their inference cost continues to grow. Model quantization offers a promising way to compress these models and accelerate inference. However, at 4 bits, activation quantization is substantially more challenging than weight quantization. Recent post-training quantization (PTQ) and quantization-aware training (QAT) methods have made progress in 4-bit activation quantization by introducing smoothing, SVD branches, rotations, mixed precision, or advanced formats such as NVFP4. These additional operators and data types impose demanding requirements on inference engines and hardware, limiting the broad adoption of low-precision models. Can quantization be achieved using only plain low-bit operators? To answer this question, we propose JustQuant, a simple yet effective framework that moves the complexity of low-bit quantization from deployment-time operators into the training process. We first revisit model quantization from the perspective of knowledge distillation and show that a key reason existing PTQ and QAT methods fail is that they typically exploit supervision at only a single level. We then introduce Theseus QAD, a quantization-aware distillation method that progressively applies multi-level supervision, analogous to the gradual replacement process in the Ship of Theseus. Extensive experiments on DiT and diffusion large language models show two distinct regimes. For smaller models, Theseus QAD can serve as a lightweight warm-up stage that substantially improves subsequent QAT with plain operators, while naive QAD may collapse in the same setting. For larger models, Theseus QAD provides a stronger distillation training path than ordinary QAD. Across both regimes, JustQuant improves low-bit quantization quality while avoiding the complex operators required by many existing PTQ methods.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization

    May 25, 2026Ke Li, Dong An, Xiaoling Zang +6Large Language Model QuantizationDistributions

  2. DynamicPTQ: Mitigating Activation Quantization Collapse via Residual-Stream Dynamics

    Jun 10, 2026Zimo Zhao, Maolin Wang, Bowen Yu +3Post-Training QuantizationMixed-Precision Quantization

  3. AAAC: Activation-Aware Adaptive Codebooks for 4-bit LLM Weight Quantization

    May 9, 2026Beshr IslamBouli, David JinLarge Language Model QuantizationModel Weights