JustQuant: You Don't Need Smoothing, SVD, or Rotation for 4-Bit Activation Quantization
Organizations: Shanghai Jiao Tong University · Tsinghua University · The Hong Kong Polytechnic University
Abstract
Recent generative models have become increasingly powerful, but their inference cost continues to grow. Model quantization offers a promising way to compress these models and accelerate inference. However, at 4 bits, activation quantization is substantially more challenging than weight quantization. Recent post-training quantization (PTQ) and quantization-aware training (QAT) methods have made progress in 4-bit activation quantization by introducing smoothing, SVD branches, rotations, mixed precision, or advanced formats such as NVFP4. These additional operators and data types impose demanding requirements on inference engines and hardware, limiting the broad adoption of low-precision models. Can quantization be achieved using only plain low-bit operators? To answer this question, we propose JustQuant, a simple yet effective framework that moves the complexity of low-bit quantization from deployment-time operators into the training process. We first revisit model quantization from the perspective of knowledge distillation and show that a key reason existing PTQ and QAT methods fail is that they typically exploit supervision at only a single level. We then introduce Theseus QAD, a quantization-aware distillation method that progressively applies multi-level supervision, analogous to the gradual replacement process in the Ship of Theseus. Extensive experiments on DiT and diffusion large language models show two distinct regimes. For smaller models, Theseus QAD can serve as a lightweight warm-up stage that substantially improves subsequent QAT with plain operators, while naive QAD may collapse in the same setting. For larger models, Theseus QAD provides a stronger distillation training path than ordinary QAD. Across both regimes, JustQuant improves low-bit quantization quality while avoiding the complex operators required by many existing PTQ methods.
Figures & tables
| Small model | Large model | |||
|---|---|---|---|---|
| Method | Quality | Deploy. | Quality | Deploy. |
| PTQ | ✓ | ✗ | ✓ | ✗ |
| QAT | ✓ | ✗ | ✗ | ✓ |
| QAD | ✗ | ✗ | ✓ | ✓ |
| JustQuant | ✓ | ✓ | ✓ | ✓ |
| 1: | |
|---|---|
| 2: | (see text) |
| 3: | for and do |
| 4: | , |
| 5: | for do |
| 6: | |
| 7: |
| Small model | Large model | |
|---|---|---|
| DiT | DiT-XL/2, 0.6B ( Peebles and Xie, 2023 ) | FLUX.1, 12B ( Black Forest Labs, 2024 ) |
| DLLM | ELF-B, 105M ( Hu et al., 2026 ) | LLaDA, 8B ( Nie et al., 2025 ) |
| Setting | Method | W/A | Type/Step | FID | sFID | IS | Additional Operator |
| 10K samples 50 steps | FP ( Peebles and Xie, 2023 ) | 32/32 | -/7000k | 6.78 | 20.56 | 243.70 | – |
| SVDQuant ( Li et al., 2025 ) | 4/4 | PTQ/0k | 81.28 | 67.27 | 29.17 | Smooth, SVD | |
| ConvRot ( Huang et al., 2025a ) | 4/4 | PTQ/0k | 18.69 | 32.06 | 139.54 | Rotation | |
| QuaRot ( Ashkboos et al., 2024 ) | 4/4 | PTQ/0k | 53.31 | 56.74 | 53.12 | Rotation | |
| VETA-DiT ( Xu et al., 2025 ) | 4/4 | PTQ/14.4k | 9.87 | 25.52 | 202.55 | Smooth, Rotation | |
| TreeQ ( Yang et al., 2025 ) | 4/4 | PTQ/36.9k | 6.92 | 20.86 | 219.66 | MP, SVD, Rotation |
| W/A | Arm | Qwen3-8B-Base | GPT-2-Large | Add. Op. | ||
|---|---|---|---|---|---|---|
| gen-PPL (4seed) | PPL vs FP | gen-PPL (4seed) | PPL vs FP | |||
| 4/4 | FP | – | – | – | ||
| JustQuant | None | |||||
| RobuQ | Hadamard, SVD | |||||
| Direct QAT | None | |||||
| 1.58/4 | FP | – | – | – | ||
| Method | LR | Steps | FID | sFID | IS |
|---|---|---|---|---|---|
| QAT | 1e-4 | 20k | 24.60 | 28.39 | 90.01 |
| QAT | 1e-6 | 20k | 27.52 | 35.68 | 84.70 |
| QAT | 2e-6 | 20k | 25.77 | 36.19 | 89.09 |
| QAT | 5e-6 | 20k | 25.64 | 34.68 | 85.60 |
| QAT | 1e-5 | 20k | 24.94 | 33.47 | 88.06 |
| QAT | 2e-5 | 20k | 18.27 | 29.78 | 114.82 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| DiT-XL/2, ImageNet | |||
| Setting | W1.58A32 | W4A4 | W1.58A4 |
| Stage 1 (Theseus QAD) | |||
| Dataset | ImageNet, 128k VAE latents | ImageNet, 128k VAE latents | ImageNet, 128k VAE latents |
| Learning rate | |||
| Batch size | 256 | 128 | 256 |
| Training steps | 10k | 10k | 10k |