AlignQuant: Tile-Aligned Mixed-Precision Quantization for Efficient LLM Generation
Organizations: LLaVi Lab, Computer Science & Engineering, University of North Texas
Abstract
Fine-grained mixed-precision quantization promises efficient large language model inference, but local precision choices can conflict with regular GPU storage and computation units. This precision-boundary mismatch limits the translation of compression into practical acceleration. We introduce AlignQuant, a post-training quantization method that uses GPU-compatible two-dimensional weight tiles as the common unit of precision allocation, compact storage, and execution. This shared partition lets precision follow sensitivity within output channels. Joint prefill/decode calibration scores precision reductions using projection-output perturbations weighted by language-model loss gradients under quantized activations. Phase-normalized scores prioritize higher precision for tiles important to either phase under a model-wide weight-storage budget. Each tile stores one selected representation, while phase-specialized kernels reuse the packed model and expand lower-bit weights for INT8 computation with 8-bit activations. Across four LLMs spanning 3B to 14B parameters, AlignQuant achieves up to generation speedup over BF16 while preserving model quality. Evaluations further cover three GPUs and contexts up to 64K tokens. These results show that local precision flexibility and regular GPU execution can coexist through a shared tile unit. The implementation is available at https://github.com/HanzhiZhang-Ulrica/AlignQuant.
Figures & tables
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Symbol | Description |
|---|---|---|
| Projection | Input activations | |
| Projection weights | ||
| Projection output | ||
| Activation rows, input features, and output features | ||
| Coordinates | Orthogonal output- and input-feature transforms | |
| Transformed coordinates and quantized reconstructions |
| Schedule | Output region | Threads | Arithmetic |
|---|---|---|---|
| Prefill M16 | 128 | Tensor Core INT8 MMA | |
| Prefill M64 | 256 | Tensor Core INT8 MMA | |
| Single-sequence decode | 256 | DP4A dot products |