GLF-Q: Global-Local Feature-based Quantization for Vision Transformers
Organizations: State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing 210023, China · School of Artificial Intelligence, Nanjing University, Nanjing 210023, China · Zhongguancun Academy, Beijing 100094, China
Abstract
Post-training quantization (PTQ) efficiently compresses Vision Transformers (ViTs) without retraining, yet suffers severe accuracy degradation at low bit-widths. Existing optimization-based PTQ methods guide block reconstruction via either soft logits or second-order Hessian proxies. Logit supervision is prone to overfitting on limited calibration data, while Hessian approximations incur structural truncation errors. To address these limitations, we propose \textbf{GLF-Q}, a novel PTQ framework guided by Global-Local Feature alignment. GLF-Q propagates quantized block outputs through downstream full-precision layers to align penultimate-layer representations under local output regularization, providing downstream feature supervision without explicitly approximating the Hessian or using a Taylor expansion. Furthermore, offline Hadamard transformations are introduced with zero runtime overhead to disperse activation outliers across channels, effectively contracting dynamic ranges and reducing quantization errors. Meanwhile, optimizing this loss via a Straight-Through Estimator (STE) achieves rapid convergence, bypassing continuous relaxation rounding formulations such as AdaRound. Extensive experiments across representative ViT architectures demonstrate that GLF-Q with standard uniform quantizers substantially outperforms state-of-the-art methods under 3-bit quantization on image classification. In addition, GLF-Q exhibits strong out-of-domain calibration robustness and achieves speedups under 8-bit GPU deployment.
Figures & tables
| Method | Opt. | Spec. | W/A | ViT-S | ViT-B | DeiT-T | DeiT-S | DeiT-B | Swin-S | Swin-B |
| Full-Prec | - | - | 32/32 | 81.39 | 84.54 | 72.21 | 79.85 | 81.80 | 83.23 | 85.27 |
| PTQ4ViT ( Yuan et al., 2022 ) | 3/3 | 0.10 | 0.10 | 3.50 | 0.10 | 31.06 | 28.69 | 20.13 | ||
| RepQ-ViT ( Li et al., 2023 ) | 3/3 | 0.10 | 0.10 | 0.10 | 0.10 | 0.10 | 0.10 | 0.10 | ||
| AdaLog ( Wu et al., 2024 ) | 3/3 | 13.88 | 37.91 | 31.56 | 24.47 | 57.47 | 64.41 | 69.75 | ||
| I&S-ViT ( Zhong et al., 2023 ) | 3/3 | 45.16 | 63.77 | 41.52 | 55.78 | 73.30 | 74.20 | 69.30 | ||
| DopQ-ViT ( Yang et al., 2024 ) | 3/3 | 54.72 | 65.76 | 44.71 | 59.26 | 74.91 | 74.77 | 69.63 |
| Method | Opt. | Spec. | W/A | Mask R-CNN | Cascade Mask R-CNN | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Swin-T | Swin-S | Swin-T | Swin-S | ||||||||
| Full-Precision | - | - | 32/32 | 46.0 | 41.6 | 48.5 | 43.3 | 50.4 | 43.7 | 51.9 | 45.0 |
| PTQ4ViT ( Yuan et al., 2022 ) | 4/4 | 6.9 | 7.0 | 26.7 | 26.6 | 14.7 | 13.5 | 0.5 | 0.5 | ||
| APQ-ViT ( Ding et al., 2022 ) | 4/4 | 23.7 | 22.6 | 44.7 | 40.1 | 27.2 | 24.4 | 47.7 | 41.1 | ||
| RepQ-ViT ( Li et al., 2023 ) | 4/4 | 36.1 | 36.0 | 44.2 | 40.2 | 47.0 | 41.1 | 49.3 | 43.1 | ||
| Setting | ViT-S | ViT-B | DeiT-T | DeiT-S | DeiT-B | Swin-S | Swin-B |
|---|---|---|---|---|---|---|---|
| LF | 60.48 | 77.06 | 53.16 | 70.16 | 76.76 | 76.66 | 78.46 |
| DPLR-FIM ( Wu et al., 2025a ) | 63.83 | 77.15 | 54.94 | 70.67 | 76.60 | 77.30 | 79.19 |
| Logits + LF ( Liu et al., 2023 ) | 65.16 | 76.89 | 55.26 | 70.71 | 76.87 | 76.80 | 78.58 |
| GF | 66.09 | 77.17 | 54.06 | 70.41 | 76.69 | 77.75 | 80.30 |
| GLF (Ours) | 67.40 | 78.18 | 56.48 | 71.55 | 77.20 | 78.10 | 80.45 |
| FIMA-Q + MR + Had. † | 66.88 | 76.13 | 55.29 | 69.81 | 76.24 | 77.37 | 79.42 |
| Method | ViT-S | ViT-B | DeiT-T | DeiT-S | DeiT-B | Swin-S | Swin-B |
|---|---|---|---|---|---|---|---|
| GLF only | 60.37 | 72.50 | 55.28 | 68.75 | 75.79 | 76.92 | 79.31 |
| GLF + MR | 63.53 | 73.60 | 55.24 | 69.00 | 75.79 | 77.96 | 80.09 |
| GLF + Hadamard | 65.04 | 77.16 | 55.69 | 71.30 | 76.91 | 77.63 | 79.81 |
| GLF-Q | 67.40 | 78.18 | 56.48 | 71.55 | 77.20 | 78.10 | 80.45 |
| Method | 3-bit | 4-bit | ||||
|---|---|---|---|---|---|---|
| ViT-S | DeiT-S | Swin-S | ViT-S | DeiT-S | Swin-S | |
| FIMA-Q ( Wu et al., 2025a ) | 41.86 | 36.03 | 69.95 | 73.37 | 75.89 | 80.72 |
| DPLR-FIM † | 47.80 | 65.58 | 70.77 | 73.88 | 76.02 | 80.51 |
| GLF-Q (Ours) | 58.72 | 68.00 | 74.71 | 75.34 | 76.46 | 81.07 |
| Method | Latency | Top-1 Accuracy | ||||
|---|---|---|---|---|---|---|
| ViT-S | DeiT-S | Swin-S | ViT-S | DeiT-S | Swin-S | |
| FIMA-Q ( Wu et al., 2025a ) | 2.95 | 2.95 | 6.83 | 79.19 | 78.72 | 83.06 |
| GLF-Q (Ours) | 2.50 | 2.49 | 5.88 | 80.35 | 79.07 | 83.08 |
| Method | ViT-S | ViT-B | DeiT-T | DeiT-S | DeiT-B | Swin-S | Swin-B |
|---|---|---|---|---|---|---|---|
| FIMA-Q ( Wu et al., 2025a ) | 48 | 87 | 44 | 49 | 88 | 115 | 135 |
| LS-ViT ( Hwang et al., 2026 ) | 40 | 56 | 41 | 40 | 56 | 101 | 104 |
| GLF-Q (Ours) | 21 | 34 | 19 | 20 | 34 | 57 | 63 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Network Position | Parameter Transformation Rule |
|---|---|
| Input Patch Embedding | |
| Class Token & Positional Embeddings | |
| QKV Projection & MLP First Layer | |
| Attention Output Projection & MLP Second Layer | |
| Final Classification Head |
| Logits + LF | GLF | |||
|---|---|---|---|---|
| Model | Calibration | Validation | Calibration | Validation |
| ViT-S | 86.91 | 65.16 | 87.70 | 67.40 |
| ViT-B | 91.11 | 76.89 | 91.11 | 78.18 |
| DeiT-T | 77.05 | 55.26 | 76.66 | 56.48 |
| DeiT-S | 90.04 | 70.71 | 89.65 | 71.55 |
| DeiT-B | 96.00 | 76.87 | 96.00 | 77.20 |
| W/A | ViT-S | ViT-B | DeiT-T | DeiT-S | DeiT-B | Swin-S | Swin-B |
|---|---|---|---|---|---|---|---|
| 3/3 | |||||||
| 4/4 |
| ViT-S | ViT-B | DeiT-T | DeiT-S | DeiT-B | Swin-S | Swin-B | Mean | |
|---|---|---|---|---|---|---|---|---|
| 1 | 67.53 | 78.29 | 56.31 | 71.47 | 77.33 | 78.22 | 80.49 | 72.81 |
| 2 | 67.40 | 78.18 | 56.48 | 71.55 | 77.20 | 78.10 | 80.45 | 72.77 |
| 4 | 67.48 | 78.18 | 56.61 | 71.59 | 77.22 | 78.14 | 80.35 | 72.80 |