Post-training quantization (PTQ) enables efficient deployment, yet transformer architectures remain challenging to quantize due to nonlinear layers. While existing methods attribute accuracy loss to insufficient numerical precision, often necessitating floating-point fallbacks, we demonstrate that degradation is actually driven by specific structural error sources. We find that learned scale parameters in normalization layers and compounded approximations in GELU are the primary error contributors, whereas SoftMax remains inherently robust to aggressive quantization. To address these bottlenecks, we introduce TR-PTQ, a unified integer-only formulation using shared Taylor Region (TR) exponential and logarithm primitives. This approach allows computationally expensive operations, including division and square roots, to be performed entirely in the log-domain via standard integer arithmetic. Combined with a calibration-free, outlier-aware optimization for LayerNorm parameters, our method eliminates the need for floating-point hardware units for nonlinearities, achieving less than 1.5% absolute accuracy degradation across vision and language benchmarks.
Figures & tables
Figure 1 : Taylor Regions for ex . Vertical dotted lines denote anchor points; dashed curves show local Taylor approximations centered at each anchor.
Figure 2 : Block-level overview of TR- ln , TR- exp , and TR-SoftMax. The same exp and log primitives are reused across attention, normalization, and activation functions to enable a unified and hardware-efficient implementation.
Figure 3 : Block-level overview of the proposed TR-Norm (top) and TR-GELU (bottom) operators, illustrating their integration within the low-precision nonlinear processing pipeline.
Method
DeiT-T
DeiT-S
DeiT-B
Swin-T
Swin-S
Swin-B
Full FP32 (baseline)
72.21
79.85
81.85
81.35
83.20
83.60
FQ-ViT ( Lin et al., 2021 )
71.61
79.17
81.20
80.51
82.71
82.97
Zhang ( Zhang et al., 2023 )
71.08
78.49
80.74
80.03
82.29
82.67
SOLE (INT8)
71.07
78.89
81.12
80.14
82.60
82.79
QUARK ( Zhao et al., 2025 )
71.29
79.40
81.52
81.06
82.81
85.03 †
TR-PTQ (Ours)
71.39
79.27
81.40
80.70
82.80
83.18
Table 1 : Comparison of Top-1 accuracy (%) across different ViT models and quantization methods.
Figure 4 : Approximation error of the proposed TR nonlinear operators (top: TR- exp , bottom: TR- ln ) under 8-bit integer arithmetic.
Method
DeiT-T
DeiT-S
DeiT-B
Swin-T
Swin-S
Swin-B
FP32 (baseline)
72.21
79.85
81.85
81.35
83.20
83.60
4-bit (rounding)
68.11
77.21
80.53
79.23
80.81
81.28
6-bit (rounding)
70.64
78.61
81.29
80.28
82.67
82.97
8-bit (rounding)
70.84
78.50
81.21
80.34
82.73
82.98
8-bit (1st order)
71.39
79.27
81.40
80.70
82.80
83.18
Table 2 : Model Accuracy Across Different Bit-Precisions
Configuration
Accuracy (%)
FP32 baseline (all layers)
72.21
Integer GELU ( α=1.702 )
71.07
Integer GELU (segmented αi )
72.13
Table 3 : Impact of GELU implementation on model accuracy. Only the GELU nonlinearity is quantized; all other operations remain in floating point.
Config
CoLA
QNLI
RTE
MNLI
QQP
SST-2
MRPC
FP32 Base
53.38
91.54
72.56
84.57
90.91
92.89
89.81
All Quant
51.27
88.30
70.76
81.94
90.07
92.78
88.43
Piecewise
54.80
88.30
70.76
82.75
90.79
92.88
89.84
Excl Linear
55.91
90.99
71.48
83.48
90.79
93.23
90.97
Norm (No Opt)
2.95
65.17
55.23
53.48
70.55
86.01
38.75
Norm (Ours)
53.18
91.05
71.48
83.65
90.82
92.63
90.01
Table 4 : Performance comparison (%) across GLUE tasks under different quantization configurations.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A.1 : Comparison of division noise between the proposed method and a low-cost comparable solution (SOLE). The plot reports the absolute error with respect to a high-precision reference across the evaluated input range.
Figure A.2 : Division noise comparison between the proposed method and a low-cost comparable solution (SOLE) across different input ranges and bit-precisions. Noise is measured as absolute error with respect to a high-precision reference.
Model
DeiT-T
DeiT-S
DeiT-B
Swin-T
Swin-S
Swin-B
Acc. (%)
70.71
78.67
81.03
78.48
82.60
82.79
Appendix
Table B.1: Vision model performance (Top-1 Acc. %) under 4-bit weight and 12-bit activation quantization.
Figure B.1 : Effect of relaxation on LayerNorm γ parameters under min-max quantization. The top row shows the original γ matrices before quantization and the corresponding 12-bit min-max quantized result, where scale distortion is observed. The bottom row illustrates the same matrices after relaxation, prior to quantization: the relaxed parameters remain closer to the original, and even under a 4-bit representation exhibit improved stability compared to the unrelaxed 12-bit case.
Many LLM applications require only narrow capabilities, yet standard post-training quantization (PTQ) methods allocate precision without considering the target task. This can waste bits on layers that are less relevant to the task signal while over-compressing layers that are critical for downstream behavior. We propose Task-Aware Quantization (TAQ), a training-free, weight-only mixed-precision PTQ framework that uses a small set of unlabeled task calibration prompts to allocate higher precision to task-relevant transformer layers under a fixed bit budget. TAQ estimates layer importance from hidden representations and output sensitivity, and we instantiate it with three scoring rules: TAQ-IS, based on activation information and stability; TAQ-KL, based on output-distribution sensitivity under a quantization-noise proxy; and TAQ-O, a label-informed oracle diagnostic for analyzing layer sensitivity. Across several benchmarks, TAQ outperforms task-agnostic baselines such in most settings, with especially strong gains in the accuracy--memory ratio. We further validate that these gains translate to real deployment behavior through hardware throughput and latency measurements, and analyze calibration robustness and residual-stream error propagation. Overall, TAQ turns mixed-precision PTQ from a model-centric compression step into a task-conditioned precision-allocation problem. A reference implementation is available at \href{https://anonymous.4open.science/r/TAQ-9217/README.md}{\includegraphics[height=1em]{imgs/github-mark.png}}.
Amit LeVi, Raz Lapid, Rom Himelstein +3
1Technion—Israel Institute of Technology · 2Intuit · 3Ben-Gurion University of the Negev +1
Large language models (LLMs) require substantial compute, and thus energy, at inference time. While quantizing weights and activations is effective at improving efficiency, naive quantization of LLMs can significantly degrade performance due to large magnitude outliers. This paper describes FPTQuant, which introduces three novel, lightweight, and expressive function-preserving transforms (FPTs) to facilitate quantization of transformers: (1) a mergeable pre-RoPE transform for queries and keys, (2) a mergeable transform for values, and (3) a cheap, dynamic per-token scaling transform. By leveraging the equivariances and independencies inherent to canonical transformer operation, we designed these FPTs to maintain the model's function while shaping the intermediate activation distributions to be more quantization friendly. FPTQuant requires no custom kernels and adds virtually no overhead during inference. The FPTs are trained both locally to reduce outliers, and end-to-end such that the outputs of the quantized and full-precision models match. FPTQuant enables static INT4 quantization with minimal overhead and shows SOTA speed-up of up to 3.9X over FP. Empirically, FPTQuant has an excellent accuracy-speed trade-off -- it is performing on par or exceeding most prior work and only shows slightly lower accuracy compared to a method that is up to 29% slower.
Boris van Breugel, Yelysei Bondarenko, Paul Whatmough +1
Post-training quantization (PTQ) is a practical way to reduce the memory footprint of large language models, but low-bit quantization is sensitive to mismatches between the quantization codebook and the empirical weight/activation distributions. We revisit Benford-like leading-digit statistics as a lightweight diagnostic of scale-broad behavior in transformer tensors. Across several model families, we observe a consistent functional dichotomy: transformational nn.Linear weights tend to be Benford-like, whereas LayerNorm parameters systematically deviate. Motivated by this observation, we propose BenQ, a data-free PTQ codebook that uses a simple log-spaced grid as a proxy for scale-broad distributions and applies it selectively to transformational layers while keeping stability-critical parameters in higher precision. In 4-bit group-wise PTQ, BenQ consistently improves over uniform RTN and trades wins with NF4 across architectures and tasks, while remaining substantially simpler than optimization-based methods. We additionally report dynamic activation quantization as an exploratory stress test: the results show that log-spaced grids can reduce RTN failures in some families, but also reveal that outlier handling remains essential for reliable low-bit activation PTQ. Code is available at https://github.com/ufopcsilab/benford-quant.
Arthur Negrão, Pedro Silva, Vander L. S. Freitas +2
Postgraduate Program in Computer Science Federal University of Ouro Preto · Computing Department Federal University of Ouro Preto