Post-training quantization (PTQ) enables efficient deployment, yet transformer architectures remain challenging to quantize due to nonlinear layers. While existing methods attribute accuracy loss to insufficient numerical precision, often necessitating floating-point fallbacks, we demonstrate that degradation is actually driven by specific structural error sources. We find that learned scale parameters in normalization layers and compounded approximations in GELU are the primary error contributors, whereas SoftMax remains inherently robust to aggressive quantization. To address these bottlenecks, we introduce TR-PTQ, a unified integer-only formulation using shared Taylor Region (TR) exponential and logarithm primitives. This approach allows computationally expensive operations, including division and square roots, to be performed entirely in the log-domain via standard integer arithmetic. Combined with a calibration-free, outlier-aware optimization for LayerNorm parameters, our method eliminates the need for floating-point hardware units for nonlinearities, achieving less than 1.5% absolute accuracy degradation across vision and language benchmarks.
Figures & tables
Figure 1 : Taylor Regions for ex . Vertical dotted lines denote anchor points; dashed curves show local Taylor approximations centered at each anchor.
Figure 2 : Block-level overview of TR- ln , TR- exp , and TR-SoftMax. The same exp and log primitives are reused across attention, normalization, and activation functions to enable a unified and hardware-efficient implementation.
Figure 3 : Block-level overview of the proposed TR-Norm (top) and TR-GELU (bottom) operators, illustrating their integration within the low-precision nonlinear processing pipeline.
Method
DeiT-T
DeiT-S
DeiT-B
Swin-T
Swin-S
Swin-B
Full FP32 (baseline)
72.21
79.85
81.85
81.35
83.20
83.60
FQ-ViT ( Lin et al., 2021 )
71.61
79.17
81.20
80.51
82.71
82.97
Zhang ( Zhang et al., 2023 )
71.08
78.49
80.74
80.03
82.29
82.67
SOLE (INT8)
71.07
78.89
81.12
80.14
82.60
82.79
QUARK ( Zhao et al., 2025 )
71.29
79.40
81.52
81.06
82.81
85.03 †
TR-PTQ (Ours)
71.39
79.27
81.40
80.70
82.80
83.18
Table 1 : Comparison of Top-1 accuracy (%) across different ViT models and quantization methods.
Figure 4 : Approximation error of the proposed TR nonlinear operators (top: TR- exp , bottom: TR- ln ) under 8-bit integer arithmetic.
Method
DeiT-T
DeiT-S
DeiT-B
Swin-T
Swin-S
Swin-B
FP32 (baseline)
72.21
79.85
81.85
81.35
83.20
83.60
4-bit (rounding)
68.11
77.21
80.53
79.23
80.81
81.28
6-bit (rounding)
70.64
78.61
81.29
80.28
82.67
82.97
8-bit (rounding)
70.84
78.50
81.21
80.34
82.73
82.98
8-bit (1st order)
71.39
79.27
81.40
80.70
82.80
83.18
Table 2 : Model Accuracy Across Different Bit-Precisions
Configuration
Accuracy (%)
FP32 baseline (all layers)
72.21
Integer GELU ( α=1.702 )
71.07
Integer GELU (segmented αi )
72.13
Table 3 : Impact of GELU implementation on model accuracy. Only the GELU nonlinearity is quantized; all other operations remain in floating point.
Config
CoLA
QNLI
RTE
MNLI
QQP
SST-2
MRPC
FP32 Base
53.38
91.54
72.56
84.57
90.91
92.89
89.81
All Quant
51.27
88.30
70.76
81.94
90.07
92.78
88.43
Piecewise
54.80
88.30
70.76
82.75
90.79
92.88
89.84
Excl Linear
55.91
90.99
71.48
83.48
90.79
93.23
90.97
Norm (No Opt)
2.95
65.17
55.23
53.48
70.55
86.01
38.75
Norm (Ours)
53.18
91.05
71.48
83.65
90.82
92.63
90.01
Table 4 : Performance comparison (%) across GLUE tasks under different quantization configurations.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A.1 : Comparison of division noise between the proposed method and a low-cost comparable solution (SOLE). The plot reports the absolute error with respect to a high-precision reference across the evaluated input range.
Figure A.2 : Division noise comparison between the proposed method and a low-cost comparable solution (SOLE) across different input ranges and bit-precisions. Noise is measured as absolute error with respect to a high-precision reference.
Model
DeiT-T
DeiT-S
DeiT-B
Swin-T
Swin-S
Swin-B
Acc. (%)
70.71
78.67
81.03
78.48
82.60
82.79
Appendix
Table B.1: Vision model performance (Top-1 Acc. %) under 4-bit weight and 12-bit activation quantization.
Figure B.1 : Effect of relaxation on LayerNorm γ parameters under min-max quantization. The top row shows the original γ matrices before quantization and the corresponding 12-bit min-max quantized result, where scale distortion is observed. The bottom row illustrates the same matrices after relaxation, prior to quantization: the relaxed parameters remain closer to the original, and even under a 4-bit representation exhibit improved stability compared to the unrelaxed 12-bit case.