cs.LGSep 27, 2026

Pretraining Transformers with Quantized Softmax in Attention

Authors: Shangzhen Zhu, Muyan Hu, Tomasz Kozlowski

Organizations: University of Illinois Urbana-Champaign

Abstract

Low-precision Transformer systems increasingly quantize attention matrix multiplications, while softmax often remains at higher precision. During pretraining, an approximate softmax changes the gradients that train the model as well as its forward computation. We study this interaction with K-interval attention, which approximates the exponential using K+1 grid values. We vary per-row grid calibration, interpolation versus hard rounding, and the placement of a straight-through surrogate relative to normalization. We derive the corresponding backward rules, including calibration derivatives, and compare these choices in pretraining experiments matched on model, data, and optimizer. Detaching the row extrema leaves the forward computation unchanged but produces a delayed increase in validation loss. With hard rounding at K=4, min-max calibration and a pre-normalization surrogate incur a large loss gap; changing either choice substantially reduces it. At 124M parameters and 2.5B training tokens, fixed-window calibration with a post-normalization surrogate yields a validation loss gap of +0.019 nats relative to softmax at K=4, and with a pre-normalization surrogate yields +0.004 nats at K=16.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. From Attention Sensitivity to Layer Role: Revisiting Mixed-Precision Quantization of Transformers

    Sep 28, 2026Nafiseh HosseinpourFardi, Negar Alihadi, Mahmoudreza Babaei +2Post-Training QuantizationMixed-Precision Quantization

  2. Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration

    Sep 27, 2026Shangzhen Zhu, Muyan Hu, Tomasz KozlowskiGumbel-Softmax RelaxationTransformer Architectures

  3. EFQ-Softmax: Exp-Free Quantization for Softmax

    Sep 9, 2026Haohui Han, Yuming Wan, Hongni Wang +4Gumbel-Softmax RelaxationBlock Sparse Flash Attention