Pretraining Transformers with Quantized Softmax in Attention
Organizations: University of Illinois Urbana-Champaign
Abstract
Low-precision Transformer systems increasingly quantize attention matrix multiplications, while softmax often remains at higher precision. During pretraining, an approximate softmax changes the gradients that train the model as well as its forward computation. We study this interaction with K-interval attention, which approximates the exponential using K+1 grid values. We vary per-row grid calibration, interpolation versus hard rounding, and the placement of a straight-through surrogate relative to normalization. We derive the corresponding backward rules, including calibration derivatives, and compare these choices in pretraining experiments matched on model, data, and optimizer. Detaching the row extrema leaves the forward computation unchanged but produces a delayed increase in validation loss. With hard rounding at K=4, min-max calibration and a pre-normalization surrogate incur a large loss gap; changing either choice substantially reduces it. At 124M parameters and 2.5B training tokens, fixed-window calibration with a post-normalization surrogate yields a validation loss gap of +0.019 nats relative to softmax at K=4, and with a pre-normalization surrogate yields +0.004 nats at K=16.
Figures & tables
| Name | Calibration | Reconstr. | Backward |
|---|---|---|---|
| softmax | — | exp | autograd |
| LERP | MinMax | LERP | exact Jacobian ( 2 ) |
| MinMax–Weight | MinMax | Nearest | Weight-STE ( 3 ) |
| MinMax–Prob | MinMax | Nearest | Prob-STE ( 4 ) |
| FWM–Weight | FWM, | Nearest | Weight-STE |
| FWM–Prob | FWM, | Nearest | Prob-STE |
| Suite | Purpose | Model | Tokens | Corpus | Seeds | Runs |
|---|---|---|---|---|---|---|
| A long-horizon | , horizon, downstream | 124M | 2.5B | FineWebEdu-3B | 1 | 15 |
| B factorial | calibration surrogate | 124M | 250M | FineWebEdu-3B | 1 | 10 |
| C scale probe | ordering at parameters | 1B | 100M | WikiText-103 | 1 | 12 |
| D seeds | grid, seed spread | 124M | 100M | WikiText-103 | 5 | 130 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Condition | Forward | Backward |
|---|---|---|
| Degenerate MinMax row ( ): single valid key or equal valid scores | weights set to 1 ( torch.where ), invalid keys 0, uniform (softmax’s forward) | zero gradient through where , extremum gradients cancel; finite |
| Very small positive span ( ) | regular branch, tiny , normal reconstruction | all gradients finite |
| FWM: single key or equal scores | for every valid key, all at the top grid value, uniform; no degenerate branch | regular formulas |
| Tied maxima/minima, positive span | unique-extremum formulas | amax / amin split the extremum gradient evenly among ties |
| Grid-boundary index switch | is a floor, piecewise constant | zero derivative a.e.; LERP weight continuous, secant identities on either side |
| Nearest midpoint tie | rounds up; hard weight discontinuous | ( 5 ) applies away from these points |
| Run (suite: condition) | mean | max |
|---|---|---|
| B: MinMax–Prob K4 | ||
| B: MinMax–Prob K16 | ||
| B: FWM–Prob K4 | ||
| B: FWM–Prob K16 | ||
| A: MinMax–Prob K4 | ||
| A: FWM–Prob K4 |
| Run (suite: condition) | mean | max |
|---|---|---|
| B: MinMax–Prob K4 | ||
| B: MinMax–Prob K16 | ||
| B: FWM–Prob K4 | ||
| B: FWM–Prob K16 | ||
| A: MinMax–Prob K4 | ||
| A: FWM–Prob K4 |
| Family | slope | 95% -CI | seed SD |
|---|---|---|---|
| LERP | -0.0222 | [-0.024, -0.020] | 0.0014 |
| MinMax–Weight | -0.0472 | [-0.051, -0.043] | 0.0033 |
| MinMax–Prob | -0.0301 | [-0.034, -0.027] | 0.0028 |
| FWM–Weight | -0.0078 | [-0.010, -0.005] | 0.0020 |
| FWM–Prob | -0.0030 | [-0.007, +0.001] | 0.0029 |
| Family | slope | 95% -CI | seed SD |
|---|---|---|---|
| LERP | -0.0222 | [-0.024, -0.020] | 0.0014 |
| MinMax–Weight | -0.0472 | [-0.051, -0.043] | 0.0033 |
| MinMax–Prob | -0.0301 | [-0.034, -0.027] | 0.0028 |
| FWM–Weight | -0.0078 | [-0.010, -0.005] | 0.0020 |
| FWM–Prob | -0.0030 | [-0.007, +0.001] | 0.0029 |
| Condition | BLiMP | HellaSwag | PIQA | ARC-E | LAMBADA | SciQ | SciQ-ns |
|---|---|---|---|---|---|---|---|
| softmax (absolute, %) | 79.5 | 30.0 | 59.7 | 45.2 | 21.6 | 73.1 | 51.2 |
| MinMax–Weight K4 | -19.4 ∗ | -3.9 ∗ | -4.1 ∗ | -9.6 ∗ | -18.4 ∗ | -19.3 ∗ | -17.6 ∗ |
| MinMax–Weight K16 | -9.5 ∗ | -2.7 ∗ | -2.2 ∘ | -5.7 ∗ | -11.6 ∗ | -9.3 ∗ | -8.7 ∗ |
| LERP K1 | -5.2 ∗ | -1.9 ∗ | -0.2 | -3.9 ∗ | -12.9 ∗ | -8.5 ∗ | -5.2 ∗ |
| MinMax–Weight K32 | -3.8 ∗ | -1.6 ∗ | -0.4 | -3.5 ∗ | -4.4 ∗ | -2.2 | -1.4 |
| FWM–Weight K2 | -5.4 ∗ | -1.1 ∗ | -1.0 | -2.6 ∗ | -3.6 ∗ | -1.2 | -1.4 |
| Diagnostic | Question | Result |
|---|---|---|
| A1 forward shift invariance | , softmax, LERP, Nearest, | Pass for all. |
| A2 gradient shift residual | per row | softmax ; LERP full ; LERP detach median 0.026–0.410, max near 1. |
| A3 all-ones directional finite difference (fp64) | autograd vs derivative along | Detach : autograd vs true 0. Full: . |
| A4 interior finite difference | local LERP slopes away from extrema | softmax, detach and full all match. |
| P1-0 projection sanity | removes exactly the common-mode gradient? | Forward bitwise equal to detach; projected ; . |
| Condition | Backward | Span | Entropy | Top-level mass | Eff. levels | TV | ||
|---|---|---|---|---|---|---|---|---|
| Nearest K2 | detach | 226.2 | 113.1 | 58.0 | 3.97 | 0.679 | 1.81 | 0.439 |
| Nearest K4 | full | 19.8 | 4.94 | 10.9 | 3.77 | 0.637 | 2.31 | 0.355 |
| Nearest K4 | detach | 903.8 | 225.9 | 88.4 | 0.66 | 0.990 | 1.06 | 0.873 |
| Nearest K8 | full | 25.5 | 3.18 | 11.3 | 3.86 | 0.506 | 3.57 | 0.259 |
| Nearest K8 | detach | 1544.7 | 193.1 | 72.8 | 1.03 | 0.927 | 1.47 | 0.702 |
| Nearest K16 | full | 29.6 | 1.85 | 12.2 | 3.78 | 0.386 | 6.70 | 0.160 |