Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration
Organizations: University of Illinois Urbana-Champaign
Abstract
On NVIDIA Blackwell B200, tensor-core throughput outpaces special-function exponential throughput by more than two orders of magnitude, exposing exponential evaluation in fused attention kernels. A pretrained Transformer, however, may not need it evaluated accurately at every element. We characterize what a pretrained model does need by approximating softmax at inference in ten frozen decoder-only models (0.5B-72B). The number of positions the softmax map assigns probability to and within-row resolution can be cut substantially, yet uniform weighting of the same positions is damaging. Where a fixed resolution budget is placed matters as much as its size, with resolution near the row maximum consistently favored. Perturbations matched on scalar distortion produce model-dependent responses of opposite sign. These findings motivate Rowmax-PoT, a coarse logarithmic weight representation anchored at each row maximum, and Rowmax-H15, its hardware specialization in FlashAttention-4. On B200, the patched FP8 attention forward is 12.4% faster at causal 8K and 25.8% faster at non-causal 8K in host-side call-latency measurements; board energy per forward falls by 8.4% at causal 16K. Measured separately on the BF16 kernel path at 2K, Rowmax-H15 increases perplexity by 0.091-0.492% across five models from three families.
Figures & tables
| Degree of freedom | Intervention (others fixed) | Endpoint | Where |
|---|---|---|---|
| Support | top- , mean threshold | § 4.1 , Fig. 1 A | |
| Within-support weighting | uniform vs. softmax, same support | § 4.1 , Fig. 1 B | |
| Intervals ; weight map | full-range grid; exp. vs. linear map | § 4.1 , Fig. 1 C, Fig. S4 | |
| Allocation | , | , | § 4.2 , Fig. 2 |
| Reconstruction | upper edge, nearest, interpolation | § 4.4 | |
| Exponential evaluation | Rowmax-PoT; Rowmax-H15 in FA4 | , latency | § 4.4 , § 5 |
| Fidelity path | Model | NLL [95% CI] | PPL change (%) |
| BF16 kernel, measured | Qwen2.5-1.5B | 0.00121 [0.000862, 0.00156] | 0.121 |
| BF16 kernel, measured | Qwen2.5-72B | 0.00101 [0.000665, 0.00139] | 0.101 |
| BF16 kernel, measured | Mistral-7B-v0.3 | 0.000914 [0.000600, 0.00125] | 0.091 |
| BF16 kernel, measured | Llama-3.1-8B | 0.00246 [0.00206, 0.00289] | 0.246 |
| BF16 kernel, measured | Llama-3.1-70B | 0.00491 [0.00402, 0.00582] | 0.492 |
| BF16 semantic simulation | Qwen2.5-1.5B | 0.00107 [0.000680, 0.00145] | 0.107 |
| Model | Params (B) | PPL inc. (%) | Rowmax-PoT PPL inc. (%) | Rowmax-PoT top-1 agree. (%) | |
|---|---|---|---|---|---|
| Qwen2.5-0.5B | 0.49 | 0.134 | 0.00611 [0.00525, 0.00697] | 0.382 | 95.8 |
| Qwen2.5-1.5B | 1.54 | 0.091 | 0.00351 [0.00282, 0.00420] | 0.310 | 96.2 |
| Qwen2.5-3B | 3.09 | 0.082 | 0.00331 [0.00256, 0.00404] | 0.289 | 96.4 |
| Qwen2.5-72B | 72.71 | 0.055 | 0.00196 [0.00137, 0.00258] | 0.252 | 97.6 |
| Llama-3.2-1B | 1.24 | 0.223 | 0.0109 [0.00976, 0.0120] | 0.950 | 94.4 |
| Llama-3.2-3B | 3.21 | 0.121 | 0.00629 [0.00540, 0.00714] | 0.704 | 95.2 |
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
| change | reduction in NLL [95% CI] |
|---|---|
| anchor: absolute rowmax, octave spacing | 0.00355 [0.00268, 0.00448] |
| anchor: absolute rowmax, half-octave spacing | 0.00122 [0.000708, 0.00171] |
| spacing: octave half-octave, absolute anchor | 0.00443 [0.00367, 0.00517] |
| spacing: octave half-octave, rowmax anchor | 0.00210 [0.00132, 0.00286] |
| Measurement (semantics) | NLL [95% CI] | PPL change | Evidence status |
|---|---|---|---|
| B200 patched FlashAttention-4 kernel (A1 A0) (BF16 path, rescale_threshold = 8.0 ) | +0.00121 [0.000862, 0.00156] | +0.121% | direct kernel measurement |
| Semantic simulation of the BF16 kernel path (BF16 semantics, threshold 8.0) | +0.00107 [+0.000680, +0.00145] | +0.107% | cross-check of the simulator against the measured kernel |
| Semantic simulation of the FP8 kernel path (FP8 semantics, per-tile running max (tile 128), threshold 0) | +0.000716 [+0.000370, +0.00106] | +0.072% | semantic simulation — not a direct FP8 kernel measurement |
| Offline whole-row Rowmax-H15 (control) (rowmax anchor, no tiling) | +0.00105 [+0.000712, +0.00139] | +0.105% | semantic control |
| Absolute-anchored lattice (control; condition ilpot_h15_v1 ) (absolute anchor) | +0.00224 [+0.00179, +0.00266] | +0.224% | semantic control (phase-misaligned) |
| PoT m = 2 (half-octave reference) (rowmax anchor) | +0.000927 [+0.000543, +0.00132] | +0.093% | reference |
| Contrast | ( NLL) [95% CI] | Reading |
|---|---|---|
| FP8 semantic Rowmax-H15 PoT m = 2 | -0.000211 [-0.000612, +0.000183] | No resolved difference at the reported precision; the paired 95% interval includes zero. |
| FP8 semantic Rowmax-H15 ILPoT-H-A | -0.000143 [-0.000590, +0.000304] | No resolved difference at the reported precision; the paired 95% interval includes zero. |
| FP8 semantic Rowmax-H15 offline whole-row Rowmax-H15 | -0.000333 [-0.000648, -0.0000278] | per-tile running-max semantics differ from whole-row rowmax |
| BF16 semantic Rowmax-H15 FP8 semantic Rowmax-H15 | +0.000358 [-0.00000769, +0.000736] | Higher point estimate under BF16 conditional-rescale semantics; the paired 95% interval includes zero. |
| Simulation vs measured kernel (BF16 path) | measured +0.00121 vs simulated +0.00107 | relative gap 11%; intervals overlap |
| Offline operator | NLL vs softmax [95% CI] | PPL change | Paired NLL vs Rowmax-H15 [95% CI] |
|---|---|---|---|
| Rowmax-H15, | +0.00105 [+0.000712, +0.00139] | +0.105% | reference |
| Rowmax-PoT (Eq. ( 18 )) | +0.000927 [+0.000543, +0.00132] | +0.093% | [ , +0.000301] |
| Unshifted Schraudolph | [ , +0.000270] | [ , ] | |
| , | +0.000329 [+0.0000270, +0.000624] | +0.033% | [ , ] |
| , | +0.0000796 [ , +0.000316] | +0.00796% | [ , ] |
| Run group | Measurements | GPU used | Driver | PyTorch / CUDA | Clock and power record |
|---|---|---|---|---|---|
| Characterization, core | Section 4 interventions on the core models; anchor controls and semantic simulations (Qwen2.5-1.5B) | one GPU, model not recorded | not recorded | 2.5.1+cu121 / 12.1 | not recorded |
| Characterization, large | Section 4 interventions on Llama-3.1-70B and Qwen2.5-72B | three A100-SXM4-80GB | 580.173.02 | 2.5.1+cu121 / 12.1 | not recorded |
| B200 S1 | patch and gates; Qwen2.5-1.5B BF16 fidelity; main matrix; sequence sweep and 1K/2K retest; FP8 non-causal control; head, batch and head-dimension-64 controls; five tuning settings and Rowmax-H15; energy loop; cuDNN and SDPA references; ncu refused | one B200 | 580.126.09 | 2.14.0+cu130 / 13.0 | lock not permitted; clock, temperature and power after each window; energy loop sampled |
| B200 S2 | gates re-run with element dumps; GQA (three rounds); ncu refused | one B200 | 580.126.09 | 2.14.0+cu130 / 13.0 | after each window (GQA at 1965 MHz throughout) |
| B200 S3 | Nsight Systems timeline (kernel duration, call latency, launch cost); ncu refused | one B200 | 595.91.07 | 2.14.0+cu130 / 13.0 a | none |
| B200 S4 | FP16 causal and FP16, BF16 non-causal (three rounds); three tuning settings | one B200 | 580.126.09 | 2.14.0 / 13.0 b | after each window; 1485–1965 MHz and 258–546 W in the three-round cells |
| seq | mask | kernel ( s) A0 / A1 | call latency ( s) A0 / A1 | kernel-only | call latency | non-kernel % (A0 / A1) | ||
|---|---|---|---|---|---|---|---|---|
| 1K | causal | 17.840 | 16.672 | 45.611 | 42.529 | +7.0% | +7.2% | 60.9 / 60.8 |
| 2K | causal | 30.688 | 28.000 | 42.916 | 45.027 | +9.6% | -4.7% | 28.5 / 37.8 |
| 4K | causal | 57.280 | 52.256 | 58.495 | 52.803 | +9.6% | +10.8% | 2.1 / 1.0 |
| 8K | causal | 200.640 | 178.368 | 201.617 | 179.274 | +12.5% | +12.5% | 0.5 / 0.5 |
| 16K | causal | 725.920 | 653.296 | 726.815 | 655.539 | +11.1% | +10.9% | 0.1 / 0.3 |
| 32K | causal | 2747.570 | 2508.674 | 2799.928 | 2548.133 | +9.5% | +9.9% | 1.9 / 1.5 |
| dtype | seq | A0 ms (pooled median) | A1 ms | speed-up (pooled) | per-round sd | rounds |
|---|---|---|---|---|---|---|
| FP8 E4M3 | 4K | 0.0621 | 0.0580 | +7.1% | 2.13 | 3 |
| FP8 E4M3 | 8K | 0.2063 | 0.1835 | +12.4% | 0.02 | 3 |
| FP8 E4M3 | 16K | 0.7320 | 0.6509 | +12.5% | 0.08 | 3 |
| BF16 | 4K | 0.0586 | 0.0564 | +3.8% | 0.93 | 3 |
| BF16 | 8K | 0.2020 | 0.1917 | +5.4% | 0.17 | 3 |
| BF16 | 16K | 0.7871 | 0.7350 | +7.1% | 0.52 | 3 |
| seq | pooled speed-up | per-round min … max | sd | 1K/2K retest (n = 9) |
|---|---|---|---|---|
| 1K | -3.6% | -5.5% … +1.1% | 2.97 | -2.2% |
| 2K | -2.6% | -6.0% … +4.5% | 4.53 | -3.7% |
| 4K | +7.9% | +7.7% … +8.3% | 0.25 | — |
| 8K | +12.1% | +12.0% … +12.1% | 0.06 | — |
| 16K | +12.5% | +12.1% … +12.7% | 0.28 | — |
| 32K | +9.9% | +9.9% … +10.1% | 0.08 | — |
| dtype | mask | seq | speed-up (pooled) | rounds |
|---|---|---|---|---|
| FP8 | non-causal | 8K | +25.8% (sd 0.03) | 3 |
| BF16 | non-causal | 8K | +13.0% (sd 0.50) | 3 |
| FP16 | non-causal | 8K | +11.8% (sd 0.70) | 3 |
| FP8 | non-causal | 4K | +27.2% | 1 |
| FP8 | non-causal | 16K | +25.4% | 1 |
| BF16 | non-causal | 4K | +11.4% | 1 |
| A0 configuration | emulation fraction | vs stock |
|---|---|---|
| freq 0 (hardware MUFU only) | 0% | +1.4% |
| freq 16, start 1 | 12.5% | +1.6% |
| freq 8, start 1 (stock) | 25% | +0.0% |
| freq 8, start 0 | 37.5% | +2.7% |
| freq 2, start 1 | 50% | +0.2% |
| freq 4, start 1 | 50% | -1.2% |
| configuration | speed-up |
|---|---|
| 8K, heads 4 | +12.6% |
| 8K, heads 8 | +12.3% |
| 8K, heads 32 | +12.1% |
| 8K, batch 4 | +11.7% |
| 8K, GQA (32 query / 8 KV heads): 0.38885 ms 0.34583 ms | +12.4% (per-round sd 0.03) |
| head dim 64, FP8: 4K / 8K / 16K | +2.0% / +2.5% / +3.7% |
| kernel | forwards | mean power (W) | time per forward (ms) | energy per forward (mJ) |
|---|---|---|---|---|
| A0 stock FA4 | 59550 | 982.3 | 0.7559 | 742.5 |
| A1 FA4 + Rowmax-H15 | 65050 | 982.9 | 0.6922 | 680.3 |
| seq | torch SDPA flash | torch SDPA mem-efficient | cuDNN 9.24 | FA4 stock (A0) | FA4 + Rowmax-H15 (A1) | Rowmax-H15 vs cuDNN | FA4 vs cuDNN |
|---|---|---|---|---|---|---|---|
| 4K | 0.2876 | 0.5596 | 0.0536 | 0.0586 | 0.0564 | -5.1% | -8.6% |
| 8K | 0.9117 | 1.9075 | 0.2230 | 0.2020 | 0.1917 | +16.4% | +10.4% |
| 16K | 3.1568 | 6.9877 | 0.8316 | 0.7871 | 0.7350 | +13.1% | +5.6% |
| Model | Context | Blocks | Predictions | achieved (rel. err. %) | achieved (rel. err. %) | |
|---|---|---|---|---|---|---|
| Qwen2.5-72B | 2K | 97 | 198,559 | 0.001228 | 0.001229 (0.084) | 0.001235 (0.591) |
| Qwen2.5-72B | 8K | 24 | 196,584 | 0.001323 | 0.001315 (0.614) | 0.001333 (0.725) |
| Qwen2.5-72B | 16K | 12 | 196,596 | 0.001362 | 0.001371 (0.629) | 0.001353 (0.666) |
| Llama-3.1-70B | 2K | 97 | 198,559 | 0.001778 | 0.001787 (0.510) | 0.001785 (0.393) |
| Llama-3.1-70B | 8K | 24 | 196,584 | 0.001137 | 0.001129 (0.687) | 0.001138 (0.100) |
| Llama-3.1-70B | 16K | 12 | 196,596 | 0.001236 | 0.001244 (0.621) | 0.001237 (0.062) |
| Condition | LAMBADA | HellaSwag | ARC-E | ARC-C | PIQA | WinoGrande | LAMBADA PPL (%) |
|---|---|---|---|---|---|---|---|
| Qwen2.5-72B | |||||||
| Softmax | 77.72 (+0.00) | 86.10 (+0.00) | 83.38 (+0.00) | 62.71 (+0.00) | 83.62 (+0.00) | 77.35 (+0.00) | — |
| Mean-threshold support | 76.73 ( 0.99) | 85.72 ( 0.38) | 84.81 (+1.43) | 62.71 (+0.00) | 81.83 ( 1.80) | 76.16 ( 1.18) | +6.727 |
| Exp. | 77.59 ( 0.14) | 86.13 (+0.03) | 83.25 ( 0.13) | 62.54 ( 0.17) | 83.62 (+0.00) | 77.35 (+0.00) | +0.037 |
| Exp. | 77.64 ( 0.08) | 86.05 ( 0.05) | 83.21 ( 0.17) | 62.63 ( 0.09) | 83.68 (+0.05) | 78.30 (+0.95) | +0.033 |
| Rowmax-PoT | 77.53 ( 0.19) | 86.06 ( 0.04) | 83.46 (+0.08) | 62.20 ( 0.51) | 83.73 (+0.11) | 77.74 (+0.39) | +0.231 |
| model | params (B) | (21) [95% CI] | resolved? | (16) | (32) | [95% CI] |
|---|---|---|---|---|---|---|
| Qwen2.5-0.5B | 0.49 | +0.00611 [+0.00525, +0.00697] | yes | +0.0105 | +0.00192 | +0.00855 [+0.00712, +0.0101] |
| Llama-3.2-1B | 1.24 | +0.0109 [+0.00976, +0.0120] | yes | — | — | — |
| Qwen2.5-1.5B | 1.54 | +0.00351 [+0.00282, +0.00420] | yes | +0.00612 | +0.00133 | +0.00480 [+0.00356, +0.00600] |
| Gemma-2-2B | 2.61 | +0.00527 [-0.00369, +0.0132] | no (interval crosses 0) | — | — | — |
| Qwen2.5-3B | 3.09 | +0.00331 [+0.00256, +0.00404] | yes | +0.00532 | +0.00113 | +0.00419 [+0.00305, +0.00535] |
| Llama-3.2-3B | 3.21 | +0.00629 [+0.00540, +0.00714] | yes | — | — | — |
| model | (J 0 ) [95% CI] | resolved | (10J 0 ) [95% CI] | resolved |
|---|---|---|---|---|
| Qwen2.5-0.5B | +0.00472 [+0.00331, +0.00607] | + | +0.0300 [+0.0258, +0.0341] | + |
| Llama-3.2-1B | -0.00313 [-0.00428, -0.00202] | (resolved negative) | -0.000204 [-0.00339, +0.00285] | unresolved |
| Qwen2.5-1.5B | +0.00170 [+0.000676, +0.00269] | + | +0.0133 [+0.0104, +0.0161] | + |
| Gemma-2-2B | +0.246 [+0.218, +0.273] | + | +0.806 [+0.767, +0.844] | + |
| Qwen2.5-3B | +0.00409 [+0.00308, +0.00512] | + | +0.0195 [+0.0164, +0.0227] | + |
| Llama-3.2-3B | +0.000976 [+0.000138, +0.00185] | + | +0.0107 [+0.00813, +0.0135] | + |
| Model | [95% CI] | [95% CI] |
|---|---|---|
| Qwen2.5-0.5B | 0.00249 [ 0.000190, 0.00518] | 0.0275 [0.0252, 0.0297] |
| Llama-3.2-1B | 0.00862 [ 0.0112, 0.00604] | 0.00842 [0.00657, 0.0103] |
| Qwen2.5-1.5B | 0.00278 [ 0.00510, 0.000565] | 0.0161 [0.0144, 0.0178] |
| Gemma-2-2B | 0.0512 [ 0.0854, 0.0175] | 0.857 [0.825, 0.889] |
| Qwen2.5-3B | 0.00204 [ 0.0000977, 0.00417] | 0.0175 [0.0156, 0.0194] |
| Llama-3.2-3B | 0.00371 [0.00142, 0.00607] | 0.00700 [0.00491, 0.00916] |
| Method | Object | Anchor | Calibration / training | Evaluation |
|---|---|---|---|---|
| FQ-ViT | post-softmax probabilities, 4-bit | absolute ( ) | none | ViT accuracy; no fused kernel |
| RepQ-ViT | post-softmax probabilities, ( at inference) | absolute, layer-wise scale | percentile calibration | ViT accuracy; no fused kernel |
| ITA | integer streaming softmax, power of two of the floor-rounded distance | row maximum (streaming) | clipping threshold by QAT | 22 nm ASIC |
| EXAQ | coarse exponential | row maximum | clipping range from activation statistics | frozen LLMs; isolated softmax timing |
| IntAttention | 32-entry exponential lookup table | row maximum | clipping range fixed offline | frozen LLMs; integer attention on Armv8 CPUs |
| FA4 | cubic exp2 emulation for part of each row | running maximum (exact exponential) | none; pointwise error | fused kernel on B200 |