Softmax Reparameterization for Output-Head Quantization
Organizations: Adobe Search, Discovery & ContentAI
Abstract
Large vocabularies make output heads a substantial inference cost in small language models. We introduce softmax reparameterization, a post-training method that searches over functionally equivalent output heads before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL. For linear-softmax heads, these shifts preserve full-precision predictions exactly and require no decoder retraining; a rank-one correction extends the construction to nonlinear logit paths. Across seven output heads and three quantizers, W4 gains are largest where baseline quantization substantially distorts predictions: test KL falls by 93% on XGLM under RTN and by 73--77% on Phi, BLOOM, and BLOOMZ under activation-weighted MSE. Heads with low baseline error change little; at W2, used as a compression stress test, benefits extend more broadly. On Phi, the gains persist under stronger GPTQ calibration; a separate untouched holdout reproduces the improvements on Phi and BLOOM. Frozen WikiText-selected coefficients also transfer without retuning to C4 and OpenWebMath. Residual analysis on Phi shows how fidelity can improve despite greater total logit error: the selected representative reduces error on likely outputs and lowers its Fisher-weighted cost. For shift-compatible heads, the shift adds no inference operation. With the decoder held in BF16, a packed W4 Phi output head reduces batch-one generation latency by 10.8%, and reparameterization preserves this speedup.
Figures & tables
| W4 KL | W2 KL | |||||
|---|---|---|---|---|---|---|
| Model | RTN | AW-MSE | GPTQ | RTN | AW-MSE | GPTQ |
| Phi-4-mini | 1.23 0.351 | 0.936 0.256 | 0.158 0.059 | 98.2 74.2 | 41.6 16.2 | 3.08 1.16 |
| Gemma 3 | 0.049 0.046 | 0.041 0.038 | 0.033 0.033 | 2.66 2.21 | 0.822 0.714 | 0.508 0.491 |
| Gemma 4 † | 0.012 0.010 | 0.008 0.008 | 0.008 0.008 | 0.911 0.911 | 0.212 0.199 | 0.165 0.162 |
| Qwen3.5 | 0.021 0.019 | 0.013 0.013 | 0.011 0.010 | 1.84 1.74 | 0.312 0.285 | 0.200 0.197 |
| BLOOM-1.7B | 1.12 0.531 | 0.594 0.136 | 0.037 0.028 | 16.7 6.15 | 4.46 1.46 | 0.427 0.335 |
| RTN | AW-MSE | GPTQ | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Head | |||||||||
| Phi-4-mini | 1.229 | 0.811 | 0.351 | 0.936 | 0.577 | 0.256 | 0.158 | 0.117 | 0.059 |
| BLOOM-1.7B | 1.117 | 0.978 | 0.531 | 0.594 | 0.451 | 0.136 | 0.037 | 0.036 | 0.028 |
| BLOOMZ-1.7B | 1.119 | 0.859 | 0.483 | 0.662 | 0.399 | 0.158 | 0.042 | 0.039 | 0.032 |
| XGLM-1.7B | 2.129 | 0.143 | 0.143 | 0.584 | 0.095 | 0.095 | 0.009 | 0.007 | 0.007 |
| Head | KL to BF16 | WikiText PPL | B1 (ms) | B16 (ms) |
|---|---|---|---|---|
| BF16 | 0.00 | 11.65 | 1136.6 | 1335.8 |
| W4 min–max | 1.23 | 40.57 | 1015.3 | 1211.8 |
| W4 min–max + shift | 0.34 | 16.28 | 1014.1 | 1210.1 |
| W4 AW-MSE | 0.97 | 31.65 | 1014.8 | 1210.2 |
| W4 AW-MSE + shift | 0.28 | 15.40 | 1013.4 | 1210.3 |
| W4 GPTQ | 0.15 | 13.67 | — | — |
| Quantizer | (%) | Common (%) | Fisher/2 | Actual KL | ||
|---|---|---|---|---|---|---|
| RTN | 0 | 0.141 | 0.239 | 0.03 | 1.271 | 1.223 |
| RTN | 1 | 0.137 | 0.189 | 0.00 | 0.720 | 0.821 |
| RTN | 4 | 0.150 | 0.707 | 5.67 | 0.347 | 0.352 |
| AW-MSE | 0 | 0.125 | 0.249 | 4.06 | 0.861 | 0.935 |
| AW-MSE | 1 | 0.121 | 0.191 | 0.13 | 0.532 | 0.583 |
| AW-MSE | 4 | 0.136 | 0.677 | 27.51 | 0.240 | 0.256 |
| Source probability | Mass (%) | Visible logit error, | Fisher/2, | Drop (%) |
|---|---|---|---|---|
| 84.766 | 71.8 | |||
| 15.037 | 27.9 | |||
| 0.196 | 0.3 |
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
| Comparison | Codes | Scale selection/storage | Reconstruction |
|---|---|---|---|
| Matrix/breadth RTN | Symmetric | FP32 | FP32 |
| Matrix/breadth AW-MSE | Signed | Score and store BF16 | BF16, then FP32 products |
| GPTQ probe | Signed | BF16 initial scales | BF16, then FP32 products |
| Matched Phi residuals | Signed | Score and store BF16 | BF16, then FP32 products |
| Packed Phi RTN/AW-MSE | Signed | Float candidates; store BF16 | Packed W4 serving |
| Head | Quantizer | KL | PPL | BF16 PPL |
|---|---|---|---|---|
| Phi | AW-MSE | 9.33 | ||
| Phi | RTN | 9.33 | ||
| BLOOM | AW-MSE | — | — | |
| BLOOM | GPTQ | — | — |
| Model | Language | RTN | AW-MSE | GPTQ |
|---|---|---|---|---|
| BLOOM-1.7B | English | |||
| French | ||||
| Spanish | ||||
| Arabic | ||||
| Hindi | ||||
| Chinese |
| Head | Quantizer | Domain | Raw KL | Centered KL | Frozen KL | KL: 95% CI | |
|---|---|---|---|---|---|---|---|
| Phi-4-mini | RTN | C4 | 1.372879 | 0.899519 | 0.393513 | ||
| Phi-4-mini | RTN | OpenWebMath | 1.105418 | 0.675704 | 0.342138 | ||
| Phi-4-mini | AW-MSE | C4 | 1.039478 | 0.657195 | 0.304324 | ||
| Phi-4-mini | AW-MSE | OpenWebMath | 0.764193 | 0.498817 | 0.252586 | ||
| Phi-4-mini | GPTQ | C4 | 0.111289 | 0.078885 | 0.041984 | ||
| Phi-4-mini | GPTQ | OpenWebMath | 0.101269 | 0.072217 | 0.044054 |
| Quantizer | WikiText (16) | C4 (2,048) | OpenWebMath (2,048) |
|---|---|---|---|
| RTN | |||
| AW-MSE | |||
| GPTQ |
| Quantizer | Domain | Selection at | Selection at | Test KL: |
|---|---|---|---|---|
| RTN | C4 | : 6/10; : 4/10 | : 10/10 | |
| RTN | OpenWebMath | : 9/10; : 1/10 | : 10/10 | |
| AW-MSE | C4 | : 10/10 | : 10/10 | |
| AW-MSE | OpenWebMath | : 3/10; : 7/10 | : 10/10 | |
| GPTQ | C4 | : 3/10; : 7/10 | : 10/10 | |
| GPTQ | OpenWebMath | : 5/10; : 5/10 | : 10/10 |
| Bits | Model | RTN | AW-MSE | GPTQ |
|---|---|---|---|---|
| W2 | Phi | 98.2 74.2 | 41.6 16.2 | 3.08 1.16 |
| W2 | Gemma 3 | 2.66 2.21 | 0.822 0.714 | 0.508 0.491 |
| W2 | Gemma 4 † | 0.911 0.911 | 0.212 0.199 | 0.165 0.162 |
| W2 | Qwen3.5 | 1.84 1.74 | 0.312 0.285 | 0.2 0.197 |
| W2 | BLOOM-1.7B | 16.7 6.15 | 4.46 1.46 | 0.427 0.335 |
| W2 | BLOOMZ-1.7B | 16.2 6.22 | 3.99 2.12 | 0.456 0.367 |
| Head | BF16 | RTN | AW-MSE | GPTQ |
|---|---|---|---|---|
| Phi | 9.73 | 111.5 | 1002 28.47 | 16.86 11.97 |
| Gemma 3 | 38.49 | 44.34 44.34 | 38.54 38.89 | 40.09 40.09 |
| Gemma 4 † | 66.41 | 71.11 70.83 | 68.71 69.06 | 69.05 68.85 |
| Qwen3.5 | 8.98 | 9.851 9.793 | 9.503 9.446 | 9.351 9.291 |
| BLOOM-1.7B | 18.47 | 282.3 100.7 | 50.78 26.89 | 20.14 20.04 |
| BLOOMZ-1.7B | 22.06 | 285.3 124.9 | 56.08 31.42 | 23.8 23.9 |
| Condition | Test KL | Test PPL | KL reduction |
|---|---|---|---|
| GPTQ: 1,024 states | |||
| GPTQ: 65,536, original ranges | |||
| GPTQ: 65,536, refitted ranges | |||
| RTN | |||
| AW-MSE | |||
| Scaled RTN |
| RTN | AW-MSE | |||||
|---|---|---|---|---|---|---|
| Direction | Test KL | Test PPL | Test KL | Test PPL | ||
| Vocabulary mean | 0.35058 | 13.4611 | 0.25551 | 12.4856 | ||
| Coordinate-wise median | 0.73863 | 20.1241 | 0.55894 | 17.0431 | ||
| Random seed 0 | 1.05676 | 27.7650 | 0.77292 | 21.8672 | ||
| Random seed 1 | 1.06612 | 28.1215 | 0.86884 | 22.6155 | ||
| Random seed 2 | 1.04971 | 27.0636 | 0.84131 | 22.1993 | ||
| Condition | test KL | interval |
|---|---|---|
| GPTQ: 1,024 states | ||
| GPTQ: 65,536, original ranges | ||
| GPTQ: 65,536, refitted ranges | ||
| RTN | ||
| AW-MSE | ||
| Scaled RTN |
| RTN | AW-MSE | |||||
|---|---|---|---|---|---|---|
| Model | Raw KL | Reparam KL | Raw KL | Reparam KL | ||
| Phi | 4 | 1.229 | 0.351 | 4 | 0.936 | 0.256 |
| Gemma 3 | 0.5 | 0.049 | 0.046 | 0.041 | 0.038 | |
| Gemma 4 † | 1 | 0.012 | 0.010 | 0.5 | 0.008 | 0.008 |
| Qwen3.5 | 0.021 | 0.019 | 0.013 | 0.013 | ||
| Best sampled | Diagnostic KL | |||||
|---|---|---|---|---|---|---|
| Model | Bits | MSE | Proj. MSE | KL | At MSE min | At KL min |
| Gemma 3 | 4 | 1 | 1 | 0.5 | 0.0580 | 0.0465 |
| Gemma 3 | 3 | 1 | 1 | 0.25 | 0.3161 | 0.2340 |
| Gemma 3 | 2 | 1 | 1 | 0 | 2.8305 | 2.5793 |
| Qwen3.5 | 4 | 1 | 1 | 0 | 0.0291 | 0.0202 |
| Qwen3.5 | 3 | 1 | 1 | 0 | 0.1599 | 0.1017 |
| Model | Shared weight energy (%) | Range ratio | Raw KL | Centered KL |
|---|---|---|---|---|
| Gemma 3 | 3.41 | 0.9612 | 0.0499 | 0.0585 |
| Qwen3.5 | 14.23 | 0.9029 | 0.0196 | 0.0294 |
| Phi | 16.40 | 1.0128 | 1.2109 | 0.8087 |
| Quantizer | Source probability | Mass (%) | Error share (%) | Error energy | Fisher share (%) | |
|---|---|---|---|---|---|---|
| RTN | 0 | 0.196 | 93.474 | 134039.57 | 0.317 | |
| RTN | 0 | 2.196 | 6.024 | 8637.93 | 4.015 | |
| RTN | 0 | 12.842 | 0.476 | 682.63 | 24.223 | |
| RTN | 0 | 84.766 | 0.026 | 36.88 | 71.444 | |
| RTN | 1 | 0.196 | 94.179 | 106621.80 | 0.392 | |
| RTN | 1 | 2.196 | 5.381 | 6092.47 | 4.994 |
| (%) | Common (%) | Fisher/2 | Actual KL | ||
|---|---|---|---|---|---|
| 0 | 0.148515 | 0.025340 | 0.4644 | 0.088771 | 0.089362 |
| 4 | 0.162178 | 0.067442 | 4.7268 | 0.034710 | 0.034909 |
| Model | KL | ||||
|---|---|---|---|---|---|
| Gemma 3 4B | 0.041 | ||||
| Qwen3.5 4B | 0.013 | ||||
| Phi-4-mini | 0.925 |
| BF16 head | WikiText PPL | NLL vs. source |
|---|---|---|
| Tied source | 11.6477 | +0.000000 |
| Untied, | 11.6455 | -0.000194 |
| Untied, | 11.6363 | -0.000983 |
| Representation | Head payload (GiB) | Relative to BF16 |
|---|---|---|
| BF16 | 1.2503 | 1.0000 |
| W8, G128 | 0.6349 | 0.5078 |
| W4, G128 | 0.3223 | 0.2578 |
| W4, G32 | 0.3516 | 0.2812 |
| Head | B1 latency (ms) | B16 throughput (tokens/s) |
|---|---|---|
| BF16 | 516.6 | 570.9 |
| W8 AW-MSE | 480.4 | 585.9 |
| W4 AW-MSE | 467.2 | 614.8 |
| Model | Vocabulary | Width | Head (M) | Nominal share (%) |
|---|---|---|---|---|
| Gemma 3 4B | 262,208 | 2,560 | 671 | 16.8 |
| Gemma 4 E4B | 262,144 | 2,560 | 671 | — |
| Qwen3.5 4B | 248,320 | 2,560 | 636 | 15.9 |
| Phi-4-mini | 200,064 | 3,072 | 615 | 16.2 |
| BLOOM-1.7B | 250,880 | 2,048 | 514 | 30.2 |
| BLOOMZ-1.7B | 250,880 | 2,048 | 514 | 30.2 |
| Family | Earlier checkpoint and rows | Selected checkpoint and rows | |
|---|---|---|---|
| Gemma | google/gemma-2b [ 2024 ]; | google/gemma-3-4b-it [ 2025 ]; | 2560 |
| XGLM | — | facebook/xglm-1.7B [ 2021 ]; | 2048 |
| BLOOM | — | bigscience/bloom-1b7 [ 2022 ]; | 2048 |
| Qwen | Qwen/Qwen-7B [ 2023 ]; | Qwen/Qwen3.5-4B [ 2026 ]; | 2560 |
| Llama | meta-llama/Llama-2-7b-hf [ 2023 ]; | meta-llama/Llama-4-Scout-17B-16E-Instruct [ 2025 ]; | 5120 |
| GPT | openai-community/gpt2 [ 2019 ]; | openai/gpt-oss-20b [ 2025 ]; | 2880 |
| Family | Model | Vocabulary |
| Phi | Phi-2 | 51,200 |
| Phi-3-mini | 32,064 | |
| Phi-4 | 100,352 | |
| Phi-4-mini | 200,064 | |
| Llama | Llama 2 | 32,000 |
| Llama 3 | 128,256 |
| Model | Head (M) | BF16 | W4 min–max | W4 AW-MSE | W8 AW-MSE |
|---|---|---|---|---|---|
| Gemma 3 4B | 671 | 60.12 | 60.84 | 59.46 | 60.12 |
| Gemma 4 E4B | 671 | 74.97 | 75.48 | 75.55 | 74.95 |
| Qwen3.5 4B | 636 | 10.89 | 11.19 | 10.98 | 10.89 |
| Phi-4-mini | 615 | 11.65 | 40.58 | 30.94 | 11.75 |
| Nine-model experiment | Matched-budget SmolLM3 study | |
|---|---|---|
| Initialization | Main 14-point scalar grid | Shared 27-point scalar grid |
| Scalar budget | 14 | (scalar refinement) |
| Grouped budget | ||
| Evaluation | Main 16 test articles | 16 previously unused articles |
| Quantizers | RTN, AW-MSE, GPTQ | RTN, AW-MSE |
| Model | Quantizer | Raw | Scalar | Grouped | Red. (%) | KL: 95% CI | |
|---|---|---|---|---|---|---|---|
| Phi-4-mini | RTN | 4 | 1.230021 | 0.350579 | 0.300249 | ||
| AW-MSE | 4 | 0.939398 | 0.255508 | 0.222474 | |||
| GPTQ | 4 | 0.160414 | 0.060511 | 0.058634 | |||
| Gemma 3 | RTN | 0.5 | 0.049208 | 0.046005 | 0.038335 | ||
| AW-MSE | -1 | 0.040232 | 0.037523 | 0.033956 | |||
| GPTQ | 0 | 0.033609 | 0.033609 | 0.032296 |
| Scalar grouped PPL | ||||
|---|---|---|---|---|
| Model | Source PPL | RTN | AW-MSE | GPTQ |
| Phi-4-mini | 9.719 | 13.461 13.042 | 12.486 12.106 | 10.286 10.287 |
| Gemma 3 | 37.977 | 38.482 38.419 | 38.086 38.176 | 38.852 38.613 |
| Gemma 4 † | 67.223 | 68.156 68.011 | 68.192 68.129 | 67.646 67.961 |
| Qwen3.5 | 8.977 | 9.131 9.143 | 9.106 9.060 | 9.065 9.054 |
| Qwen3 | 12.577 | 13.020 12.839 | 12.747 12.617 | 12.636 12.639 |