cs.LGSep 25, 2026

Softmax Reparameterization for Output-Head Quantization

Authors: Asim Kadav, Christian Flores, Chirag Arora, Varun Kotte, Hongbo Zheng, Lan Yan, Priya Shanmugasundaram, Tracy Holloway King

Organizations: Adobe Search, Discovery & ContentAI

Abstract

Large vocabularies make output heads a substantial inference cost in small language models. We introduce softmax reparameterization, a post-training method that searches over functionally equivalent output heads before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL. For linear-softmax heads, these shifts preserve full-precision predictions exactly and require no decoder retraining; a rank-one correction extends the construction to nonlinear logit paths. Across seven output heads and three quantizers, W4 gains are largest where baseline quantization substantially distorts predictions: test KL falls by 93% on XGLM under RTN and by 73--77% on Phi, BLOOM, and BLOOMZ under activation-weighted MSE. Heads with low baseline error change little; at W2, used as a compression stress test, benefits extend more broadly. On Phi, the gains persist under stronger GPTQ calibration; a separate untouched holdout reproduces the improvements on Phi and BLOOM. Frozen WikiText-selected coefficients also transfer without retuning to C4 and OpenWebMath. Residual analysis on Phi shows how fidelity can improve despite greater total logit error: the selected representative reduces error on likely outputs and lowers its Fisher-weighted cost. For shift-compatible heads, the shift adds no inference operation. With the decoder held in BF16, a packed W4 Phi output head reduces batch-one generation latency by 10.8%, and reparameterization preserves this speedup.

Figures & tables

Appendix figures & tables27 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SoftWater: Class-Aware Rate Allocation for Softmax Quantization

    Aug 12, 2026Joao V. Cavalcanti, Ashia C. WilsonGumbel-Softmax RelaxationKullback-Leibler Divergence

  2. ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads

    Aug 3, 2026Şuayp Talha Kocabay, Talha Rüzgar Akkuş, Kamer Ali YukselLarge Language Model QuantizationMatched Fp16 Intermediate