cs.LGOct 4, 2026

Underscoring the Problem: Why Softpick Fails at Initialization

Authors: Aryan Sood, Jaikaran Singh, Ishaan Bansal

Organizations: IIT Roorkee

Abstract

Softmax attention gives every token a nonzero weight, which in trained models concentrates into attention sinks and massive activations that widen the dynamic range low-precision inference must cover. Softpick removes this constraint by rectifying scores, eliminating sinks and lowering hidden-state kurtosis, but its advantage fades at scale. We reframe this failure as a normalization problem. Softpick's denominator splits into positive- and negative-shifted sums D+D^+ and D−D^-, used identically in the forward and backward pass, preventing their roles from being isolated. We separate them into a family of operators that independently choose each denominator. The failure originates at initialization: every layer contains rows where D+D^+ is exactly zero, while near-dead rows produce gradient norms above 101210^{12} regardless of the backward denominator. Only Softpick and a stop-gradient variant, which keeps D++D−D^+ + D^- forward but backpropagates through D+D^+ alone, train from scratch. At 230M parameters, the stop-gradient operator matches Softpick on quantization, has fewer dead heads, and retrieves passkeys more reliably, trailing only on peak attention-weight kurtosis.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Pretraining Transformers with Quantized Softmax in Attention

    Sep 27, 2026Shangzhen Zhu, Muyan Hu, Tomasz KozlowskiTransformer AttentionGumbel-Softmax Relaxation

  2. EFQ-Softmax: Exp-Free Quantization for Softmax

    Sep 9, 2026Haohui Han, Yuming Wan, Hongni Wang +4Gumbel-Softmax RelaxationBlock Sparse Flash Attention

  3. From Attention Sensitivity to Layer Role: Revisiting Mixed-Precision Quantization of Transformers

    Sep 28, 2026Nafiseh HosseinpourFardi, Negar Alihadi, Mahmoudreza Babaei +2Post-Training QuantizationMixed-Precision Quantization