Softmax attention gives every token a nonzero weight, which in trained models concentrates into attention sinks and massive activations that widen the dynamic range low-precision inference must cover. Softpick removes this constraint by rectifying scores, eliminating sinks and lowering hidden-state kurtosis, but its advantage fades at scale. We reframe this failure as a normalization problem. Softpick's denominator splits into positive- and negative-shifted sums D+ and D−, used identically in the forward and backward pass, preventing their roles from being isolated. We separate them into a family of operators that independently choose each denominator. The failure originates at initialization: every layer contains rows where D+ is exactly zero, while near-dead rows produce gradient norms above 1012 regardless of the backward denominator. Only Softpick and a stop-gradient variant, which keeps D++D− forward but backpropagates through D+ alone, train from scratch. At 230M parameters, the stop-gradient operator matches Softpick on quantization, has fewer dead heads, and retrieves passkeys more reliably, trailing only on peak attention-weight kurtosis.
Figures & tables
Operator
Forward Df
Backward Db
Role
Softpick
D++D−
D++D−
Coupled baseline
Stop-gradient
D++D−
D+ (stop-grad on D− )
The survivor
Rectified-forward
D+
D++D−
Honest gradient, rectified forward
Partial ( λ )
D++λD−
matched
λ∈{0.11,0.3}
Floored
max(D+,δ)
D++D−
Hard forward floor
Table 1: The operator family. Each row sets a forward and a backward denominator independently.
Figure 1: (a) Softpick keeps one attention spike per query row (max 0.0625); the D+ -only forward scatters spurious spikes at a higher peak (0.4333). (b) The stop-gradient operator ends 230M training with 23 dead heads of 238 against softpick’s 29.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Metric
ctx ∼ 1024
ctx ∼ 2048
ctx ∼ 3200
Exact match
0%
0%
0%
Target digits appear (substring)
0%
0%
0%
First generated number correct
0%
0%
0%
Digit recall (LCS / target length)
23.4%
32.7%
0.0%
Mean normalized edit distance
0.77
0.67
1.00
Generations containing any digit
32.0%
43.0%
0.0%
Appendix
Table 2: Digit-level analysis of softpick’s 230M generations, n=500 per context length. All strict digit metrics are 0% ; the low recall is chance overlap, and most generations contain no digits.
Condition
ctx ∼ 1024
ctx ∼ 2048
ctx ∼ 3200
abs (softpick baseline)
96.7%
93.3%
86.7%
relu ( D+ only)
0%
0%
0%
relu, rescaled
0%
0%
0%
damped c=0.01
0%
0%
0%
damped c=0.05
0%
0%
0%
damped c=0.10
0%
0%
0%
Appendix
Table 3: Full inference-only denominator sweep, 340M checkpoint, n=30 per cell.
Checkpoint
c
ctx ∼ 1024
ctx ∼ 2048
ctx ∼ 3200
Baseline
1.000
96.7%
93.3%
86.7%
step-500
0.309
100.0%
53.3%
40.0%
step-700
0.111
96.7%
90.0%
63.3%
step-900
0.012
83.3%
80.0%
30.0%
step-1000
0.000
0%
0%
0%
Appendix
Table 4: Passkey retrieval along the anneal, 340M. Step-700 is the last checkpoint before collapse.
Table 7: Passkey retrieval, 130M, exact match, n=500 per cell.
Operator
Forward
Backward
Outcome (from scratch)
Softpick
D++D−
D++D−
Trains stably to 28k+ steps
Stop-gradient
D++D−
D+
Trains stably to 28k+ steps
Rectified-forward
D+
D++D−
Loss flat ∼ 10.39, grad norm 1012 – 1014
Partial (0.11, 0.3)
D++λD−
matched
NaN before 1000 steps
Floored
max(D+,δ)
D++D−
Learns, oscillates ∼ 4 nats, grad norm spikes
Matched-rectified
D+
D+
Not run (Section J )
Appendix
Table 8: Every operator trained from scratch, and how it failed. Softpick and the stop-gradient operator are the only two that train.
Figure 2: Loss (left) and gradient norm (right) over training at 230M, for softpick, the stop-gradient operator, and the floored operator. The floored operator learns instead of pinning at random-guess loss, then oscillates well above the two stable operators and spikes to nearly 6×109 in gradient norm near step 19,000.
Figure 3: Attention-weight surface at layer 2, companion to Figure 1 . Softpick and the D+ -only forward reach the same maximum weight, 0.1875, but the D+ -only forward redistributes it across several spurious peaks.
Figure 4: Full training panel for the c:1→0 anneal at 340M: loss, gradient norm (log scale), learning rate, and the c -schedule, all against step.
130M
230M
340M
Parameters
130M
228.9M
340M
Hidden size
768
896
1024
Layers
12
17
24
Attention heads (total)
144 (12/layer)
238 (14/layer)
384 (16/layer)
Head dim
64
64
64
MLP
SwiGLU, ratio 4
SwiGLU, ratio 4
SwiGLU, ratio 4
Appendix
Table 9: Training configuration, read from our run configurations and training scripts and the publicly released 340M checkpoint.
Sharif University of Technology, Tehran, Iran · GISMA University of Applied Sciences, Potsdam, Germany · BRAINS, Brandenburg Research Center for Applied Intelligent Systems, Potsdam, Germany +3