Softmax attention has two structural gaps. A head cannot abstain, because its weights sum to one, so it outputs something even when nothing is relevant. Nor can it filter what it reads, because its output is a weighted average of value vectors, passing interference as faithfully as signal. We call these missing primitives abstention and noise filtering. Recent studies report that gating the value pathway improves pretraining but attribute the gain to different causes. We show that a value gate partly supplies both primitives, which unifies the reported causes as views of one gain. We give each primitive its own mechanism in matched models of 10M to 350M parameters and measure what each contributes. The gain from gating is almost entirely abstention at 10M, whereas by 350M filtering contributes as much as abstention, so what a study observes depends on its scale. The two benefits are largely additive, with a small overlap. A gate determined by each value alone leaves the attention sink in place, whereas a query-controlled mechanism removes it. Injecting interference into the value reads shows that abstention and filtering protect against it in distinguishable ways. The same patterns appear in pretrained models up to 20B parameters.
Figures & tables
Figure 1 : The two primitives across scale. (a,b) For each gate form, the abstention increment (sink logit over baseline) and the filtering increment (combined variant over sink logit). The filtering increment grows with scale, and the abstention increment shrinks from 10M to 124M. Bars are standard errors over three seeds. (c) The average attention mass the sink-logit model places on its phantom grows with scale (bars are standard deviations over seeds). This is the phantom’s softmax weight, averaged over every query, head, layer, and validation sequence at the tier’s training context. (d) Gate selectivity, the attention-weighted standard deviation of the gate values among the reads a query attends to, averaged over queries, heads, and layers (seed 0 per tier). It is zero for a gate that attenuates every attended read equally.
Variant
10M
50M
124M
350M
Improvement over the baseline
sink logit
0.0185±0.0033
0.0068±0.0003
0.0052±0.0025
0.0103±0.0084
norm gate
0.0037±0.0026
0.0069±0.0008
0.0059±0.0012
–
projection gate
0.0113±0.0029
0.0094±0.0011
0.0091±0.0011
–
sink+norm
0.0177±0.0037
0.0106±0.0015
0.0087±0.0021
0.0187±0.0060
sink+proj
0.0221±0.0029
0.0133±0.0018
0.0106±0.0022
0.0193±0.0057
Table 1 : Paired validation-loss improvements (nats), mean over seeds with its standard error beneath. The lower block is the filtering increment, the improvement of each combined variant over the sink logit.
Figure 2 : Where the sink lives at 124M (1024-token context, seed 0). (a) Attention mass on the first key position, as a fraction of the mass on real keys, and on the phantom, as a fraction of all mass. The baseline and the norm gate park mass on the first key position, every variant with a phantom routes it there instead, and the projection gate keeps a position-0 sink of the baseline’s size. (b) The norm ratio divides the mean norm of the value vector at position 0, taken before any gate is applied, by the mean norm of the value vectors at positions 4 and later in the same head and the same sequences, averaged over heads and layers on fixed validation blocks. The baseline crushes that read, the norm gate crushes it less, and every variant with a phantom or a projection gate keeps it ordinary.
Tier
Sink
Gate
Both
Overlap
Filter
Share
Norm gate
10M
0.0185
0.0037
0.0177
0.0045
−0.0008
−5%
50M
0.0068
0.0069
0.0106
0.0030
0.0030
28%
124M
0.0052
0.0059
0.0087
0.0024
0.0036
41%
350M
0.0103
–
0.0187
–
0.0084
45%
Projection gate
Table 2 : Two-by-two decomposition of the benefit of each gate form (validation-loss improvement over the baseline, nats, paired per seed). Sink, gate, and both are the improvements from the sink logit alone, the gate alone, and the combined variant. Overlap is sink plus gate minus both, filter is both minus sink, and share is filter as a percentage of both.
mass (%)
loss increase at dose ϵ
Variant
quiet
phantom
0.2
0.4
0.8
1.6
124M
baseline
36
0
0.018
0.110
2.63
4.95
sink logit
20
31
0.005
0.021
0.13
3.16
norm gate
44
0
0.004
0.042
4.40
6.82
projection gate
22
0
0.006
0.029
0.26
1.89
Table 3 : Injected interference at 124M and 350M, seed 0. Quiet is the share of a query’s attention mass, phantom included, that lands on reads below the head’s 25th-percentile norm in that sequence, which is the share exposed to the injection, and phantom is the share on the phantom, zero by construction for variants without one, each averaged over all queries, heads, and layers on eight validation batches at the training context. The dose columns give the mean increase in validation loss (nats) on the same batches under structured junk injected into those reads at dose ϵ . The gate-only variants were not trained at 350M.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Tier
Layers
Heads
Width
Context
Batch
Iterations
Warmup
Tokens
10M
6
6
384
512
64×1
6,000
200
0.20B
50M
10
10
640
1024
16×2
29,000
500
0.95B
124M
12
12
768
1024
8×4
48,800
800
1.60B
350M
24
16
1024
1024
8×4
177,000
2,000
5.80B
Appendix
Table 4 : Per-tier training configuration. Batch is the number of sequences per optimizer step, written as micro-batch size times accumulation steps.
Variant
Abstention
Filter
Params
baseline
none
none
0
offbyone
phantom, fixed
none
0
sinklogit
phantom, learned
none
H
sinktoken
token, approximate
none
dmodel
normgate
none; Zi scaling
norm threshold
H
projgate
none; Zi scaling
learned direction
H(d+1)
Appendix
Table 5 : The eight variants. Parameters are counted per layer for H heads of dimension d . Every variant is compatible with the key-value cache, because each gate is a function of the read it multiplies. The gates have no routing-side abstention. Their row scaling Zi (Lemma 3.1 ) is noted because it can shrink a row’s output.
Variant
10M
50M
124M
350M
Baseline loss
baseline
4.2119
3.5399
3.3710
3.0293
Improvement over baseline (paired per seed)
off-by-one
−0.0160±0.0027
–
–
–
sink token
−0.0025±0.0041
–
–
–
sink logit
−0.0185±0.0033
−0.0068±0.0003
−0.0052±0.0025
−0.0103±0.0084
Appendix
Table 6 : Final validation loss (nats) across four model sizes. Each improvement is the mean of per-seed paired differences with its standard error, and negative is better. The last block measures what each gate adds once a sink logit is in place. Variants absent at a tier are marked with a dash.
Tier
Variant
seed 0
seed 1
seed 2
seed 3
seed 4
mean
FineWeb 10M
baseline
4.2141
4.2060
4.2155
–
–
4.2119
off-by-one
4.1988
4.1944
4.1945
–
–
4.1959
sink token
4.2037
4.2061
4.2184
–
–
4.2094
sink logit
4.1981
4.1916
4.1904
–
–
4.1934
norm gate
4.2099
4.2071
4.2076
–
–
4.2082
projection gate
4.2070
4.1890
4.2057
–
–
4.2006
Appendix
Table 7 : Final validation loss (nats) of every run. Dashes mark seeds that were not trained.
Tier
improvement over baseline
mean gate value
reads below 0.5
10M
−0.0017±0.0014
0.999
0.0%
50M
+0.0017±0.0002
0.999
0.0%
124M
−0.0001±0.0005
0.987
0.7%
Appendix
Table 8 : The norm gate in the routing slot. Improvement is the mean of per-seed paired differences from the baseline with its standard error, and positive is better. The gate value and the fraction of reads below one half are averaged over the three seeds.
Figure 3 : Robustness to injected junk. (a,b) Increase in validation loss (log scale) under structured junk injected into the quietest quarter of each head’s reads, at 124M and 350M. Every variant with a sink logit or a gate is several times more robust than the baseline at low and moderate doses. At high doses the norm gate, and the sink+norm variant that contains it, become worse than the baseline, whereas the projection gate does not. (c) At 124M, junk along a random fixed direction hurts the projection gate no more than the baseline, but junk along the gate’s own learned direction collapses it, and collapses sink+proj further.
Tier
Junk type (cutoff)
Variant
ϵ=0.2
ϵ=0.4
ϵ=0.8
ϵ=1.6
FineWeb 10M
structured (other context) (q25)
baseline
+0.032
+0.138
+0.771
+2.695
sink logit
+0.009
+0.041
+0.178
+0.778
norm gate
+0.027
+0.124
+0.877
+3.162
projection gate
+0.026
+0.100
+0.408
+1.368
sink+norm
+0.010
+0.047
+0.252
+1.685
sink+proj
+0.007
+0.030
+0.133
+0.594
Appendix
Table 9 : Full injection battery at the 10M and 50M tiers. Entries are the increase in validation loss (nats) over the uncorrupted model.
Tier
Junk type (cutoff)
Variant
ϵ=0.2
ϵ=0.4
ϵ=0.8
ϵ=1.6
124M
structured (other context) (q25)
baseline
+0.018
+0.110
+2.632
+4.948
sink logit
+0.005
+0.021
+0.134
+3.155
norm gate
+0.004
+0.042
+4.403
+6.824
projection gate
+0.006
+0.029
+0.263
+1.891
sink+norm
+0.004
+0.020
+0.513
+5.544
sink+proj
+0.004
+0.020
+0.183
+1.774
Appendix
Table 10 : Full injection battery at the 124M and 350M tiers.
50M
124M
350M
Variant
median
q25
phantom
median
q25
phantom
median
q25
phantom
baseline
0.551
0.349
–
0.563
0.365
–
0.632
0.458
–
sink logit
0.361
0.215
0.273
0.337
0.197
0.311
0.224
0.132
0.530
norm gate
0.604
0.418
–
0.619
0.443
–
–
–
–
projection gate
0.431
0.232
–
0.419
0.223
–
–
–
–
sink+norm
0.458
0.299
0.156
0.451
0.298
0.186
0.384
0.267
0.346
Appendix
Table 11 : Exposure to injected junk. Share of each attention row’s total mass placed on quiet reads (below the per-head median or 25th percentile norm) and on the phantom, averaged over heads and rows at seed 0 and measured at the training context, as in the injection experiments. Mass on the phantom is mass not exposed to corruption.
Variant
LAMBADA acc
HellaSwag acc
WikiText-2 ppl
baseline
0.182±0.006
0.323±0.002
58.1±0.5
sink logit
0.179±0.004
0.325±0.003
57.0±0.5
norm gate
0.185±0.001
0.323±0.002
57.3±0.2
projection gate
0.186±0.002
0.325±0.003
57.3±0.8
sink+norm
0.186±0.003
0.323±0.002
56.8±0.1
sink+proj
0.187±0.007
0.326±0.003
56.6±0.3
Appendix
Table 12 : Zero-shot evaluation of the 124M models, mean and standard deviation over three seeds, each model evaluated at its 1,024-token context with the GPT-2 tokenizer. LAMBADA is last-word accuracy on the 5,153 passages of the OpenAI variant of the test set: the final word of each passage is the target, and a passage counts as correct only if every token of the target is the greedy prediction at its position. HellaSwag is accuracy on the first 3,000 validation examples, choosing among the four endings by the mean per-token log-probability of the ending given the context. WikiText-2 is perplexity on the raw validation split, concatenated and cut into non-overlapping 1,024-token windows, computed as the exponential of the mean next-token loss.
Figure 4 : The channel-blind superposition task. Task error against superposition load for the baseline, off-by-one abstention, the norm gate, and a trained oracle. The norm gate approaches the oracle’s error, whereas off-by-one abstention does not help.
Model
Junk
ϵ=0.2
ϵ=0.4
ϵ=0.8
ϵ=1.6
Pythia-160M
other context
+0.13
+0.52
+2.24
+5.35
same context
+0.12
+0.51
+2.42
+5.43
Gaussian
+0.08
+0.25
+1.18
+4.78
Pythia-410M
other context
+0.15
+0.94
+5.47
+7.89
same context
+0.16
+1.09
+6.08
+7.79
Gaussian
+0.09
+0.43
+2.31
+6.72
Appendix
Table 13 : Injection replication on pretrained Pythia models. Increase in WikiText-2 NLL (nats) over the clean model when the quietest quarter of each head’s reads receives junk of the head’s median read norm scaled by ϵ , transplanted from another context, transplanted from the same sequence, or drawn as Gaussian noise.
Layer type
phantom mass
mass on position 0
r0 ratio
sliding-window layers (12)
0.48
0.003
0.98
full-attention layers (12)
0.34
0.064
1.31
Appendix
Table 14 : Sink measurement on gpt-oss-20b, mean over heads and layers of each type: attention mass on the phantom, share of the remaining mass on the first key position, and the position-0 value-norm ratio.
124M
350M
Variant
pos. 0 (%)
phantom (%)
norm ratio
pos. 0 (%)
phantom (%)
norm ratio
baseline
6.0
0
0.42
11.1
0
0.36
norm gate
3.7
0
0.65
–
–
–
projection gate
7.1
0
1.16
–
–
–
sink logit
0.5
31.1
0.99
0.7
53.0
1.01
sink+norm
0.4
18.6
0.96
0.7
34.6
0.99
Appendix
Table 15 : Sink analysis at the training context (1024 tokens), seed 0: attention mass on the first key position as a percentage of the mass on real keys, mass on the phantom as a percentage of all mass, and the norm ratio of the position-0 read. The phantom share is zero by construction for variants without one, and dashes mark variants not trained at 350M.
When attention concentrates on a single token, a sink, what is the model actually computing? Attention sinks are ubiquitous in softmax transformers, yet this shared visual signature can hide fundamentally different algorithms. We show that visually similar sink patterns can reflect two distinct mechanisms: {i} adaptive nop, where a head suppresses its update by routing to a null token, and {ii} broadcast, where a sink aggregates and redistributes global information. In that case, sinks serve an analogous role: a safe destination when there is nothing useful to compute. Proposed interventions like gating or registers work because they implicitly target one or the other, revealing a duality between method and assumed mechanism: gating implicitly assumes nop; registers implicitly assume broadcast. Each mechanism leaves distinct traces (nop sinks exhibit negligible value norms; broadcast sinks induce low-rank outputs) which we formalize on synthetic tasks and use to derive practical diagnostics. Applied to pretrained vision transformers, these diagnostics reveal that both mechanisms exist at scale: sinks transition from CLS in early layers to patches in deeper layers, and concentrate in specialized heads. Strikingly, register tokens, designed for broadcast, are repurposed to also serve nop, confirming that neither intervention alone suffices. Combining gating with registers yields complementary gains in stability and performance. Overall, we find that the same attention pattern can reflect two very different computations and effective intervention requires first asking what the model is actually computing.
Lukas Fesser, Mozes Jacobs, Thomas Fel +2
Kempner Institute Harvard University Cambridge, MA
Transformers commonly exhibit an attention sink: disproportionately high attention to the first position. We study this behavior in GPT-2-style models with learned query biases and absolute positional embeddings. Combining structural analysis with causal interventions, validated across natural-language, mathematical, and code inputs, we find that the sink arises from the interaction among (i) a learned query bias, (ii) the first-layer MLP transformation of the positional encoding, and (iii) structure in the key projection. Crucially, each component we identify is individually dispensable: architectures omitting each of them robustly exhibit sinks. This indicates that attention sinks may arise through distinct circuits across architectures. These findings inform mitigation of sinks, and motivate broader investigation into why sinks emerge.
The attention mechanism forms the foundation of many modern AI models such as the Transformer. In one subclass of problems where attention is used, inputs and outputs are bound to the probability simplex so that all outputs sum to one. In this setting, softmax attention admits an exact, component-by-component quantum realization. Attention scores are Hadamard-test statistics on block-encoded projections of amplitude-encoded inputs. The exponential softmax is the interior of a cosine-squared family generated by Born-rule measurement under an exact bijection, whose boundary expresses sparse attention with exact zeros at finite parameter values. The softmax temperature is a repetition count where post-selected measurement rounds realize discretized inverse temperature exactly. Value aggregation is a deterministic column-loading channel that dilates the column-stochastic value matrix. The gated residual is the preparation angle of a single ancilla, with the additive identity at a mixing angle of π/2. Every learnable parameter is a rotation-gate angle. The composed layer is exact in the infinite-shot limit with one measure-and-reload step per attention score; a fully-coherent variant is ε-approximate via quantum singular value transformation in the infinite depth limit. The algebraic core is machine-checked in Lean 4.
Eric A. F. Reinhardt, Adam J. Hauser
Department of Physics and Astronomy University of Alabama, Tuscaloosa, AL 35487, USA