Massive activation features (MAs) in Transformers are extreme-value residual-stream features that persist across layers despite the model's ability to suppress them. Why do they survive? Our investigation using an operator-level mechanistic analysis of attention and feed-forward (FFN) blocks reveals that these blocks systematically ignore MA coordinates while reading, but not while writing; creating a read-write asymmetry that blocks corrective feedback while allowing continued accumulation. We find that both attention and feed-forward layers have this read-blindness, and contribute to the emergence and persistence of MAs. To validate prior work that hypothesized that FFN's amplification abilities is the primary reason for MAs (Sun et al., 2026), we analyze the model checkpoints during learning. Contrary to our expectation, read-blindness emerges before FFN amplification, suggesting that it acts upstream in the MA mechanism. We further contribute gradient analysis to link this behavior to surprising asymmetries in the loss landscape, concluding that the model actively maintains this read-blindness. Finally, we find that removing read-blocking at different locations induces compensatory shifts elsewhere, but MAs still persist.
Figures & tables
Llama 135M
Llama 1.28B
Llama 2.56B
Qwen3 1.7B
Massive coordinates ∣M∣
10
10
5
7
Maximum activation magnitude
280
2.27×103
3.92×103
2.23×103
Input availability:
Attention normalized input, X~
0.58 (-0.60)
0.68 (-0.73)
0.55 (-0.64)
0.58 (-0.37)
FFN normalized input, Y~
0.56 (-0.65)
0.56 (-0.13)
0.50 (-0.63)
0.54 (-0.48)
Read side weight view:
Table 1: Erasure results. Cells report the median eraM with rank-biserial rb against the full complement ¬M in parentheses. Bold cells have rb>0.95 ; stars denote one-sided Mann–Whitney tests of eraM>era¬M with p<0.05 .
Signed alignment
Llama 135M
Llama 1.28B
Llama 2.56B
Qwen3 1.7B
Attention, Xk with Δkattn
+0.03 (+0.45)
+0.09 (+0.29)
+0.02 (-0.27)
+0.06 (+0.16)
FFN, Yk with Δkffn
+0.05 (+0.57)
+0.25 (+0.59)
-0.02 (-0.20)
+0.10 (+0.14)
Table 2: Signed-alignment results. For each coordinate, S is median-aggregated across layers. Cells report the median SM with rank-biserial rb comparing M and ¬M in parentheses.
Diagnostic
value M
rb
∥Uk∥F
7.74
+0.78
∣λ⋆(k)∣
2.19
+0.98
∣λ⋆(k)∣/∥Sk∥F
0.39
+1.00
Within-set ∣cos(s⋆)∣
0.53/0.46
—
Table 3: s⋆ amplifier diagnostics for Llama 1.28 Base. Per-coordinate entries are medians over M and a deterministic, size-matched ¬M ; the final row is the layer-mean of the within-set median pairwise absolute cosine. rb>0 means larger values on M .
Figure 1: FFN read-blindness precedes amplifier specialization. We assign massive and control coordinates at the final checkpoint and trace these fixed sets through 21 checkpoints from 0 to 20 k training steps. (a) Maximum-over-layer FFN erasure for eventual massive coordinates (orange) and non-massive controls (gray). (b) Maximum-over-layer leading amplifier gain for the same sets. Lines and shaded regions show medians and interquartile ranges; vertical dotted lines mark tsep , the first checkpoint at which the rank-biserial effect reaches rb≥0.95 . FFN erasure separates at 2.3 k steps, before amplifier gain at 4.5 k steps. (c) Final-checkpoint amplifier norm ∥Uk∥F versus the inverse participation ratio of the corresponding FFN write row Wdown[k,:] . Colored points are massive coordinates, with size denoting massiveness and color denoting attention-input erasure; gray points are non-massive coordinates.
Figure 2: Directional curvature on M versus cM in the 1.28B base model (30 batches, 64 directions). Read-side slices are stiffer on M ; write-side slices are sloppier. Error bars are bootstrap 95% confidence intervals.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Llama 135M
Llama 1.28B
Llama 2.56B
Qwen3 1.7B
Massive coordinates ∣M∣
10
10
5
7
Maximum activation magnitude
280
2.27×103
3.92×103
2.23×103
Input availability:
Attention normalized input, X~
0.52 (-0.40)
0.52 (-0.24)
0.52 (-0.56)
0.76 (+0.19)
FFN normalized input, Y~
0.51 (-0.41)
0.80 (+1.00) *
0.65 (+0.93) *
0.80 (+0.68) *
Read side weight view:
Appendix
Table 4: Null occupancy results. Cells report the median ηM with rank-biserial rb in parentheses. Bold cells have rb>0.95 ; stars denote one-sided Mann–Whitney tests of ηM>η¬M with p<0.05 .
WV Frozen
WV Reparam
Wup,Wgate Frozen
Massive coordinates ∣M∣
7
10
11
Maximum activation magnitude
2.08×103
2.12×103
681
Input availability:
Attention normalized input, X~
0.51 (-0.52)
0.56 (-0.57)
0.61 (-0.17)
FFN normalized input, Y~
0.58 (+0.70) *
0.52 (-0.28)
0.57 (-0.17)
Read side weight view:
Appendix
Table 5: Llama 1.28B Erasure results for intervention settings. Bold cells have rb>0.95 ; stars denote one-sided Mann–Whitney tests of eraM>era¬M with p<0.05 .
WV Frozen
WV Reparam
Wup,Wgate Frozen
Massive coordinates ∣M∣
7
10
11
Maximum activation magnitude
2.08×103
2.12×103
681
Input availability:
Attention normalized input, X~
0.57 (+0.30)
0.54 (-0.13)
0.55 (+0.02)
FFN normalized input, Y~
0.89 (+1.00) *
0.80 (+1.00) *
0.68 (+0.77) *
Read side weight view:
Appendix
Table 6: Llama 1.28B Null occupancy results for intervention settings. Bold cells have rb>0.95 ; stars denote one-sided Mann–Whitney tests of eraM>era¬M with p<0.05 .
Property
Llama 135M
Llama 1.28B
Llama 2.56B
Qwen 1.7B
Number of Transformer layers ( L )
10
16
27
28
Residual-stream dimension ( dmodel )
768
2,048
2,560
2,048
Feed-forward dimension ( dff )
1,536
8,192
7,680
6,144
Number of query heads ( nh )
8
32
40
16
Number of key–value heads ( nkv )
8
32
40
16
Attention-head dimension ( dh )
96
64
64
128
Appendix
Table 7: Llama ( Touvron et al., 2023 ) and Qwen3 ( Team, 2025 ) Architectures of the models used in our experiments.
Figure 3: Emergence of M discrimination during pre-training of the 1.28B model (21 checkpoints, 0 – 20 k steps; per-coordinate maximum across 16 layers). Solid colored and dashed gray lines show medians and interquartile ranges for M ( n=10 ) and sampled cM controls ( n=100 ), respectively. The dotted line marks tsep , the first checkpoint with rank-biserial effect >0.95 . FFN read-blindness preceeds its amplifier.
Figure 4: FFN write-row concentration vs. amplifier gain (left) and s⋆ mass on M (right). Color denotes attention-input erasure, size denotes massiveness R , and grey points are non-MA coordinates.
Figure 5: Curvature and radial loss-gradient projection for Wdown[k,:] in deep layers L∈{9,12,15} . Negative ∇LCE⋅w^ means that the gradient vector points inward, so the loss-descent direction points outward. Strong-amplifier M coordinates concentrate in the negative-curvature, negative-gradient quadrant, unlike the matched non-massive coordinates.
Figure 6: Local first-order AdamW pressure on activation magnitude in the Llama 1.28B model. Lines show the batch mean and shaded regions show the batch 10th–90th percentile spread. Orange denotes M and dashed blue denotes the matched non-massive control. The full AdamW step increases the magnitude of M while leaving the control near zero. This effect is dominated by the saved optimizer state. The current gradient conditioned on that state is smaller, and weight decay acts in the opposite direction.
R=2
R=5
R=10
R=20
Massive coordinates ∣M∣
15
10
6
4
Maximum activation magnitude
2.27×103
2.27×103
2.27×103
2.27×103
Input availability:
Attention normalized input, X~
0.68 (-0.66)
0.68 (-0.73)
0.62 (-0.74)
0.61 (-0.77)
FFN normalized input, Y~
0.56 (-0.14)
0.56 (-0.13)
0.56 (-0.04)
0.56 (-0.12)
Read side weight view:
Appendix
Table 8: Erasure results at different MA selection thresholds for Llama 1.28B base model. Cells report the median eraM with rank-biserial rb in parentheses. Bold cells have rb>0.95 ; stars denote one-sided Mann–Whitney tests of eraM>era¬M with p<0.05 .
Trained transformers reliably develop massive activations, a small number of hidden dimensions whose magnitude is far above the median and which concentrate on the sequence-start token. Whether these outliers are a removable artifact of the residual stream's overloaded read and write role, or instead a functional necessity, is actively debated. We test the artifact hypothesis directly, with an architectural intervention. Our architecture, Ledger Residuals, splits the residual stream into a mutable scratch stream (Deliberation) that intermediate computation may freely overwrite and a protected, decode-only accumulator (Commitment) that holds the representation the model reads out. If massive activations exist only because one stream is forced to be both scratchpad and answer, then a dedicated answer channel should remove the need for them. We find that it does not. In matched-loss language models at the 160M and 290M scales, the model rebuilds the canonical fixed-dimension, start-token outlier inside the protected channel. The rebuilt feature is smaller in magnitude than in a standard transformer but more sharply concentrated on the start token, and a stronger sparsity penalty makes it more persistent and more concentrated still, rather than removing it. Massive activations therefore look architecturally robust: they re-emerge in whichever representation the model decodes from, which is what we would expect if they are functional rather than incidental. We release our architecture and measurement code.
Feedforward network (FFN) blocks account for a large fraction of the parameters and computation in Transformer architectures, yet their internal structure remains difficult to interpret due to the additive superposition induced by the residual stream. We examine whether the activation of an FFN neuron can be explained by a sparse set of preceding neuron activations and attention outputs. We introduce a training-free attribution method that estimates the relative influence of upstream neurons and attention outputs on a target neuron's activation. Empirically, across models and layers, we find that small subsets of preceding activations and attention outputs suffice to preserve neuron activations with high fidelity when all remaining inputs are masked with their average values. Effective sparsity is even greater when accounting for the inherent activation sparsity of upstream layers. Moreover, applying the neuron-specific masks in all layers simultaneously, such that the induced deviations propagate through the network, leaves model perplexity largely unchanged at moderate sparsity levels. These results demonstrate that, despite dense parameterization, FFNs exhibit sparse and structured inter-layer dependencies at the neuron level. Our method provides a practical, scalable tool for circuit-level interpretability and identifies candidate sparse pathways with potential implications for efficient inference.
A transformer can be built from operators that are legible by construction -- bounded, named units that read as fuzzy set operations rather than dense activations -- but legibility must be pressed for during training, and the pressure has a failure mode. A crispness penalty meant to sharpen a bounded operator into a decisive detector instead collapses it into a dead constant. An identity, E[v(1-v)] = mu(1-mu) - var, shows why -- the penalty is a variance-minimizer blind to the difference between a live detector and a constant -- and names the fix: a per-channel variance floor, the target legibility metric written as a loss, which recovers both legibility and quality. A learned per-unit fraction then retires the hand-set reserved-GELU partition of prior work: given the choice the model keeps no unit as pure GELU and routes 87% of its load-bearing computation through crisp operators. The result is the most legible transformer we have built -- 78% of its feed-forward operands and 50% of its attention value channels are crisp-and-contextual detectors, and per-head legibility rises from 18% in shallow layers to 78% in deep ones. Read in the correct rotated per-layer frame, these units separate a clean detection (what a unit responds to) from a harder naming (what its output decodes to); and because the objective makes each unit crisp and sparse, edits to them are far more local -- 50-184x in the deep layers where the edit sites concentrate -- and can target explicit conjunctions a single neuron cannot express. Finally, a between-unit decorrelation pressure exposes a legibility dial: it trades a circuit's reuse for independence at no quality cost, turning concepts into single, surgically editable units and a prediction into a short explanation read off a handful of named operations. Quality holds at parity with a conventional baseline throughout.
Mark Oskin
School of Computer Science and Engineering University of Washington