Massive activation features (MAs) in Transformers are extreme-value residual-stream features that persist across layers despite the model's ability to suppress them. Why do they survive? Our investigation using an operator-level mechanistic analysis of attention and feed-forward (FFN) blocks reveals that these blocks systematically ignore MA coordinates while reading, but not while writing; creating a read-write asymmetry that blocks corrective feedback while allowing continued accumulation. We find that both attention and feed-forward layers have this read-blindness, and contribute to the emergence and persistence of MAs. To validate prior work that hypothesized that FFN's amplification abilities is the primary reason for MAs (Sun et al., 2026), we analyze the model checkpoints during learning. Contrary to our expectation, read-blindness emerges before FFN amplification, suggesting that it acts upstream in the MA mechanism. We further contribute gradient analysis to link this behavior to surprising asymmetries in the loss landscape, concluding that the model actively maintains this read-blindness. Finally, we find that removing read-blocking at different locations induces compensatory shifts elsewhere, but MAs still persist.
Figures & tables
Llama 135M
Llama 1.28B
Llama 2.56B
Qwen3 1.7B
Massive coordinates ∣M∣
10
10
5
7
Maximum activation magnitude
280
2.27×103
3.92×103
2.23×103
Input availability:
Attention normalized input, X~
0.58 (-0.60)
0.68 (-0.73)
0.55 (-0.64)
0.58 (-0.37)
FFN normalized input, Y~
0.56 (-0.65)
0.56 (-0.13)
0.50 (-0.63)
0.54 (-0.48)
Read side weight view:
Table 1: Erasure results. Cells report the median eraM with rank-biserial rb against the full complement ¬M in parentheses. Bold cells have rb>0.95 ; stars denote one-sided Mann–Whitney tests of eraM>era¬M with p<0.05 .
Signed alignment
Llama 135M
Llama 1.28B
Llama 2.56B
Qwen3 1.7B
Attention, Xk with Δkattn
+0.03 (+0.45)
+0.09 (+0.29)
+0.02 (-0.27)
+0.06 (+0.16)
FFN, Yk with Δkffn
+0.05 (+0.57)
+0.25 (+0.59)
-0.02 (-0.20)
+0.10 (+0.14)
Table 2: Signed-alignment results. For each coordinate, S is median-aggregated across layers. Cells report the median SM with rank-biserial rb comparing M and ¬M in parentheses.
Diagnostic
value M
rb
∥Uk∥F
7.74
+0.78
∣λ⋆(k)∣
2.19
+0.98
∣λ⋆(k)∣/∥Sk∥F
0.39
+1.00
Within-set ∣cos(s⋆)∣
0.53/0.46
—
Table 3: s⋆ amplifier diagnostics for Llama 1.28 Base. Per-coordinate entries are medians over M and a deterministic, size-matched ¬M ; the final row is the layer-mean of the within-set median pairwise absolute cosine. rb>0 means larger values on M .
Figure 1: FFN read-blindness precedes amplifier specialization. We assign massive and control coordinates at the final checkpoint and trace these fixed sets through 21 checkpoints from 0 to 20 k training steps. (a) Maximum-over-layer FFN erasure for eventual massive coordinates (orange) and non-massive controls (gray). (b) Maximum-over-layer leading amplifier gain for the same sets. Lines and shaded regions show medians and interquartile ranges; vertical dotted lines mark tsep , the first checkpoint at which the rank-biserial effect reaches rb≥0.95 . FFN erasure separates at 2.3 k steps, before amplifier gain at 4.5 k steps. (c) Final-checkpoint amplifier norm ∥Uk∥F versus the inverse participation ratio of the corresponding FFN write row Wdown[k,:] . Colored points are massive coordinates, with size denoting massiveness and color denoting attention-input erasure; gray points are non-massive coordinates.
Figure 2: Directional curvature on M versus cM in the 1.28B base model (30 batches, 64 directions). Read-side slices are stiffer on M ; write-side slices are sloppier. Error bars are bootstrap 95% confidence intervals.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Llama 135M
Llama 1.28B
Llama 2.56B
Qwen3 1.7B
Massive coordinates ∣M∣
10
10
5
7
Maximum activation magnitude
280
2.27×103
3.92×103
2.23×103
Input availability:
Attention normalized input, X~
0.52 (-0.40)
0.52 (-0.24)
0.52 (-0.56)
0.76 (+0.19)
FFN normalized input, Y~
0.51 (-0.41)
0.80 (+1.00) *
0.65 (+0.93) *
0.80 (+0.68) *
Read side weight view:
Appendix
Table 4: Null occupancy results. Cells report the median ηM with rank-biserial rb in parentheses. Bold cells have rb>0.95 ; stars denote one-sided Mann–Whitney tests of ηM>η¬M with p<0.05 .
WV Frozen
WV Reparam
Wup,Wgate Frozen
Massive coordinates ∣M∣
7
10
11
Maximum activation magnitude
2.08×103
2.12×103
681
Input availability:
Attention normalized input, X~
0.51 (-0.52)
0.56 (-0.57)
0.61 (-0.17)
FFN normalized input, Y~
0.58 (+0.70) *
0.52 (-0.28)
0.57 (-0.17)
Read side weight view:
Appendix
Table 5: Llama 1.28B Erasure results for intervention settings. Bold cells have rb>0.95 ; stars denote one-sided Mann–Whitney tests of eraM>era¬M with p<0.05 .
WV Frozen
WV Reparam
Wup,Wgate Frozen
Massive coordinates ∣M∣
7
10
11
Maximum activation magnitude
2.08×103
2.12×103
681
Input availability:
Attention normalized input, X~
0.57 (+0.30)
0.54 (-0.13)
0.55 (+0.02)
FFN normalized input, Y~
0.89 (+1.00) *
0.80 (+1.00) *
0.68 (+0.77) *
Read side weight view:
Appendix
Table 6: Llama 1.28B Null occupancy results for intervention settings. Bold cells have rb>0.95 ; stars denote one-sided Mann–Whitney tests of eraM>era¬M with p<0.05 .
Property
Llama 135M
Llama 1.28B
Llama 2.56B
Qwen 1.7B
Number of Transformer layers ( L )
10
16
27
28
Residual-stream dimension ( dmodel )
768
2,048
2,560
2,048
Feed-forward dimension ( dff )
1,536
8,192
7,680
6,144
Number of query heads ( nh )
8
32
40
16
Number of key–value heads ( nkv )
8
32
40
16
Attention-head dimension ( dh )
96
64
64
128
Appendix
Table 7: Llama ( Touvron et al., 2023 ) and Qwen3 ( Team, 2025 ) Architectures of the models used in our experiments.
Figure 3: Emergence of M discrimination during pre-training of the 1.28B model (21 checkpoints, 0 – 20 k steps; per-coordinate maximum across 16 layers). Solid colored and dashed gray lines show medians and interquartile ranges for M ( n=10 ) and sampled cM controls ( n=100 ), respectively. The dotted line marks tsep , the first checkpoint with rank-biserial effect >0.95 . FFN read-blindness preceeds its amplifier.
Figure 4: FFN write-row concentration vs. amplifier gain (left) and s⋆ mass on M (right). Color denotes attention-input erasure, size denotes massiveness R , and grey points are non-MA coordinates.
Figure 5: Curvature and radial loss-gradient projection for Wdown[k,:] in deep layers L∈{9,12,15} . Negative ∇LCE⋅w^ means that the gradient vector points inward, so the loss-descent direction points outward. Strong-amplifier M coordinates concentrate in the negative-curvature, negative-gradient quadrant, unlike the matched non-massive coordinates.
Figure 6: Local first-order AdamW pressure on activation magnitude in the Llama 1.28B model. Lines show the batch mean and shaded regions show the batch 10th–90th percentile spread. Orange denotes M and dashed blue denotes the matched non-massive control. The full AdamW step increases the magnitude of M while leaving the control near zero. This effect is dominated by the saved optimizer state. The current gradient conditioned on that state is smaller, and weight decay acts in the opposite direction.
R=2
R=5
R=10
R=20
Massive coordinates ∣M∣
15
10
6
4
Maximum activation magnitude
2.27×103
2.27×103
2.27×103
2.27×103
Input availability:
Attention normalized input, X~
0.68 (-0.66)
0.68 (-0.73)
0.62 (-0.74)
0.61 (-0.77)
FFN normalized input, Y~
0.56 (-0.14)
0.56 (-0.13)
0.56 (-0.04)
0.56 (-0.12)
Read side weight view:
Appendix
Table 8: Erasure results at different MA selection thresholds for Llama 1.28B base model. Cells report the median eraM with rank-biserial rb in parentheses. Bold cells have rb>0.95 ; stars denote one-sided Mann–Whitney tests of eraM>era¬M with p<0.05 .