Large language models (LLMs) can sometimes report perturbations to their internal activations---even when the input provides no evidence that an intervention occurred. How do models detect and localize such internal changes? We study this question using a controlled task that keeps the input text fixed. We either inject a concept vector into the hidden state at one of ten token positions or apply no intervention. The model is asked to identify the perturbed position or report that no intervention occurred. Across three model families, we identify two small groups of attention heads with distinct roles in introspective reporting. Middle-layer gate heads influence whether the model reports a change, while router heads in a later layer help select the position to report. Interventions on gate heads can suppress position reports even when router heads supply location information. We further examine why reporting accuracy varies across concepts. Concept vectors that are localized more accurately produce stronger attention-score and output responses in gate heads, which is associated with better alignment of the induced key and value changes in their QK and OV computations. Together, these findings identify attention-head mechanisms supporting introspective detection and localization.
Figures & tables
Figure 1: Overview of the task and the two-component circuit. Left : the experimental pipeline. With the input text held fixed, either a concept vector is injected into the hidden state at one of ten candidate token positions or no intervention is applied. The model is asked to report the intervened position or answer none ; its answer is read from the logits at the first output position. Right : the two processes we identify. Gate heads influence whether the model reports a position or none , while router heads separate the ten injection positions from one another. Each point is one injected concept, and points sharing a color share the same injected position.
Ordered labels
Shuffled labels
Model
Digits
Letters
Words
Digits
Letters
Words
Mean
(a) Injected trials localization accuracy ↑
Qwen3-4B-IT
48.34
62.14
65.31
44.45
49.19
57.50
54.49
LLaMA-3.1-8B-IT
71.63
45.77
55.48
35.70
30.14
27.15
44.31
Gemma-3-12B-IT
84.16
70.46
82.11
62.37
57.63
38.96
65.95
(b) Clean trials none -response accuracy ↑
Table 1: Test accuracy across label sets and orderings. Scores are percentages; higher is better.
Figure 2: Layerwise patterns of report outcome and injection position. (a) Held-out injected validation trials, grouped by whether the model reports none or a position, projected onto the position–none direction, with normalization and density estimates given in Appendix A.6 . (b) Validation accuracy of K-means clusters matched to injection positions using training trials. Stars mark layers 24, 17, and 29; the dotted line marks the 10% chance level.
Figure 3: Effects of gate and router interventions (%, mean over six label settings). (a) Gate-off: none -response rate. (b) Gate-on: position-report rate. Outlined bars show the unmodified source run. (c, d) Combined gate and router patches in a clean (c) and an injected (d) run (Section 4.4 ); the legend gives the source of each head set. 95% intervals: Appendix A.17 . Details in Appendix A.18 .
Figure 4: Individual-head patching in LLaMA-3.1-8B-IT. Reduction in correct-position accuracy, Δh , in percentage points for each head.
Model
Output i
Output j
Other
Qwen3-4B-IT
2.9
38.0
1.1
LLaMA-3.1-8B-IT
3.2
37.8
4.1
Gemma-3-12B-IT
4.9
40.1
2.0
Table 2: Cross-position patching (% of trials, mean over six label settings). Gate outputs come from an injection at i , router outputs from j=i . Other: mean rate per remaining position. 95% intervals: Appendix A.17 . Details in Appendix A.18 .
Model
Size, QK σ1
Size, OV μk
Alignment, QK ∣v1⊤qI∣
Alignment, OV ∣cos(aI,xk)∣
Qwen3-4B-IT
1.22×
1.12 – 1.22×
1.80×
1.49 – 3.85×
LLaMA-3.1-8B-IT
1.08×
1.04 – 1.08×
1.20×
1.13 – 1.48×
Gemma-3-12B-IT
1.14×
1.09 – 1.18×
1.55×
1.37 – 2.90×
Table 3: Gate-head responses, Cintro relative to Cnonintro (ratio of group means; OV: range over the five leading modes; details in Appendix C.1 ).
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Total layers
Sweep range
Selected (layer, α )
Qwen3-4B-IT
36
0–35
(3, 3)
LLaMA-3.1-8B-IT
32
0–31
(0, 6)
Gemma-3-12B-IT
48
0–47
(0, 5)
Appendix
Table 4: Injection-layer and strength calibration. The extraction layer matches the injection layer. Strengths α∈{1,…,8} are tested at every layer in the sweep range; the selected setting is the best-ranked pair within this grid (correct-position accuracy, then mean correct-position probability, then lower α , then earlier layer).
Localization accuracy
Position reports
Model
Seed 701
Seed 702
Seed 703
Pooled
Injected
Clean
Qwen3-4B-IT
0.86
0.86
0.80
0.84
3.37
3.33
LLaMA-3.1-8B-IT
6.21
6.76
7.17
6.71
24.90
30.00
Gemma-3-12B-IT
0.00
0.00
0.00
0.00
4.59
6.67
Appendix
Table 5: Localization under norm-matched Gaussian injection (%). Pooled rates use all 180,000 injected trials per model. Position reports count any position answer, correct or not; on clean prompts they are false positives.
Ordered labels
Shuffled labels
Model
Replacement
Digits
Letters
Words
Digits
Letters
Words
Qwen3-4B-IT
Concept
1.00
1.81
3.12
1.17
1.43
3.75
Random
0.95
1.63
2.67
1.13
1.11
2.52
LLaMA-3.1-8B-IT
Concept
9.55
2.33
4.41
10.72
2.83
9.13
Random
9.35
2.15
4.49
11.13
3.38
9.95
Gemma-3-12B-IT
Concept
3.17
1.49
3.10
3.52
1.89
4.14
Appendix
Table 6: Localization after lexical replacement without activation injection. Accuracy (%) for corresponding concept words and random vocabulary words.
Figure 5: Selecting the STE mask cardinality on validation, against a random- k control. Solid coloured: the trained STE mask. Dashed grey: k heads picked at random within the same layer range the search covered (Qwen3-4B-IT 17–23, LLaMA-3.1-8B-IT 13–16, Gemma-3-12B-IT 21–28), ten draws per k , shaded with the 95% interval of their mean; nothing else changes. Gate on (red): a clean run gives a position report once the selected heads take their injected values. Gate off (blue): an injected run outputs none once they are restored to their clean values. Dotted: the unmodified reference rate. At k=0 no head is modified, so the two curves meet.
Model
Layer
Gate-on heads
Gate-off heads
Qwen3-4B-IT
17
15
5
18
–
14
19
8, 10 , 19, 23
0, 10 , 23 , 25
20
2 , 3 , 8 , 10 , 15 , 23, 29
2 , 3 , 5, 6, 8 , 10 , 11, 15 , 29
21
0 , 9, 11, 13, 15, 16, 18, 19 , 24
0 , 19
22
4 , 5 , 7 , 10 , 11 , 13, 15, 17 , 27
0, 1, 4 , 5 , 7 , 10 , 11 , 17 , 18, 25, 28
Appendix
Table 7: Gate heads selected by the Top-32 STE masks. Each cell lists the head indices selected in that layer; “–” means none. Bold heads are selected by both the gate-on and the gate-off mask.
Figure 6: Head-wise activation patching. Accuracy drop Δh=Accinj−Accpatch(h) (percentage points) for every attention head, obtained by patching the head’s final-position output in the injected run with its clean-run value. Red cells lower correct-index accuracy; annotated heads are the router heads.
Ordered labels
Shuffled labels
Model
Outcome
Digits
Letters
Words
Digits
Letters
Words
Mean
Qwen3-4B-IT
→j
0.480
0.679
0.704
0.582
0.585
0.760
0.632
→i
0.011
0.011
0.029
0.004
0.018
0.023
0.016
Other
0.004
0.003
0.022
0.013
0.031
0.065
0.023
none
0.505
0.307
0.245
0.401
0.366
0.152
0.329
LLaMA-3.1-8B-IT
→j
0.759
0.530
0.680
0.829
0.676
0.725
0.700
Appendix
Table 8: Cross-position attention redirection: output proportions. For each label arm (the digit, letter, and word-name label sets, each under the identity and the shuffled permutation) and each of the 90 injection–readout position pairs with i=j , we set the selected router heads’ final-position attention to a one-hot distribution on the successor token of position j and record where the output lands: the redirected index ( →j ), the original injection index ( →i ), any other index (Other), or none . Proportions are pooled over all 90 pairs per arm (270,000 trials) and sum to 1 in each column. Router heads: Qwen3-4B-IT L24 H{29,31}; LLaMA-3.1-8B-IT L17 H{24}; Gemma-3-12B-IT L29 H{1,11}. The Mean column averages each outcome’s proportion over the six label arms. 95% intervals are in Appendix A.17 .
Figure 7: Injection-induced attention changes across model families. All panels show injected-minus-clean attention from the final prompt token T : red indicates an increase and blue a decrease, with darker colors indicating larger changes in magnitude. The blue outline marks the query token. Color scales are set per panel. The following TOKEN marks the successor position of the injected candidate. Prompt text is shown with each model’s chat-template delimiters.
Figure 8: Final-token residual PCA around the transition layer in all three models. Each colored point is one injected run at the final prompt token, colored by the perturbed candidate position; the black point is the clean run. Principal components are fitted separately for each panel. Panels are ordered transition layer minus one, transition layer, transition layer plus one, and the last layer.
Figure 9: Three-dimensional PCA of attention-head outputs across models. Each panel shows one attention head; red frames highlight the selected router heads. Colors distinguish injected indices, and black points mark clean references. Only correctly answered injected runs are shown.
Figure 10: Three-dimensional PCA of attention-head outputs across models (continued). Qwen3-4B-IT, heads 16–31. Colors and red frames follow panel (a).
Figure 11: Three-dimensional PCA of attention-head outputs across models (continued). LLaMA-3.1-8B-IT, heads 0–15. Colors and red frames follow panel (a).
Figure 12: Three-dimensional PCA of attention-head outputs across models (continued). LLaMA-3.1-8B-IT, heads 16–31. Colors and red frames follow panel (a).
Figure 13: Three-dimensional PCA of attention-head outputs across models (continued). Gemma-3-12B-IT. Colors and red frames follow panel (a).
Qwen3-4B-IT
LLaMA-3.1-8B-IT
Gemma-3-12B-IT
Condition
%
95% CI
%
95% CI
%
95% CI
(a) Gate off: none rate
Injected run
37.7
[37.5, 37.9]
25.3
[25.1, 25.5]
23.8
[23.6, 24.0]
Gate patch
81.5
[81.4, 81.7]
60.7
[60.5, 60.9]
80.7
[80.5, 80.9]
Target (clean run) ∗
89.4
[84.1, 93.1]
62.8
[55.5, 69.5]
87.8
[82.2, 91.8]
(b) Gate on: position rate
Appendix
Table 9: 95% intervals for Figure 3 . Plotted rate (%) and its Wilson interval. Each rate is the mean over the six label settings of Table 1 , which equals the rate pooled over them. Unless marked, a rate pools 180,000 trials (six settings × 100 concepts × 30 prompts × 10 positions). ∗ Unmodified clean run: 180 prompts (30 per setting). † 1,800,000 trials (all 100 gate–router position pairs in each setting). Bold : the gate heads alone are patched.
Qwen3-4B-IT
LLaMA-3.1-8B-IT
Gemma-3-12B-IT
Outcome
%
Wilson
Pos. boot.
%
Wilson
Pos. boot.
%
Wilson
Pos. boot.
Output i
2.9
[2.9, 2.9]
[2.2, 3.7]
3.2
[3.2, 3.2]
[1.5, 5.8]
4.9
[4.9, 5.0]
[1.2, 10.3]
Output j
38.0
[37.9, 38.1]
[32.3, 42.6]
37.8
[37.7, 37.9]
[35.4, 39.4]
40.1
[40.0, 40.1]
[31.0, 47.4]
Other
1.1
[1.1, 1.1]
[0.9, 1.3]
4.1
[4.1, 4.1]
[3.6, 4.4]
2.0
[2.0, 2.0]
[1.5, 2.3]
Appendix
Table 10: 95% intervals for Table 2 . Rate (%) averaged over the six label settings of Table 1 , each pooling the 90 ordered pairs i=j (1,620,000 trials in total), with two intervals: Wilson , a score interval over trials, and Pos. boot. , a bootstrap that resamples the ten gate-source positions i together with all of their pairs, shared across the six settings. Other is the average rate per position other than i and j , and its intervals are those of the eight-position total divided by eight. Output j exceeds output i and Other under both intervals in every model.
Ordered labels
Shuffled labels
Outcome
Digits
Letters
Words
Digits
Letters
Words
Mean
Qwen3-4B-IT
→j
0.480 [0.361, 0.587]
0.679 [0.560, 0.776]
0.704 [0.624, 0.769]
0.582 [0.491, 0.652]
0.585 [0.532, 0.631]
0.760 [0.734, 0.784]
0.632 [0.552, 0.695]
→i
0.011 [0.001, 0.027]
0.011 [0.005, 0.019]
0.029 [0.016, 0.045]
0.004 [0.001, 0.008]
0.018 [0.012, 0.024]
0.023 [0.016, 0.030]
0.016 [0.013, 0.020]
Other
0.004 [0.002, 0.005]
0.003 [0.002, 0.004]
0.022 [0.014, 0.031]
0.013 [0.009, 0.017]
0.031 [0.025, 0.035]
0.065 [0.053, 0.076]
0.023 [0.019, 0.027]
none
0.505 [0.396, 0.627]
0.307 [0.209, 0.427]
0.245 [0.176, 0.333]
0.401 [0.330, 0.494]
0.366 [0.315, 0.427]
0.152 [0.129, 0.179]
0.329 [0.263, 0.413]
Appendix
Table 11: 95% position-bootstrap intervals for Table 8 . Each cell gives the proportion from Table 8 and, below it, an interval from resampling the ten injection positions i with all of their readout positions j=i ; the Mean column resamples positions jointly across the six arms. Wilson half-widths are at most 0.002 for every a score interval over trials, and arm and at most 0.001 for the Mean.
Figure 14: Gate and router interventions with ordered labels. Panels as in Figure 3 ; gate and router heads are the ones selected under ordered digit labels.
Figure 15: Gate and router interventions with shuffled labels. Panels as in Figure 3 ; gate and router heads are the ones selected under ordered digit labels.
Table 12: Cross-position patching under each label setting. Setup as in Table 2 , whose values are the unweighted mean of these six tables. Values are percentages of all trials: output i , output j , Other (the average rate per position other than i and j , as in Table 2 ), and none .
KL: − query / − key
Correct-position rate
Model
Concept set
Successor
Full context
Native
− query / − key
Qwen3-4B-IT
Cintro
0.0197 / 0.3108
0.0351 / 0.1581
48.28%
43.62% / 32.07%
Cnonintro
0.0025 / 0.0384
0.0043 / 0.0300
1.67%
1.49% / 0.96%
LLaMA-3.1-8B-IT
Cintro
0.0093 / 0.0585
0.0134 / 0.0411
69.34%
67.27% / 65.38%
Cnonintro
0.0041 / 0.0322
0.0079 / 0.0266
20.97%
19.03% / 18.09%
Gemma-3-12B-IT
Cintro
0.0519 / 0.5847
0.1205 / 0.3979
83.67%
83.07% / 58.23%
Appendix
Table 13: Ablating the key term versus the query term. Attention-distribution change and reporting accuracy under each ablation, by model and concept set.
Concept set
Query norm ∥qI∥
Magnitude σ1
Readout ∣v1⊤qI∣
First mode R1
Target loading Uˉi1
Target response sˉi
(a) Qwen3-4B-IT
Cintro
14.2870
9.7021
2.0431
1.8697
0.8847
1.7000
Cnonintro
14.2356
7.9384
1.1333
0.8099
0.3832
0.3203
Intro / Non-intro
1.00×
1.22×
1.80×
2.31×
2.31×
5.31×
(b) LLaMA-3.1-8B-IT
Cintro
12.5969
9.4412
1.1510
0.9493
0.6945
0.6774
Appendix
Table 14: QK responses of introspective and non-introspective concepts. Group statistics and between-set ratios in gate heads; the ratio rows extend the QK columns of Table 3 .
Attention
Change matrix M
Output MaI
Concept set
∥aI∥
∥Δa∥
∥M∥F
η5
∥MaI∥
Top-5 norm
ρ5
(a) Qwen3-4B-IT
Cintro
0.367
0.085
19.3
80.1
0.912
0.841
68.7
Cnonintro
0.360
0.036
16.6
82.0
0.313
0.236
49.3
Intro / Non-intro
1.02×
2.40×
1.16×
0.98×
2.92×
3.56×
1.39×
(b) LLaMA-3.1-8B-IT
Appendix
Table 15: Attention and output statistics of the ΔV term in gate heads. Group means over the full context; energies are percentages.
Size μk
Alignment ∣cos(aI,xk)∣
Concept set
k=1
2
3
4
5
k=1
2
3
4
5
(a) Qwen3-4B-IT
Cintro
11.3
8.24
6.45
5.33
4.48
0.149
0.126
0.104
0.097
0.087
Cnonintro
10.1
7.05
5.47
4.43
3.68
0.039
0.059
0.059
0.061
0.058
Intro / Non-intro
1.12×
1.17×
1.18×
1.20×
1.22×
3.85×
2.16×
1.76×
1.58×
1.49×
(b) LLaMA-3.1-8B-IT
Appendix
Table 16: Size and alignment of the five leading modes of M in gate heads. Group means over the full context.
Terms, 30 clusters
Modes of M , 30 clusters
Model
Set
Injected
−Δa
−ΔV
− both
Injected
− top 5
− others
−ΔV
Qwen3-4B-IT
Cintro
48.3
38.8
4.4
1.4
48.3
5.1
44.6
4.4
Cnonintro
1.7
1.3
0.7
0.7
1.7
0.7
1.5
0.7
LLaMA-3.1-8B-IT
Cintro
69.3
67.8
54.3
49.8
69.3
57.4
67.1
54.3
Cnonintro
21.0
19.3
13.8
11.9
21.0
15.3
19.0
13.8
Gemma-3-12B-IT
Cintro
83.7
74.1
13.9
2.6
83.7
13.5
81.2
13.9
Appendix
Table 17: Correct-position rate (%) after removing parts of the gate-head output in the injected run. −Δa and −ΔV remove the corresponding terms of Equation 1 ; − top 5 and − others remove the five leading modes of M or all remaining modes.
Model
Set
Injected
− top 5
− others
−ΔV
Qwen3-4B-IT
Cintro
47.99
5.12
44.39
4.40
Cnonintro
1.66
0.67
1.47
0.65
LLaMA-3.1-8B-IT
Cintro
40.22
33.69
39.95
32.93
Cnonintro
13.37
11.80
12.96
11.32
Gemma-3-12B-IT
Cintro
83.12
13.53
80.63
13.91
Cnonintro
18.04
1.83
15.15
1.43
Appendix
Table 18: Mean correct-position token probability (%) after OV mode ablation. Full-context modes, all 30 clusters, and 30,000 paired trials per concept set and model. Probabilities use the full vocabulary.
Model
Set
Injected
− top 1
− top 2
− top 5
− top 10
−ΔV
(a) Mean correct-position token probability
Qwen3-4B-IT
Cintro
47.99
24.89
10.28
5.12
4.42
4.40
Cnonintro
1.66
1.06
0.78
0.67
0.64
0.65
LLaMA-3.1-8B-IT
Cintro
40.22
38.57
36.64
33.69
33.16
32.93
Cnonintro
13.37
13.16
12.82
11.80
11.46
11.32
Gemma-3-12B-IT
Cintro
83.12
54.86
27.82
13.53
14.24
13.91
Appendix
Table 19: OV ablation as a function of the number of removed modes. Modes of the full-context matrix M are ordered by singular value. Each entry uses all 30 clusters and 30,000 trials per concept set and model. All values are percentages; probabilities use the full-vocabulary softmax. −ΔV removes the entire content-change term.