Finding the Heads and the Neurons Responsible for Network Information Retrieval in Language Models
Authors: Abdul Kadir, Md Mohasin Hossain, Daniel Sonntag
Organizations: University of Oldenburg, Oldenburg, Germany · German Research Center for Artificial Intelligence · Saarland University, Saarbrucken, Germany
We ask whether specific attention heads, and more finely specific neurons inside those heads, are responsible for recognizing that a language model's context contains network infrastructure information (a hostname paired with its IP address), and whether that responsibility can be validated causally rather than by correlation alone. At the head level the answer is yes, across five models spanning three architecture families: in every model, a small set of heads (1 to 9 out of 128 to 1152 candidates), found by causal ablation screening and tested for selectivity against matched negative and context-free controls, supports a detector with 99.5--100% held-out accuracy. We then ask whether a head's responsibility concentrates into one neuron or stays spread across its dimensions; this is model-specific. In one model, the top head's signal concentrates into a single neuron, found independently by both a causal intervention and a correlational ranking, which agree exactly (AUC = 1.000, matching the full head). In another, the single clean head works as a whole (AUC = 1.000) but the best causally ranked neuron inside it does not (AUC = 0.665), so the responsibility there is spread across the head. The remaining three models fall in between. On an independent dataset collected by a different institution (reverse-DNS records rather than the discovery data), every model's full-head detector flags 100% of positive records; the single-neuron versions transfer less reliably, and in one model score below chance. Causal head-finding for a specific network-information entity works across models and architectures; how far that finding can be pushed down to individual neurons varies, and needs to be checked for each model.
Figures & tables
Figure 1: Pipeline overview. Two independent algorithms discover candidate heads and are cross-validated against each other; the surviving clean_subset heads then feed a second, structurally identical cross-validation at neuron granularity (correlational information-gain ranking vs. causal ablation), whose agreement or disagreement is this paper’s central result (Section 4.2 ).
Condition
Context
Question
True answer
P(internet)
Flagged
network
record
What is the IP address of this website?
62.151.181.154
1.000
yes
paraphrase
record
Resolve the IPv4 address associated with this domain.
62.151.181.154
1.000
yes
entity
record
What is the category of this website? Answer with one word: benign, mal, or phishing.
benign
0.000
no
vocabulary
record
In one short phrase, what does the abbreviation “IP” stand for?
Internet Protocol
0.000
no
ordinary
empty
What is 67 + 14?
81
0.001
no
question_only
empty
What is the IP address of this website?
62.151.181.154
0.024
no
Table 1: A worked example: one held-out record under every condition, with the predicted probability from the fixed Qwen3.5-4B detector. The network , paraphrase , entity and vocabulary rows share the same context ( al-mostafa.com , 62.151.181.154 ); the other rows have an empty context.
Figure 2: Where the interventions sit inside one attention layer. Every head’s output Oh (Eq. ( 2 )) is an additive term (Eq. ( 3 )); ablating head h zeroes its whole term, while ablating neuron j zeroes only one column of Oh . The forward-pre-hook on WO ’s input ( HeadSuppressor / NeuronSuppressor ) implements both.
Model
AUC
question_only flag rate
(both)
HE (20)
clean_subset
Qwen3.5-4B
1.000
1.00
0.00
Qwen3-8B
1.000
1.00
0.00
Qwen3.5-9B
1.000
1.00
0.00
Gemma-4-E4B
1.000
1.00
1.00
Gemma-4-E2B
1.000
1.00
1.00
Table 2: Unfiltered HE vs. filtered clean_subset : held-out AUC is identical either way, but the context-free wording confound ( question_only flag rate) differs: the filter lowers it to 0 only in the Qwen models.
Model
Cand.
Clean
Acc.
AUC
heads
heads
(held-out)
Qwen3.5-4B
128
9
100.0%
1.000
Qwen3-8B
1152
1
99.5%
1.000
Qwen3.5-9B
128
4
100.0%
1.000
Gemma-4-E4B
280
5
99.9%
1.000
Gemma-4-E2B
224
2
99.9%
1.000
Table 3: A causally-validated, selectivity-confirmed head set exists in every model tested, regardless of architecture or candidate-pool size.
Single neuron
Full head
Model
Acc.
AUC
Acc.
AUC
Qwen3.5-4B
99.4%
1.000
100.0%
1.000
Qwen3-8B
53.9%
0.665
99.5%
1.000
Qwen3.5-9B
99.4%
1.000
100.0%
1.000
Gemma-4-E4B
99.9%
1.000
99.9%
1.000
Gemma-4-E2B
86.6%
0.969
99.9%
1.000
Table 4: Single-neuron (best causally ranked neuron) vs. full-head detection performance (held-out accuracy at the 0.5 threshold, and AUC).
Model
Causal top / IG top
Agreement
Qwen3.5-4B
(15,1,0)/(15,1,0)
exact match
Qwen3-8B
(29,12,82)/(29,12,27)
same head, diff. dim
Qwen3.5-9B
(23,15,8)/(19,4,1)
different head
Gemma-4-E4B
(32,2,141)/(32,2,7)
same head, diff. dim
Gemma-4-E2B
(30,0,183)/(30,0,48)
same head, diff. dim
Table 5: Top-ranked neuron (layer, head, dimension) under the causal ranking and under the information-gain (IG) ranking.
Model
Actual
Random (same head)
Random (any neuron)
Qwen3.5-4B
1.000
0.727 ± 0.155
0.667 ± 0.186
Qwen3-8B
0.665
0.633 ± 0.148
0.728 ± 0.179
Qwen3.5-9B
1.000
0.680 ± 0.128
0.654 ± 0.195
Gemma-4-E4B
1.000
0.748 ± 0.180
0.744 ± 0.195
Gemma-4-E2B
0.969
0.675 ± 0.149
0.724 ± 0.188
Table 6: AUC of the actual single_best_causal_neuron classifier vs. two random-neuron baselines (mean ± std over 20 repeats each).
Model
Full head
Pos.
Neg.
Full head
Single neuron
acc.
acc.
acc.
AUC
AUC
Qwen3.5-4B
99.9%
100.0%
99.8%
1.000
0.535
Qwen3-8B
99.8%
100.0%
99.6%
1.000
0.422
Qwen3.5-9B
100.0%
100.0%
100.0%
1.000
0.285
Gemma-4-E4B
91.9%
100.0%
83.8%
0.995
0.866
Gemma-4-E2B
91.5%
100.0%
83.0%
0.971
0.463
Table 7: Cross-validation on an independent, real-world dataset (CAIDA). The full-head detector transfers with high AUC; the single-neuron detector does not. Pos. and Neg. are the accuracies on the CAIDA records and on the trivia negatives.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Condition
Qwen3.5-4B
Qwen3-8B
Qwen3.5-9B
Gemma-4-E4B
Gemma-4-E2B
network
1.000*
0.999*
0.999*
1.000*
0.975*
paraphrase
1.000*
1.000*
0.999*
1.000*
0.988*
entity
0.000
0.000
0.000
0.000
0.009
vocabulary
0.000
0.004
0.001
0.000
0.008
ordinary
0.001
0.002
0.001
0.000
0.002
question_only
0.024
0.006
0.111
0.996*
0.982*
Appendix
Table 8: Worked example, ask_ip direction: record al-mostafa.com / 62.151.181.154 , question “What is the IP address of this website?” (answer 62.151.181.154 for network / paraphrase ). Cells are P(internet) from each model’s own fixed detector; * marks flagged .
Condition
Qwen3.5-4B
Qwen3-8B
Qwen3.5-9B
Gemma-4-E4B
Gemma-4-E2B
network
1.000*
1.000*
1.000*
0.999*
1.000*
paraphrase
1.000*
1.000*
1.000*
0.999*
0.988*
entity
0.000
0.000
0.000
0.000
0.025
vocabulary
0.000
0.004
0.001
0.001
0.009
ordinary
0.001
0.004
0.001
0.000
0.004
question_only
0.035
0.011
0.117
0.998*
0.999*
Appendix
Table 9: Worked example, ask_url direction: record jjeducare.co.kr / 211.169.73.24 , question “What is the website address for this IP?” (answer jjeducare.co.kr for network / paraphrase ). Same layout as Table 8 .
Model
DNS-TTL flag rate
DNS-TTL mean P
MAC flag rate
MAC mean P
Qwen3.5-4B
0.990
0.884
1.000
0.994
Qwen3-8B
0.365
0.471
0.985
0.931
Qwen3.5-9B
0.440
0.576
1.000
0.993
Gemma-4-E4B
0.000
0.066
1.000
0.986
Gemma-4-E2B
0.020
0.129
0.275
0.355
Appendix
Table 10: Generalization to two new network-entity sub-types, never seen during discovery or training: 200 held-out records per model, each fixed detector unmodified (same classifier as Section 4 ). Flag rate is the fraction of the 200 items the detector labels network-positive; mean P is the mean predicted probability.
Model
Condition
N (clean)
N (random)
S (clean)
S (random)
Qwen3.5-4B
network
0.103
− 0.003
9.35
− 1.03
Qwen3.5-4B
dns_ttl
0.001
0.000
10.02
− 0.03
Qwen3.5-4B
mac
0.013
0.000
10.79
− 0.83
Qwen3-8B
network
0.000
0.000
1.22
0.06
Qwen3-8B
dns_ttl
0.000
0.000
0.42
0.23
Qwen3-8B
mac
0.000
0.000
1.13
0.12
Appendix
Table 11: Joint necessity ( N ) and sufficiency ( S ), in nats of log-probability, of clean_subset versus a random-head control of equal size, for the true IP/URL (network), TTL (dns_ttl), or MAC (mac) answer. S is the cleaner signal: sufficiency compares patching clean_subset ’s activations from the target condition into an unrelated entity prompt against patching a same-size random-head set the same way.
Genre
Should flag
Qwen3.5-4B
Qwen3-8B
Qwen3.5-9B
Gemma-4-E4B
Gemma-4-E2B
server log
yes
0.999
1.000
1.000
1.000
0.996
nginx config
yes
0.999
0.999
0.999
0.995
0.883
security incident report
yes
0.999
1.000
0.998
1.000
0.997
IT support email
yes
0.999
0.905
0.999
0.996
1.000
casual user chat
yes
1.000
0.993
1.000
0.999
1.000
server log (non-retrieval)
no
0.028
0.029
0.365
0.927 †
0.944 †
Appendix
Table 12: Use-case stress test: P(internet) from each model’s own fixed clean_subset detector on ten naturalistic items never seen during discovery or training, five positives, one per genre (real address + retrieval question), and five matched negatives. A dagger † (also in bold) marks a flag that disagrees with the intended label.
Attention-head ablation, zeroing a head and measuring the resulting change in task performance, is a common method for inferring which components of a language model are causally responsible for a behavior. We show using GPT-2 small that this inference can be fragile unless the intervention semantics, evaluation metric, and controls are carefully validated. A natural post-projection implementation of "zeroing a head" is nearly uncorrelated with a corrected pre-projection ablation (Pearson r = 0.057) and selects a completely disjoint top-5 set of important heads. We also show that binary accuracy can hide effects at behavioral floors and near ceilings, whereas gold-token log-probability remains graded. Using a discovery/held-out split and 1,000 matched random-head and layer-matched-head control draws, the corrected per-head effect ranking is highly stable across splits (Spearman rho = 0.974), and the top-5 selected heads significantly exceed both control distributions (Monte Carlo p = 0.001). However, evidence for task specificity is not robust on GPT-2. Replication on DistilGPT2 preserves the intervention-semantic and matched-control findings. These results show that single-head ablation does not by itself justify a causal claim; defensible interpretation requires correct intervention placement, a non-saturated continuous metric, and matched held-out controls.
Large language models (LLMs) can sometimes report perturbations to their internal activations---even when the input provides no evidence that an intervention occurred. How do models detect and localize such internal changes? We study this question using a controlled task that keeps the input text fixed. We either inject a concept vector into the hidden state at one of ten token positions or apply no intervention. The model is asked to identify the perturbed position or report that no intervention occurred. Across three model families, we identify two small groups of attention heads with distinct roles in introspective reporting. Middle-layer gate heads influence whether the model reports a change, while router heads in a later layer help select the position to report. Interventions on gate heads can suppress position reports even when router heads supply location information. We further examine why reporting accuracy varies across concepts. Concept vectors that are localized more accurately produce stronger attention-score and output responses in gate heads, which is associated with better alignment of the induced key and value changes in their QK and OV computations. Together, these findings identify attention-head mechanisms supporting introspective detection and localization.
Jiahong Zou, Xiangkun Sun, Lingkai Kong +1
Shandong University · Tsinghua University · Northeastern University +1
Contextualization, the core operation of language modeling, transmits information across words to build sentence-specific word representations. Prior works mainly study contextualization, focusing on individual words and attention heads as a growing discrete dictionary, lacking a global view of their general behavior. Therefore, we propose a general principle: Globally, we find and estimate that different words carry different amounts of information, and less-informative words tend to absorb more contextual information. Specifically, these low-information words do not absorb contextual words uniformly, and finer-grained selectivity enables more precise routing to promote information transmission between matched words. Moreover, to find what mechanism causes such processing, we reinterpret attention heads as channels gated by their singular vectors and find that: (1) these singular vectors point to the hidden states of more informative words, allowing such words to write their information to others more strongly to act as information sources, and vice versa; and (2) these singular vectors can be viewed equally as hidden state features, enabling automated interpretation of attention heads beyond prior heuristic head discovery, also embedding heads into a continuous space rather than treating them as discrete, independent dictionary entries.
Hakaze Cho, Haolin Yang, Zhun Sun +3
RIKEN · Tohoku University · New York University +2