Finding the Heads and the Neurons Responsible for Network Information Retrieval in Language Models
Authors: Abdul Kadir, Md Mohasin Hossain, Daniel Sonntag
Organizations: University of Oldenburg, Oldenburg, Germany · German Research Center for Artificial Intelligence · Saarland University, Saarbrucken, Germany
We ask whether specific attention heads, and more finely specific neurons inside those heads, are responsible for recognizing that a language model's context contains network infrastructure information (a hostname paired with its IP address), and whether that responsibility can be validated causally rather than by correlation alone. At the head level the answer is yes, across five models spanning three architecture families: in every model, a small set of heads (1 to 9 out of 128 to 1152 candidates), found by causal ablation screening and tested for selectivity against matched negative and context-free controls, supports a detector with 99.5--100% held-out accuracy. We then ask whether a head's responsibility concentrates into one neuron or stays spread across its dimensions; this is model-specific. In one model, the top head's signal concentrates into a single neuron, found independently by both a causal intervention and a correlational ranking, which agree exactly (AUC = 1.000, matching the full head). In another, the single clean head works as a whole (AUC = 1.000) but the best causally ranked neuron inside it does not (AUC = 0.665), so the responsibility there is spread across the head. The remaining three models fall in between. On an independent dataset collected by a different institution (reverse-DNS records rather than the discovery data), every model's full-head detector flags 100% of positive records; the single-neuron versions transfer less reliably, and in one model score below chance. Causal head-finding for a specific network-information entity works across models and architectures; how far that finding can be pushed down to individual neurons varies, and needs to be checked for each model.
Figures & tables
Figure 1: Pipeline overview. Two independent algorithms discover candidate heads and are cross-validated against each other; the surviving clean_subset heads then feed a second, structurally identical cross-validation at neuron granularity (correlational information-gain ranking vs. causal ablation), whose agreement or disagreement is this paper’s central result (Section 4.2 ).
Condition
Context
Question
True answer
P(internet)
Flagged
network
record
What is the IP address of this website?
62.151.181.154
1.000
yes
paraphrase
record
Resolve the IPv4 address associated with this domain.
62.151.181.154
1.000
yes
entity
record
What is the category of this website? Answer with one word: benign, mal, or phishing.
benign
0.000
no
vocabulary
record
In one short phrase, what does the abbreviation “IP” stand for?
Internet Protocol
0.000
no
ordinary
empty
What is 67 + 14?
81
0.001
no
question_only
empty
What is the IP address of this website?
62.151.181.154
0.024
no
Table 1: A worked example: one held-out record under every condition, with the predicted probability from the fixed Qwen3.5-4B detector. The network , paraphrase , entity and vocabulary rows share the same context ( al-mostafa.com , 62.151.181.154 ); the other rows have an empty context.
Figure 2: Where the interventions sit inside one attention layer. Every head’s output Oh (Eq. ( 2 )) is an additive term (Eq. ( 3 )); ablating head h zeroes its whole term, while ablating neuron j zeroes only one column of Oh . The forward-pre-hook on WO ’s input ( HeadSuppressor / NeuronSuppressor ) implements both.
Model
AUC
question_only flag rate
(both)
HE (20)
clean_subset
Qwen3.5-4B
1.000
1.00
0.00
Qwen3-8B
1.000
1.00
0.00
Qwen3.5-9B
1.000
1.00
0.00
Gemma-4-E4B
1.000
1.00
1.00
Gemma-4-E2B
1.000
1.00
1.00
Table 2: Unfiltered HE vs. filtered clean_subset : held-out AUC is identical either way, but the context-free wording confound ( question_only flag rate) differs: the filter lowers it to 0 only in the Qwen models.
Model
Cand.
Clean
Acc.
AUC
heads
heads
(held-out)
Qwen3.5-4B
128
9
100.0%
1.000
Qwen3-8B
1152
1
99.5%
1.000
Qwen3.5-9B
128
4
100.0%
1.000
Gemma-4-E4B
280
5
99.9%
1.000
Gemma-4-E2B
224
2
99.9%
1.000
Table 3: A causally-validated, selectivity-confirmed head set exists in every model tested, regardless of architecture or candidate-pool size.
Single neuron
Full head
Model
Acc.
AUC
Acc.
AUC
Qwen3.5-4B
99.4%
1.000
100.0%
1.000
Qwen3-8B
53.9%
0.665
99.5%
1.000
Qwen3.5-9B
99.4%
1.000
100.0%
1.000
Gemma-4-E4B
99.9%
1.000
99.9%
1.000
Gemma-4-E2B
86.6%
0.969
99.9%
1.000
Table 4: Single-neuron (best causally ranked neuron) vs. full-head detection performance (held-out accuracy at the 0.5 threshold, and AUC).
Model
Causal top / IG top
Agreement
Qwen3.5-4B
(15,1,0)/(15,1,0)
exact match
Qwen3-8B
(29,12,82)/(29,12,27)
same head, diff. dim
Qwen3.5-9B
(23,15,8)/(19,4,1)
different head
Gemma-4-E4B
(32,2,141)/(32,2,7)
same head, diff. dim
Gemma-4-E2B
(30,0,183)/(30,0,48)
same head, diff. dim
Table 5: Top-ranked neuron (layer, head, dimension) under the causal ranking and under the information-gain (IG) ranking.
Model
Actual
Random (same head)
Random (any neuron)
Qwen3.5-4B
1.000
0.727 ± 0.155
0.667 ± 0.186
Qwen3-8B
0.665
0.633 ± 0.148
0.728 ± 0.179
Qwen3.5-9B
1.000
0.680 ± 0.128
0.654 ± 0.195
Gemma-4-E4B
1.000
0.748 ± 0.180
0.744 ± 0.195
Gemma-4-E2B
0.969
0.675 ± 0.149
0.724 ± 0.188
Table 6: AUC of the actual single_best_causal_neuron classifier vs. two random-neuron baselines (mean ± std over 20 repeats each).
Model
Full head
Pos.
Neg.
Full head
Single neuron
acc.
acc.
acc.
AUC
AUC
Qwen3.5-4B
99.9%
100.0%
99.8%
1.000
0.535
Qwen3-8B
99.8%
100.0%
99.6%
1.000
0.422
Qwen3.5-9B
100.0%
100.0%
100.0%
1.000
0.285
Gemma-4-E4B
91.9%
100.0%
83.8%
0.995
0.866
Gemma-4-E2B
91.5%
100.0%
83.0%
0.971
0.463
Table 7: Cross-validation on an independent, real-world dataset (CAIDA). The full-head detector transfers with high AUC; the single-neuron detector does not. Pos. and Neg. are the accuracies on the CAIDA records and on the trivia negatives.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Condition
Qwen3.5-4B
Qwen3-8B
Qwen3.5-9B
Gemma-4-E4B
Gemma-4-E2B
network
1.000*
0.999*
0.999*
1.000*
0.975*
paraphrase
1.000*
1.000*
0.999*
1.000*
0.988*
entity
0.000
0.000
0.000
0.000
0.009
vocabulary
0.000
0.004
0.001
0.000
0.008
ordinary
0.001
0.002
0.001
0.000
0.002
question_only
0.024
0.006
0.111
0.996*
0.982*
Appendix
Table 8: Worked example, ask_ip direction: record al-mostafa.com / 62.151.181.154 , question “What is the IP address of this website?” (answer 62.151.181.154 for network / paraphrase ). Cells are P(internet) from each model’s own fixed detector; * marks flagged .
Condition
Qwen3.5-4B
Qwen3-8B
Qwen3.5-9B
Gemma-4-E4B
Gemma-4-E2B
network
1.000*
1.000*
1.000*
0.999*
1.000*
paraphrase
1.000*
1.000*
1.000*
0.999*
0.988*
entity
0.000
0.000
0.000
0.000
0.025
vocabulary
0.000
0.004
0.001
0.001
0.009
ordinary
0.001
0.004
0.001
0.000
0.004
question_only
0.035
0.011
0.117
0.998*
0.999*
Appendix
Table 9: Worked example, ask_url direction: record jjeducare.co.kr / 211.169.73.24 , question “What is the website address for this IP?” (answer jjeducare.co.kr for network / paraphrase ). Same layout as Table 8 .
Model
DNS-TTL flag rate
DNS-TTL mean P
MAC flag rate
MAC mean P
Qwen3.5-4B
0.990
0.884
1.000
0.994
Qwen3-8B
0.365
0.471
0.985
0.931
Qwen3.5-9B
0.440
0.576
1.000
0.993
Gemma-4-E4B
0.000
0.066
1.000
0.986
Gemma-4-E2B
0.020
0.129
0.275
0.355
Appendix
Table 10: Generalization to two new network-entity sub-types, never seen during discovery or training: 200 held-out records per model, each fixed detector unmodified (same classifier as Section 4 ). Flag rate is the fraction of the 200 items the detector labels network-positive; mean P is the mean predicted probability.
Model
Condition
N (clean)
N (random)
S (clean)
S (random)
Qwen3.5-4B
network
0.103
− 0.003
9.35
− 1.03
Qwen3.5-4B
dns_ttl
0.001
0.000
10.02
− 0.03
Qwen3.5-4B
mac
0.013
0.000
10.79
− 0.83
Qwen3-8B
network
0.000
0.000
1.22
0.06
Qwen3-8B
dns_ttl
0.000
0.000
0.42
0.23
Qwen3-8B
mac
0.000
0.000
1.13
0.12
Appendix
Table 11: Joint necessity ( N ) and sufficiency ( S ), in nats of log-probability, of clean_subset versus a random-head control of equal size, for the true IP/URL (network), TTL (dns_ttl), or MAC (mac) answer. S is the cleaner signal: sufficiency compares patching clean_subset ’s activations from the target condition into an unrelated entity prompt against patching a same-size random-head set the same way.
Genre
Should flag
Qwen3.5-4B
Qwen3-8B
Qwen3.5-9B
Gemma-4-E4B
Gemma-4-E2B
server log
yes
0.999
1.000
1.000
1.000
0.996
nginx config
yes
0.999
0.999
0.999
0.995
0.883
security incident report
yes
0.999
1.000
0.998
1.000
0.997
IT support email
yes
0.999
0.905
0.999
0.996
1.000
casual user chat
yes
1.000
0.993
1.000
0.999
1.000
server log (non-retrieval)
no
0.028
0.029
0.365
0.927 †
0.944 †
Appendix
Table 12: Use-case stress test: P(internet) from each model’s own fixed clean_subset detector on ten naturalistic items never seen during discovery or training, five positives, one per genre (real address + retrieval question), and five matched negatives. A dagger † (also in bold) marks a flag that disagrees with the intended label.