cs.LGOct 6, 2026

Finding the Heads and the Neurons Responsible for Network Information Retrieval in Language Models

Authors: Abdul Kadir, Md Mohasin Hossain, Daniel Sonntag

Organizations: University of Oldenburg, Oldenburg, Germany · German Research Center for Artificial Intelligence · Saarland University, Saarbrucken, Germany

Abstract

We ask whether specific attention heads, and more finely specific neurons inside those heads, are responsible for recognizing that a language model's context contains network infrastructure information (a hostname paired with its IP address), and whether that responsibility can be validated causally rather than by correlation alone. At the head level the answer is yes, across five models spanning three architecture families: in every model, a small set of heads (1 to 9 out of 128 to 1152 candidates), found by causal ablation screening and tested for selectivity against matched negative and context-free controls, supports a detector with 99.5--100% held-out accuracy. We then ask whether a head's responsibility concentrates into one neuron or stays spread across its dimensions; this is model-specific. In one model, the top head's signal concentrates into a single neuron, found independently by both a causal intervention and a correlational ranking, which agree exactly (AUC = 1.000, matching the full head). In another, the single clean head works as a whole (AUC = 1.000) but the best causally ranked neuron inside it does not (AUC = 0.665), so the responsibility there is spread across the head. The remaining three models fall in between. On an independent dataset collected by a different institution (reverse-DNS records rather than the discovery data), every model's full-head detector flags 100% of positive records; the single-neuron versions transfer less reliably, and in one model score below chance. Causal head-finding for a specific network-information entity works across models and architectures; how far that finding can be pushed down to individual neurons varies, and needs to be checked for each model.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. A mechanistic study of language model introspection

    Sep 28, 2026Jiahong Zou, Xiangkun Sun, Lingkai Kong +1Introspection AdaptersLarge Language Models Fail

  2. Rethinking Contextualization by Reinterpreting Attention Head Channels

    Sep 27, 2026Hakaze Cho, Haolin Yang, Zhun Sun +3Efficient Long-Context InferenceChannel