Which Attention Heads are like the Human Head? Not the Ones that Compute
Authors: Christopher Pinier, Gustaw Opiełka, Hannes Rosenbusch, Taylor Webb, Michael D. Nunez, Claire E. Stevenson
Organizations: Psychological Methods, University of Amsterdam, Amsterdam, The Netherlands · Princeton Neuroscience Institute, Princeton University, Princeton, NJ, USA
Brain-AI alignment is often interpreted as a sign that model and brain perform similar computations. Whether the aligned units are causally involved in model computation is rarely checked. On an abstract pattern-completion task (AAABAAA → B), we compare LLM attention-head representations with human EEG and test how ablating those heads affects task performance. Alignment and causation dissociate: brain-aligned heads contribute to performance, but their removal is substantially less disruptive than removal of heads selected via attribution patching. We compare two head sets that prior interpretability work defines without reference to the brain: concept vectors (CVs), which represent abstract patterns across formats, and function vectors (FVs), selected for their contribution to correct-answer prediction. Brain alignment shows little association with FV scores, while its association with CV scores varies across models. Among brain-aligned heads, we find recurring attention profiles: one emphasizes distinctive elements (novelty heads), the other repeating elements (repetition heads). The novelty family tracks salience and attends to the same elements that humans look at, yet its removal is less damaging than random ablation on average. Repetition heads contribute modestly to performance and are associated with abstract-pattern representation (CVs). Across 17 models spanning 3B-72B parameters, FV-ranked removal is substantially more disruptive than brain-ranked removal. Brain alignment thus captures how the model reads the stimulus, and only faintly captures how it represents the pattern and solves the task.
Figures & tables
Figure 1: Schematic description of the findings. (a) What each attention head population does with one item of the task, the pattern AAABAAA whose answer is B (in words, for example, “hat hat hat risk hat hat hat” with answer “risk”), beside its brain alignment and its causal contribution. Colored symbols mark what each population attends to or encodes. (b) The same populations, schematically, in the plane of brain alignment against damage when the heads are removed; the gray band marks the damage from ablating random heads. Measured ablation curves are in Figure 3 .
Figure 2: The task prompt and the three head scores. (a) A query as presented to the model, preceded by three solved demonstrations in the same format (not shown here). The seven query symbols are based on one of eight abstract patterns (here, ABABCDC → D ) and the model’s next token after “Answer:” is scored. (b) Brain score: Spearman correlation between a head’s representational dissimilarity matrix (RDM), built from its output averaged within each pattern, and the human frontal-FRP RDM. (c) Patching score: a head’s output, averaged over clean prompts, is patched into a corrupted prompt from a different item. In the corrupted prompt, the elements of the first block (e.g., AAAB in AAABAAA ) of every in-context demonstration and of the query are re-drawn at random from the item’s symbols (red). This destroys the pattern while the remaining positions stay intact (green), and the demonstration answers are random. The gain in the probability of the clean answer (here “pen”, which completes AAABAAAB ), averaged over items from all eight patterns, is the head’s average indirect effect (AIE); the heads with the highest AIE are the function-vector (FV) heads. (d) Concept score: RSA of head outputs over 1,200 items that vary in pattern (8), alphabet (3; German, English, Symbols) and response format (2; open-ended, multiple choice), against a same-pattern design; a same-alphabet design serves as a control.
Figure 3: Cumulative ablation relative to random removal, across 17 models. Heads are zero-ablated cumulatively in each family’s ranking order; values are accuracy minus the mean of five size-matched random ablations, in percentage points (pp), so negative values mean greater impairment than random. (a) Mean over models, with head counts interpolated on a percentage-of-heads axis; bands are model-bootstrap 95% intervals. FV-ranked ablations are the most damaging ( −43 pp at 12.5% of heads), whereas removing novelty-ranked heads is less damaging than random removal ( 14 pp at 30%). (b) The same difference for each model, averaged over the first 30% of heads removed (shaded in a). Filled points are base models, open points instruction-tuned or distilled, and the diamond is Llama - 3.1 - 70B; bars are means over models, and counts give the models below random controls.
Figure 4: Two attention profiles discovered among brain-aligned heads, and their recurrence across models. (a) The 20 highest brain-scoring heads in Llama - 3.1 - 70B, clustered by their attention profiles. Three clusters emerge: the two larger ones are named by their profiles, the third resembles both. (b) Mean attention profiles of the repetition and novelty cluster over the seven query-symbol positions for each pattern. These means are the templates used to identify carriers across models. (c) Label check on all 5,120 heads of Llama - 3.1 - 70B: each head’s novelty-template loading against its attention share on the query’s distinctive symbol (the symbol shown once, averaged over the five patterns that have one). Dashed vertical lines mark the mean of the repetition carriers, of all heads, and of the novelty carriers. (d) For each of the 17 replication models, the number of its top-20 brain-scoring heads whose profile loads ≥.50 on the repetition template, on the novelty template, or on neither; the bottom row is the mean over models.
Figure 5: Human gaze and head-family attention. (a) Group fixation map of human participants over the seven visible query positions; each row is one pattern with its symbols overlaid, normalized to the share of that pattern’s total fixation duration. (b) Attention of the top 1% novelty heads to the same positions, normalized per pattern and averaged over the 17 models. (c) Cell-wise agreement between (a) and (b): both maps are centered per position, z-scored, and multiplied, so the cells average to the Pearson correlation between (a) and (b). Boxes in (a–c) mark the five patterns whose only unrepeated symbol is the correct continuation; 70% of the agreement fell in these five cells and 95% in their rows. The first-position case ( ABCDDCB → A ) drew neither gaze nor attention. (d) Spearman correlation between the human map and the mean attention profile of each model’s top 1% of heads per family, after centering both per position (pattern-relative). One point per model (diamond: Llama - 3.1 - 70B); bars are means over the 17 models, and counts give the models with ρ>0 . Novelty heads correlated positively with gaze in all 17 models and repetition heads negatively in all 17.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Example trial from the human task. Seven icons instantiate an abstract pattern; the participant selects the missing eighth icon from four alternatives. Here the pattern is ABCAABCA , and the star completes it. The FRP analysis retains fixations within the eight sequence-position regions, including the question-mark position, but not the response-option regions. Adapted from Pinier et al. (2025) .
Figure 7: Frontal electrodes used for the human FRP target. The 17 electrodes listed in the text are highlighted in red; other displayed montage electrodes are black. The scalp is viewed from above, with the nose at the top. Adapted from Pinier et al. (2025) .
Stage
Model trials
Used to construct
Head RSA
1,200 trials across six conditions
CV and control scores
Brain RSA
200 English MC trials (RSA subset)
human-RDM scores
Attention scan
200 English MC trials
fingerprints and templates
Attribution patching
200 English open-ended trials
FV/AIE ranking
Independent clustering
saved fingerprints
unlabeled attention groups
Set construction
saved scores
full rankings and intervention grid
Appendix
Table 1: Stage-specific trial conditions in the replication pipeline. MC denotes multiple choice. The six conditions cross English words, German words, and symbols with MC and open-ended completion, with 25 items per pattern and 200 trials per condition. Brain scores reuse the English MC subset of the head-RSA dataset; the attention scan and ablation sweep use a separately generated English MC set.
Model family and size
Base
Instruct / post-trained
Llama-3.2-3B
51.0
39.5
Llama-3.1-8B
62.5
59.5
Llama-3.1-70B
70.5
77.5
Llama-3.3-70B
—
82.5
Qwen2.5-3B
29.5
25.0
Qwen2.5-14B
52.5
75.0
Appendix
Table 2: Clean English multiple-choice performance. Exact next-token accuracy (%) on 200 trials per model under raw, three-shot prompting. Each populated cell is one of the 17 models. Dashes indicate no included model, not missing trials. Phi-4 and DeepSeek-R1-Distill-Llama are post-trained models; their placement in the right column does not imply identical training procedures.
Figure 8: Brain alignment and head-selection scores. Each point represents one model’s Spearman correlation across heads between frontal-FRP RSA and concept-RSA (CV) or attribution-patching (FV) scores; colors identify model families. Boxes show medians and interquartile ranges, with whiskers extending to the most extreme observations within 1.5 interquartile ranges. All model points are displayed. Undefined head scores are excluded pairwise. The 17 models are not independent samples of model families; these descriptive correlations are not participant-level inference.
Figure 9: Brain score against concept, patching and alphabet scores. Per cohort model, the Spearman correlation across heads between brain score and concept (CV), attribution-patching (FV) or alphabet (control) score. Circles: Llama, including the DeepSeek distill; squares: Qwen2.5; triangle: Phi-4. Open markers are tuned models, the outlined circle is Llama - 3.1 - 70B, and bars are family means. The brain–concept correlation was positive in every Llama and base Qwen model and negative in every tuned Qwen model and Phi-4; the other correlations stayed near zero.
Figure 10: Template carriers among the top 20 FRP heads. Each point is one of the 17 models; colors identify model families. Boxes show medians and interquartile ranges, with whiskers extending to the most extreme observations within 1.5 interquartile ranges. (A) Percentage of the top 20 heads matching repetition, novelty, or neither template at r≥.50 . (B) Within-model difference between top-20 carrier share and its prevalence among all model heads; zero indicates equal prevalence. Positive values indicate overrepresentation, not a significance test.
Figure 11: Templates and independently discovered cluster profiles. Top row: the original repetition and novelty templates. Bottom row: equal-model mean profiles of the three largest clusters among each model’s top 20 FRP heads. Rows are patterns; columns are query positions, with the sequence symbols overlaid. Color intensity shows within-query attention share on a common scale; pink and blue link the first two aligned profiles to their display references, not to a categorical template assignment. For display, profiles were assigned one-to-one to the original cluster profiles by maximizing total Pearson similarity, without changing memberships. Title correlations compare each displayed mean profile with the indicated template using Pearson correlation after centering each position across patterns; they are not averages of model-wise correlations. This alignment is not independent evidence of recurrence or of three clusters per model. Post-hoc template comparisons found novelty-like matches in 17/17 models and repetition-like matches in 13/17 at r≥.50 .
Figure 12: Repetition-head removal reduced accuracy but preserved measured pattern information (Llama - 3.1 - 70B). (a) Repetition-template loading against concept score for all heads (Spearman ρ ). Pink: the 130 ablated repetition heads; purple: the 150 highest-scoring concept (CV) heads; gray: all other heads. (b) Layer distribution of the same two sets. The shaded band marks the readout heads, the concept heads in layers 34 and above. (c) Change from the unablated model after zero-ablating the 130 repetition heads (rep.), five damage-matched random 25-head sets (rand.; bar: mean), or 130 layer-matched novelty carriers (nov.). Left: multiple-choice task accuracy. Right: cross-format concept RSA over the readout heads. A pattern classifier on the readout heads stayed at 100% accuracy in every condition.
Figure 13: Attribution price of every novelty carrier on every task (z against layer-matched non-carriers) beside the gate column.
Figure 15: Template loading, OV copying and causal price of the future novelty population across Pythia-1.4B training.
Figure 16: Exploratory Pythia pattern-task ablation, KL and template-population trajectories. These concern the tested interventions, not universal functionality.
Attention-head ablation, zeroing a head and measuring the resulting change in task performance, is a common method for inferring which components of a language model are causally responsible for a behavior. We show using GPT-2 small that this inference can be fragile unless the intervention semantics, evaluation metric, and controls are carefully validated. A natural post-projection implementation of "zeroing a head" is nearly uncorrelated with a corrected pre-projection ablation (Pearson r = 0.057) and selects a completely disjoint top-5 set of important heads. We also show that binary accuracy can hide effects at behavioral floors and near ceilings, whereas gold-token log-probability remains graded. Using a discovery/held-out split and 1,000 matched random-head and layer-matched-head control draws, the corrected per-head effect ranking is highly stable across splits (Spearman rho = 0.974), and the top-5 selected heads significantly exceed both control distributions (Monte Carlo p = 0.001). However, evidence for task specificity is not robust on GPT-2. Replication on DistilGPT2 preserves the intervention-semantic and matched-control findings. These results show that single-head ablation does not by itself justify a causal claim; defensible interpretation requires correct intervention placement, a non-saturated continuous metric, and matched held-out controls.
Artificial vision models are often evaluated against the human visual cortex by measuring how accurately their internal representations predict brain responses. However, prediction accuracy alone does not indicate which dimensions of the target brain's response space are recovered. Here, we introduce a unified framework for evaluating both model-brain and brain-brain alignment by identifying the response dimensions recovered by prediction. Using repeated fMRI measurements, we first identify target-brain response dimensions that can be reproducibly predicted across independent trial splits. We then predict target-brain responses from either another subject's brain responses or a vision model's internal representations, and quantify how strongly each of these reproducible response dimensions is recovered. Applying this framework to a subset of the Natural Scenes Dataset, in which eight subjects viewed the same natural images during fMRI, we find that the early-to-intermediate visual-cortex responses contain a low-dimensional set of reproducible dimensions. Brain-to-brain comparisons identify which of these dimensions are consistently recoverable from other subjects' brains, providing a diagnostic human reference rather than only a scalar benchmark. In some cases, pretrained and randomly initialized models achieve similar prediction accuracy while showing distinct recovery profiles across these response dimensions. These results show that prediction accuracy alone can mask model-brain mismatches. By making explicit which reproducible brain response dimensions are recovered by prediction, our framework provides a more diagnostic evaluation of alignment between artificial vision models and the human visual cortex.
Ken Nakamura, Tomoya Nakai, Ryuto Yashiro +2
1The University of Tokyo · 2Osnabrück University and Freie Universität Berlin · 3Kobe University
Interpretability increasingly treats groups of components, not individual units, as the basic object, and proposes to find them by clustering co-activation statistics. We ask whether such a cheap signal actually identifies an attention-head circuit. Adapting a sparse-autoencoder clustering recipe to attention heads -- but validating by causal ablation rather than reconstruction -- we cluster heads and then run a closure test: ablate the discovered community and compare per-example damage to matched-random controls. Across two dense 1B-scale models (Pythia 1B, OLMo 1B) and two input distributions, the communities pass closure. In a Mixture-of-Experts model (OLMoE-1B-7B), route-conditional clustering recovers a statistically real signal that nonetheless does not survive closure -- ablation improves loss, the wrong direction. Extending closure across training, attention-target selectivity and participation ratio decouple from function in both directions. We conclude that a cheap signal is a circuit proposal, not a confirmed circuit; closure is what separates them.