When Do Attention-Head Ablations Support Causal Claims? Projection-Level Confounds, Floor Effects, and Matched Controls
Organizations: Stanford University
Abstract
Attention-head ablation, zeroing a head and measuring the resulting change in task performance, is a common method for inferring which components of a language model are causally responsible for a behavior. We show using GPT-2 small that this inference can be fragile unless the intervention semantics, evaluation metric, and controls are carefully validated. A natural post-projection implementation of "zeroing a head" is nearly uncorrelated with a corrected pre-projection ablation (Pearson r = 0.057) and selects a completely disjoint top-5 set of important heads. We also show that binary accuracy can hide effects at behavioral floors and near ceilings, whereas gold-token log-probability remains graded. Using a discovery/held-out split and 1,000 matched random-head and layer-matched-head control draws, the corrected per-head effect ranking is highly stable across splits (Spearman rho = 0.974), and the top-5 selected heads significantly exceed both control distributions (Monte Carlo p = 0.001). However, evidence for task specificity is not robust on GPT-2. Replication on DistilGPT2 preserves the intervention-semantic and matched-control findings. These results show that single-head ablation does not by itself justify a causal claim; defensible interpretation requires correct intervention placement, a non-saturated continuous metric, and matched held-out controls.
Figures & tables
| Quantity | Value | Interval | / draws |
|---|---|---|---|
| Copying-task accuracy (ceiling) | 0.995 | [0.985, 1.000] | 200 |
| Arithmetic accuracy (floor) | 0.000 | [0.000, 0.000] | 200 |
| Copying-task log-prob (nats) | [ , ] | 200 | |
| Arithmetic log-prob (nats) | [ , ] | 200 | |
| Candidate , held-out | 0.412 | [0.353, 0.474] | 100 |
| Random-control , held-out | 0.029 | [ , 0.113] † | 1000 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Quantity | Value | Interval | / draws |
| Copying-task accuracy | 0.900 | [0.855, 0.940] | 200 |
| Arithmetic accuracy (floor) | 0.000 | [0.000, 0.000] | 200 |
| Copying-task log-prob (nats) | [ , ] | 200 | |
| Arithmetic log-prob (nats) | [ , ] | 200 | |
| RQ0: Pearson (naive, correct) | 0.001 | – | 72 heads |
| RQ0: top-5 head overlap | 0/5 | – | – |