Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair
Organizations: Lexsi Labs
Abstract
Ablate a component of a language model, and other components often appear to adjust and compensate. This phenomenon, termed self-repair, has been observed repeatedly, but its mechanism remains unclear. The most systematic study to date concluded that self-repair is noisy and unlikely to have a single explanation. We argue that it has one: a gain already present before any ablation. Any intervention on a causally important component can be viewed as a point on a coordinate axis , the signed strength of a counterfactual contrast. Hence, conventional ablation methods are uncalibrated points on this axis. We show that the causal repair response for a fine-grained unit is governed by an affine law, . The slope is a fixed coefficient that consistently influences the model, with or without ablation, and its sign determines whether the unit counteracts or reinforces the removed signal. On a factual-verdict task across four models from distinct families (Gemma, Qwen, LLaMA, and Mistral), we identify components including MLP neurons, OV neurons, and singular directions that follow this affine law, 68 of 81 downstream directions in all. Moreover, we can anticipate the magnitude of from the fixed weights. On the IOI circuit of GPT-2 Small, seven of the ten heads the intervention can reach follow the law, and all seven are counterweights. From this perspective, what may appear as self-repair is a counterweight performing its usual operation when the contrastive signal emerges at the core.
Figures & tables
| Model | Core(s) | core share | ||
|---|---|---|---|---|
| Llama-3-8B-Instruct | L18.n11065 | |||
| Gemma-2-9B-it | L28.n2046 L33.n4294 | |||
| Qwen2.5-7B-Instruct | L22.n13149 L24.n14758 | |||
| Mistral-7B-Instruct-v0.3 | L20.n14286 L19.n8228 |
Appendix figures & tables61 assets
Supplementary material from the paper’s appendix.
Appendix
| Main-text claim | Supporting appendix section |
|---|---|
| Contribution 1 (compensation is standing) | § F (recruitment), § G (other primaries), § H (base checkpoints), § I (not a LayerNorm artefact) |
| Contribution 2 (a dose axis and a nested intervention) | § E , § N |
| Contribution 3 (the coupling has an affine form) | § J and § K , with IOI in § L |
| Contribution 4 (named counterweights and relays, coupling legible in the weights) | § D (catalogues, clean census), § M |
| Symbol | Meaning |
|---|---|
| The prompt (SCM exogenous variable). | |
| The matched counterfactual partner of (opposite class, identical phrasing). | |
| World index: (clean), (core dosed to ; the main text writes for the same world), (learned perturbation). | |
| A weight-space direction (an mlp_neuron , ov_neuron , ov_svd , or mlp_svd unit; § B ). | |
| ’s fixed write vector. | |
| ’s activation coefficient on prompt in world . |
| Gated MLP, neuron of layer : | |||
|---|---|---|---|
| Model | MLP input | Neuron scalar | Residual write |
| Llama-3-8B-Instruct | |||
| Gemma-2-9b-it | |||
| Qwen2.5-7B-Instruct | |||
| Mistral-7B-Instruct-v0.3 | |||
| Experiment | Section | |
| Llama single-core swap ( L18.n11065 ) | 532 | § F |
| Gemma joint-core swap ( L28.n2046+L33.n4294 ) | 583 | § F |
| Qwen joint-core swap ( L22.n13149+L24.n14758 ) | 478 | § F |
| Mistral joint-core swap ( L20.n14286+L19.n8228 ) | 598 | § F |
| Gemma / Qwen dose-resolved single-core comparisons | 100 | § F |
| Learned perturbation ( perfect_on_pairs , all models) | – (full pool) | § G |
| Axis | Promoted (top-4) | Suppressed (bottom-4) |
|---|---|---|
| L23.n8972 | iller, rello, LC, etz | lie , lies , lying , Lie |
| L19.n2738 | ledged, Alo, ADDE , くだ (Jp.) | illusion , illusions , false , illusion |
| L23.n9811 | truth , Truth , truth , Truth | oby, imore, Tell, tell |
| L19.n3754 | Edition, onte , フ (Jp.) , rale | sure , yes , yes , Sure |
| L22.h1.svd117 | miss, missing, chances, misses | true , true , Important , TRUE |
| L24.n4543 | yes , true , True , Yes | OrNil, uelle, illard, izh (Cyr.) |
| Axis | Promoted (top-4) | Suppressed (bottom-4) |
|---|---|---|
| L38.n7088 | wrong , wrong , Wrong , Wrong | __). , BarStyle , preferably , AutoSize |
| L39.n10085 | fake , Fake , pseudo , pseud | argout , cyd , MessageOf , MigrationBuilder |
| L38.n854 | ModelExpression , __": , ituary , orghini | true , true , True , True |
| L41.n8771 | SOUNDBITE , estekak , protoimpl , ChrTalk | True , true , True , true |
| L39.n10210 | incorrect , incorrect , Incorrect , WRONG | IBOutlet , setopt , openzeppelin , Geos |
| L41.n3789 | haikusbot , writeFieldEnd , complexType , getDescription | real , skuto , truly , genuine |
| Axis | Promoted (top-4) | Suppressed (bottom-4) |
|---|---|---|
| L27.n4228 | , -Headers , ómo , ’}}> | truly , 真正 (genuine) , true , 真正的 (genuine) |
| L22.n13149 | true , True , _true , true | false , False , false , False |
| L20.n2073 | False , false , False , false | true , True , true , True |
| L22.n4538 | false , false , False , False | true , True , True , true |
| L27.n13033 | fa , ase , л (Cyr.) , fa | false , False , False , false |
| L24.n3278 | true , true , 真实 (real) , True | ulfill , .Gray , naz , 泄露 (Zh.) |
| Axis | Promoted (top-4) | Suppressed (bottom-4) |
|---|---|---|
| L15.n11835 | fake , invalid , false , False | riel , lib , trends , HAL |
| L16.n1170 | ensures , ensuring , confirmed , ingo | STO , FALSE , false , |
| L16.n5090 | ugno , iten , Graf , rolog | pret , pretend , fake , nomin |
| L16.n9991 | false , fals , False , false | , amon , bert , bis |
| L17.n11073 | integrity , respons , genuine , proper | artificial , stere , unsafe , false |
| L17.n13852 | true , real , True , TRUE | asp , ape , dom , Mend |
| (clean) | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| secondary | family | detects | flip@ | ||||||||||
| L23.n8972 | mlp neuron | lie | |||||||||||
| L23.n9811 | mlp neuron | truth | |||||||||||
| L19.n3754 | mlp neuron | sure | |||||||||||
| L21.n4049 | mlp neuron | wrong | |||||||||||
| L19.n2738 | mlp neuron | illusion | |||||||||||
| (clean) | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| secondary | family | detects | flip@ | ||||||||||
| L27.n13033 | mlp neuron | false | |||||||||||
| L27.n4228 | mlp neuron | truly | |||||||||||
| L25.n4929 | mlp neuron | false | |||||||||||
| L23.n8341 | mlp neuron | illusion | |||||||||||
| L24.n12848 | mlp neuron | false | |||||||||||
| (clean) | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| secondary | family | detects | flip@ | ||||||||||
| L38.n7088 | mlp neuron | wrong | |||||||||||
| L39.n10085 | mlp neuron | fake | |||||||||||
| L38.n854 | mlp neuron | true | |||||||||||
| L41.n8771 | mlp neuron | True | |||||||||||
| L39.n10210 | mlp neuron | incorrect | |||||||||||
| (clean) | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| secondary | family | detects | flip@ | ||||||||||
| L31.n13669 | mlp neuron | true | |||||||||||
| L31.n8773 | mlp neuron | false | |||||||||||
| L28.n4079 | mlp neuron | true | |||||||||||
| L22.h23.svd2 | ov svd | wrong | |||||||||||
| L22.n8735 | mlp neuron | false | |||||||||||
| Label Assignment Rule ( ) | Class Query ( ) | Class Query ( ) |
|---|---|---|
| (Class maps to ) | Prompt: | Prompt: |
| Expected Output: | Expected Output: | |
| (Class maps to ) | Prompt: | Prompt: |
| Expected Output: | Expected Output: |
| true | fake | strict | interp. backups | |
|---|---|---|---|---|
| model | clean sup | clean fab | flip? | recruited |
| Llama-3-8B | 10 | |||
| Gemma-2-9B | 5 | |||
| Qwen2.5-7B | ✓ | 6 | ||
| Mistral-7B-v0.3 | 3 |
| (clean) | (perturb) | |||||||
| axis | family | detects | ||||||
| Llama-3-8B (clean : T /F prim T /F ) | ||||||||
| L22.h1.svd117 | ov svd | true | ||||||
| L19.n3754 | mlp neuron | sure | ||||||
| L22.h1.c60 | ov neuron | valid | ||||||
| L21.n4049 | mlp neuron | wrong | ||||||
| clean | Swap | Perturb | Swap Perturb | ||||||
| axis | detects | T | F | T | F | T | F | all | |
| Llama-3-8B | |||||||||
| L19.n3754 | sure | anti | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | |
| L21.n4049 | wrong | anti | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | |
| L22.h1.c60 | valid | null | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | |
| L22.h1.svd116 | correct | null | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | |
| Model | Class | profile in | |||||
|---|---|---|---|---|---|---|---|
| Llama-3-8B-Instruct | T | 0.990 | 1.023 | 1.060 | 1.105 | 1.157 | increasing |
| F | 0.930 | 0.992 | 1.000 | 1.010 | 1.032 | flat/slight | |
| Gemma-2-9b-it | T | 28.99 | 0.997 | 0.957 | 0.934 | 0.926 | decreasing |
| F | 31.07 | 1.080 | 1.103 | 1.089 | 1.086 | flat | |
| Qwen2.5-7B-Instruct | T | 6.09 | 0.996 | 0.978 | 0.941 | 0.890 | decreasing |
| F | 5.92 | 1.024 | 1.033 | 1.017 | 0.985 | flat |
| Model | Cls | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Llama-3-8B | T | |||||||||
| F | ||||||||||
| Gemma-2-9b | T | |||||||||
| F | ||||||||||
| Qwen2.5-7B | T | |||||||||
| Unit (detects) | Cls | dose | share of | |||
|---|---|---|---|---|---|---|
| L23.n8972 ( lie ) | T | wrong sign | ||||
| Llama-3-8B | T | wrong sign | ||||
| F | ||||||
| F | wrong sign | |||||
| L38.n7088 ( wrong ) | T | |||||
| Gemma-2-9b | T |
| Intervention | Clamp | Dose | In expectation | Used by |
|---|---|---|---|---|
| Counterfactual patch | (partner) | Vig et al. [20] , Geiger et al. [45] | ||
| Neutral clamp | this paper | |||
| Dosed swap | chosen | this paper | ||
| Zero ablation | unit-set; iff | Michel et al. [46] ; Gong et al. [32] | ||
| Mean ablation | ; if balanced | Wang et al. [1] | ||
| Resample ablation | , | , sd | Chan et al. [21] , McGrath et al. [5] , Rushing and Nanda [6] |
| coupling sign | median | flip err. (doses) | |||||||
|---|---|---|---|---|---|---|---|---|---|
| model | ds. | cert. | mlp | ov/svd | inh. | in-window | beyond | ||
| Llama-3-8B-Instruct | 20 | 18 | 18 | 0 | 0.84 | 0.74 | 0.67 (6) | 0.02 (13) | — |
| Gemma-2-9b-it | 22 | 17 | 14 | 3 | 0.78 | 0.72 | 0.38 (15) | 0.05 (2) | 0.12 (12) |
| Qwen2.5-7B-Instruct | 16 | 14 | 9 | 5 | 0.79 | 0.24 | 0.40 (13) | 0.02 (3) | 0.22 (9) |
| Mistral-7B-Instruct-v0.3 | 23 | 19 | 11 | 8 | 0.58 | 0.45 | 0.23 (13) | 0.03 (3) | 0.22 (11) |
| pooled | 81 | 68 | 52 | 16 | 0.76 | 0.68 | 0.39 (47) | 0.02 (21) | 0.19 (32) |
| direction | inh. | lin. | ||||||
| Llama-3-8B-Instruct | ||||||||
| L23.n8972 | -0.015 | -0.032 | 0.95 | 0.68 | 0.74 | 0.73 | +4.7 | |
| L23.n9811 | -0.009 | -0.012 | 0.95 | 0.59 | 0.85 | 0.85 | +5.1 | |
| L21.n4049 | -0.000 | -0.011 | 0.99 | 0.97 | 0.52 | 0.51 | +0.5 | |
| L19.n3754 | -0.005 | -0.009 | 0.76 | 0.66 | 0.76 | 0.76 | +7.0 | |
| L22.h1.svd117 | +0.003 | -0.007 | 0.95 | – | 0.31 | 0.29 | -0.8 | |
| direction | inh. | lin. | ||||||
| Mistral-7B-Instruct-v0.3 | ||||||||
| L31.n8773 | -0.393 | -0.398 | 0.98 | 0.50 | 0.99 | 0.96 | +18.0 | |
| L31.n13669 | -1.850 | -0.208 | 0.58 | 0.10 | 4.96 | 1.88 † | +14.3 | |
| L28.n4079 | -0.239 | -0.158 | 0.55 | 0.40 | 1.25 | 1.22 † | +13.4 | |
| L24.n4907 | +0.095 | -0.093 | 0.90 | – | – | – | +13.7 | |
| L22.h20.svd0 | +0.302 | -0.083 | 0.44 | – | – | – | +9.1 | |
| intercepts | shared slope | half rungs | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Model | Direction | (SE) | max | % span | track | |||||
| Llama-3-8B | L23.n8972 | |||||||||
| Gemma-2-9B | L38.n7088 | |||||||||
| Qwen2.5-7B | L27.n13033 | |||||||||
| Mistral-7B-v0.3 | L31.n8773 | |||||||||
| : 5 rungs vs 3 | shift | quadratic | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | sign | med. | med. | misfit rem. | OOS | ||||
| Llama-3-8B | 20 | ||||||||
| Gemma-2-9B | 22 | ||||||||
| Qwen2.5-7B | 16 | ||||||||
| Mistral-7B-v0.3 | 23 | ||||||||
| Model | Downstream directions | Correction family | Prompt pairs |
|---|---|---|---|
| Llama-3-8B-Instruct | 20 | 20 | 532 |
| Gemma-2-9B-it | 22 | 22 | 583 |
| Qwen2.5-7B-Instruct | 16 | 22 | 478 |
| Mistral-7B-Instruct-v0.3 | 23 | 33 | 598 |
| Total | 81 | 97 |
| Status | Count | (pooled fit) |
|---|---|---|
| Slope indistinguishable from zero | 7 | – |
| Excluded at the margin of the correction | 6 | – |
| in-sample median | held-out | |||
|---|---|---|---|---|
| Model | split (reported) | single | split (reported) | single |
| Llama | 0.874 | 0.755 | 0.905 | 0.858 |
| Gemma | 0.911 | 0.599 | 0.803 | 0.586 |
| Qwen | 0.821 | 0.665 | 0.904 | 0.867 |
| Mistral | 0.903 | 0.499 | 0.653 | |
| intercepts | shared slope | half rungs | |||||
|---|---|---|---|---|---|---|---|
| Model | Direction | (SE) | % span | ||||
| Llama-3-8B | L23.n8972 | ||||||
| Gemma-2-9B | L38.n7088 | ||||||
| Qwen2.5-7B | L27.n13033 | ||||||
| Mistral-7B-v0.3 | L31.n8773 | ||||||
| Withheld | points | affine line | bracket interpolation | quadratic |
|---|---|---|---|---|
| half rungs | 324 | 7501 | 1929 | 1188 |
| neutral rung | 162 | 3221 | 530 | 390 |
| Withheld rung | median | single-intercept | median % of span | frac. | |
|---|---|---|---|---|---|
| 2.97 | 4.14 | 17.9 | 0.62 | endpoint | |
| 1.39 | 2.36 | 9.1 | 0.36 | interior | |
| (neutral) | 1.36 | 2.19 | 9.0 | 0.36 | interior |
| 1.64 | 1.94 | 8.4 | 0.41 | interior | |
| 3.04 | 2.73 | 16.5 | 0.59 | endpoint |
| Head | Role | (SE) | cert. | |
|---|---|---|---|---|
| L10.h7 | negative name mover | yes | ||
| L11.h2 | backup name mover | yes | ||
| L10.h2 | backup name mover | yes | ||
| L11.h6 | backup name mover | yes | ||
| L10.h0 | name mover | yes | ||
| L11.h10 | negative name mover | – |
| Model | one-sided perm. | ||
|---|---|---|---|
| Llama-3-8B | 20 | ||
| Gemma-2-9B | 22 | (floor) | |
| Qwen2.5-7B | 16 | ||
| Mistral-7B-v0.3 | 23 |
| count | ||||||||
| Family | bench | down | cert. | cw / relay | (sign) | |||
| mlp_neuron | MLP neuron | 67 | 54 | 49 | 35 / 14 | (49/49) | ||
| ov_neuron | OV channel | 16 | 16 | 10 | 10 / 0 | (10/10) | ||
| ov_svd | OV SVD | 11 | 10 | 8 | 6 / 2 | (8/8) | ||
| mlp_svd | MLP SVD | 3 | 1 | 1 | 1 / 0 | (1/1) | ||
| All | 97 | 81 | 68 | sign preserved on all 68 certified directions | ||||
| direction | family | layer | clean | rel. | ||
| Qwen2.5-7B-Instruct cores at L22, L24 | ||||||
| L22.n4538 | mlp neuron | 22 | anti | |||
| L22.n9609 | mlp neuron | 22 | contribute | |||
| L20.h1.svd6 | ov svd | 20 | contribute | |||
| L22.n13025 | mlp neuron | 22 | anti | |||
| L11.svd690 | mlp svd | 11 | null | |||
| three-point fit | full ladder | ||||||
|---|---|---|---|---|---|---|---|
| Model | Direction | fold | fold | ||||
| Qwen2.5-7B | L22.n4538 | ||||||
| Qwen2.5-7B | L22.n9609 | ||||||
| Qwen2.5-7B | L20.h1.svd6 | ||||||
| Qwen2.5-7B | L22.n13025 | ||||||
| Qwen2.5-7B | L11.svd690 | ||||||