Adversarial perturbations can reduce state-of-the-art toxicity classifiers to near-zero accuracy, yet existing defences treat models as black boxes. We apply mechanistic interpretability to toxicity classification for the first time, identifying the internal attention-head circuits responsible for both correct classification and adversarial vulnerability. Across a 2×2 factorial study (BERT × RoBERTa) × (Jigsaw × ToxiGen), extended to Llama Guard~2 (8B), we show that zeroing a single attention head recovers up to 70.4 pp of adversarial accuracy for RoBERTa on Jigsaw and 37.3 pp for Llama Guard 2 on ToxiGen, at ≤0.6 pp clean cost. Vulnerable heads generalise to held-out examples within ≤1 pp, and a class-imbalance sweep confirms they act as selective toxic-class detectors. Head suppression matches or outperforms adversarial training on Jigsaw; data augmentation dominates on ToxiGen: a dataset-specific reversal explained by whether the classifier encodes a concentrated bottleneck or a distributed circuit. Demographic analysis across 20 Jigsaw and 13 ToxiGen minority groups reveals structurally unequal adversarial vulnerability, exposing mechanistically traceable fairness gaps in current toxicity classifiers.
Figures & tables
Input
Text
Label
Prediction
Regular
“Your observations are consistent with those of most imbeciles”
Toxic
Toxic ✓
Adversarial
“your terms are consistent with those of most imbeciles”
Toxic
Non-toxic ✗
Regular
“Homosexuality is a sin and a choice, not a genetic factor. Homosexuality is a behavioral disorder conflicting with one’s God-given sex”
Toxic
Toxic ✓
Adversarial
“homosexuality is a sin and a choice, not a genetic factor. homosexual is a behavioral disorder conflicting with one ’ s god - given sex”
Toxic
Non-toxic ✗
Regular
“the immigration system in the USA is full of absurdities”
Non-toxic
Non-toxic ✓
Adversarial
“the immigration system in the us is full of absurdities”
Non-toxic
Toxic ✗
Table 1: Adversarial input example showing semantic equivalence but different model predictions.
Figure 1: End-to-end pipeline from adversarial attack to targeted head suppression across six model–corpus combinations. Fine-tuned classifiers ( Stage 1 ) are attacked via PGD BERT-Attack ( Stage 2 ); activation patching identifies the bottleneck head set H⋆ ( Stage 3 ); zero-ablating H⋆ improves adversarial accuracy while preserving clean performance ( Stage 4 ). The dashed arc denotes evaluation of the suppressed model f~θ on both clean and adversarial inputs.
Figure 2: Clean-input activation patching: per-head Δ Loss (increase = head is crucial) across all four runs in a 2 × 2 grid (rows: BERT/RoBERTa; columns: Jigsaw/ToxiGen). Warmer colours indicate higher impact. Jigsaw models show a single dominant hot cell (L9H5, L6H1); ToxiGen models are diffuse.
Run
Architecture
Clean Acc (balanced)
bert_jigsaw
BERT
78.35%
roberta_jigsaw
RoBERTa
79.88%
bert_toxigen
BERT
82.40%
roberta_toxigen
RoBERTa
85.98%
Table 2: Best checkpoint accuracy on the 10k balanced test set after 3 training epochs.
Run
Top head
Δ Loss
Acc drop
bert_jigsaw
L9H5
+0.132
− 5.03 pp
roberta_jigsaw
L6H1
+0.178
− 11.58 pp
bert_toxigen
L11H2
+0.020
− 0.55 pp
roberta_toxigen
L4H1
+0.011
− 0.54 pp
Table 3: Top crucial head per run identified by clean activation patching.
Run
Top head
Δ Loss
k=1
k=5
bert_jigsaw
L9H5
− 0.562
+50.5 pp
+84.9 pp
roberta_jigsaw
L6H1
− 0.721
+70.4 pp
+76.7 pp
bert_toxigen
L9H5
− 0.069
+23.6 pp
+34.1 pp
roberta_toxigen
L6H1
− 0.174
+23.5 pp
+38.9 pp
Table 4: Top vulnerable head and adversarial accuracy recovery ( Δ from near-zero baseline) at k =1 and k =5 heads suppressed.
Figure 3: Adversarial accuracy recovery as k heads are suppressed. Jigsaw curves (left) rise steeply at k =1; ToxiGen curves (right) are shallower across all k .
Run
Heads
Base
Supp.
Δ
bert_jigsaw
L9H5, L8H8
0.17%
67.22%
+67.1 pp
roberta_jigsaw
L6H1, L5H2
0.09%
56.79%
+56.7 pp
bert_toxigen
L9H5, L11H3
0.16%
25.66%
+25.5 pp
roberta_toxigen
L6H1, L4H10
0.33%
27.46%
+27.1 pp
Table 5: Overall adversarial accuracy before and after top-2 head suppression on the demographic adversarial pool.
Figure 4: Per-group adversarial accuracy after top-2 head suppression. Groups sorted by suppressed accuracy; ∗ = n<30 .
Figure 5: Adversarial (solid) and clean (hatched) accuracy for all robustness methods. Head suppression wins on Jigsaw; augmentation wins on ToxiGen.
Run
Method
Clean
Adv
bert_jig.
Baseline
94.88%
0.00%
Head supp.
94.71%
66.00%
FGM
94.96%
42.74%
Augmentation
88.74%
65.81%
rob._jig.
Baseline
94.45%
0.00%
Head supp.
93.89%
73.14%
Table 6: Clean and adversarial accuracy for all robustness methods. Bold = best adv. accuracy per run. Clean accuracy on unbalanced natural test.
Run
Full (k=1)
Held-out B
Diff
bert_jigsaw
50.54%
50.90%
+0.36 pp
roberta_jigsaw
70.64%
71.61%
+0.97 pp
bert_toxigen
23.79%
23.10%
− 0.69 pp
roberta_toxigen
23.62%
24.02%
+0.40 pp
Table 7: Split-sample validation: adversarial accuracy at k =1 on full set vs. held-out Set B. All differences ≤ 1 pp.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: PGD BERT-Attack pipeline: gradients are computed w.r.t. input token embeddings and used to guide substitution via a BERT MLM, iterating within a norm-bounded perturbation ball.
Figure 7: Adversarial-input activation patching: per-head Δ Loss on successfully-attacked examples. Red cells = heads whose ablation helps recover adversarial accuracy (negative Δ Loss). L9H5 (BERT) and L6H1 (RoBERTa) are the dominant bottleneck heads for Jigsaw; ToxiGen shows a more distributed pattern.
Figure 8: FGM adversarial training curves (10 epochs). Jigsaw models collapse catastrophically after the best epoch; ToxiGen models are stable throughout.
BERT Jigsaw
Ep.
Loss
Clean
Adv
1
0.122
94.98%
30.82%
2
0.118
95.10%
35.98%
3
0.104
94.96%
42.74%
4
0.092
94.63%
10.93%
5
0.082
94.01%
5.96%
Appendix
Table 8: FGM training curves: BERT Jigsaw (best ep. 3) and RoBERTa Jigsaw (best ep. 1). Bold = best epoch.
BERT ToxiGen
Ep.
Loss
Clean
Adv
1
0.354
82.57%
31.84%
2
0.328
82.93%
33.98%
3
0.276
83.10%
32.42%
4
0.228
83.16%
34.96%
5
0.184
83.04%
33.59%
Appendix
Table 9: FGM training curves: BERT ToxiGen (best ep. 4) and RoBERTa ToxiGen (best ep. 9). Bold = best epoch.
Run
k
Heads (Set A)
Held-out Adv Acc
bert_jigsaw
1
L9H5
50.90%
2
+L8H8
65.74%
3
+L11H6
76.60%
5
+L6H7, L10H11
84.68%
roberta_jigsaw
1
L6H1
71.61%
2
+L5H2
75.08%
Appendix
Table 10: Set B held-out cumulative adversarial accuracy (exp06).
Figure 9: Class-imbalance sweep: clean cost (left) and adversarial benefit (right) of top-2 suppression at five toxic ratios. Jigsaw clean cost rises monotonically (22 × spread for RoBERTa), confirming selective toxic-class detection.
Run
10%
30%
50%
70%
90%
bert_jigsaw
+0.88
+5.15
+9.62
+14.28
+18.63
roberta_jigsaw
+1.67
+10.87
+19.53
+27.93
+36.88
bert_toxigen
− 0.63
− 0.02
+0.90
+1.45
+1.82
roberta_toxigen
+2.35
+1.33
+0.45
− 0.22
− 1.03
Appendix
Table 11: Clean accuracy cost ( Δclean , pp) of top-2 head suppression at each toxic ratio.
Run
10%
30%
50%
70%
90%
bert_jigsaw
+66.51
+51.31
+36.38
+22.31
+7.28
roberta_jigsaw
+87.06
+67.59
+48.56
+29.58
+9.90
bert_toxigen
+31.72
+27.28
+23.92
+19.57
+15.52
roberta_toxigen
+27.30
+29.67
+32.67
+35.47
+38.35
Appendix
Table 12: Adversarial accuracy benefit ( Δadv , pp) of top-2 head suppression at each toxic ratio.
Source
Orig
Adv
Transfer
Toxic cls
bert_jigsaw
90.5%
88.3%
3.1%
26.3%
roberta_jigsaw
90.5%
90.0%
1.8%
18.9%
bert_toxigen
68.5%
60.6%
14.6%
42.2%
roberta_toxigen
68.5%
63.6%
8.3%
30.8%
Appendix
Table 13: Transfer attack results: BERT/RoBERTa adversaries vs. Llama Guard 2.
Figure 10: Llama Guard 2 activation patching heatmaps (32 layers × 32 heads). Left: clean-input Δ Loss; right: adversarial-input Δ Loss. The clean circuit concentrates in layers 7–14; the adversarial circuit concentrates in layers 0–1. The two circuits are fully decoupled.
Figure 11: Bottleneck layer depth (as a percentile of model depth) for the top vulnerable head across all five model runs. BERT: 75th percentile (L9/12); RoBERTa: 50th percentile (L6/12); Llama Guard 2: 3rd percentile (L1/32). The bottleneck migrates to earlier layers in larger, decoder-based models.