Adversarial perturbations can reduce state-of-the-art toxicity classifiers to near-zero accuracy, yet existing defences treat models as black boxes. We apply mechanistic interpretability to toxicity classification for the first time, identifying the internal attention-head circuits responsible for both correct classification and adversarial vulnerability. Across a 2×2 factorial study (BERT × RoBERTa) × (Jigsaw × ToxiGen), extended to Llama Guard~2 (8B), we show that zeroing a single attention head recovers up to 70.4 pp of adversarial accuracy for RoBERTa on Jigsaw and 37.3 pp for Llama Guard 2 on ToxiGen, at ≤0.6 pp clean cost. Vulnerable heads generalise to held-out examples within ≤1 pp, and a class-imbalance sweep confirms they act as selective toxic-class detectors. Head suppression matches or outperforms adversarial training on Jigsaw; data augmentation dominates on ToxiGen: a dataset-specific reversal explained by whether the classifier encodes a concentrated bottleneck or a distributed circuit. Demographic analysis across 20 Jigsaw and 13 ToxiGen minority groups reveals structurally unequal adversarial vulnerability, exposing mechanistically traceable fairness gaps in current toxicity classifiers.
Figures & tables
Input
Text
Label
Prediction
Regular
“Your observations are consistent with those of most imbeciles”
Toxic
Toxic ✓
Adversarial
“your terms are consistent with those of most imbeciles”
Toxic
Non-toxic ✗
Regular
“Homosexuality is a sin and a choice, not a genetic factor. Homosexuality is a behavioral disorder conflicting with one’s God-given sex”
Toxic
Toxic ✓
Adversarial
“homosexuality is a sin and a choice, not a genetic factor. homosexual is a behavioral disorder conflicting with one ’ s god - given sex”
Toxic
Non-toxic ✗
Regular
“the immigration system in the USA is full of absurdities”
Non-toxic
Non-toxic ✓
Adversarial
“the immigration system in the us is full of absurdities”
Non-toxic
Toxic ✗
Table 1: Adversarial input example showing semantic equivalence but different model predictions.
Figure 1: End-to-end pipeline from adversarial attack to targeted head suppression across six model–corpus combinations. Fine-tuned classifiers ( Stage 1 ) are attacked via PGD BERT-Attack ( Stage 2 ); activation patching identifies the bottleneck head set H⋆ ( Stage 3 ); zero-ablating H⋆ improves adversarial accuracy while preserving clean performance ( Stage 4 ). The dashed arc denotes evaluation of the suppressed model f~θ on both clean and adversarial inputs.
Figure 2: Clean-input activation patching: per-head Δ Loss (increase = head is crucial) across all four runs in a 2 × 2 grid (rows: BERT/RoBERTa; columns: Jigsaw/ToxiGen). Warmer colours indicate higher impact. Jigsaw models show a single dominant hot cell (L9H5, L6H1); ToxiGen models are diffuse.
Run
Architecture
Clean Acc (balanced)
bert_jigsaw
BERT
78.35%
roberta_jigsaw
RoBERTa
79.88%
bert_toxigen
BERT
82.40%
roberta_toxigen
RoBERTa
85.98%
Table 2: Best checkpoint accuracy on the 10k balanced test set after 3 training epochs.
Run
Top head
Δ Loss
Acc drop
bert_jigsaw
L9H5
+0.132
− 5.03 pp
roberta_jigsaw
L6H1
+0.178
− 11.58 pp
bert_toxigen
L11H2
+0.020
− 0.55 pp
roberta_toxigen
L4H1
+0.011
− 0.54 pp
Table 3: Top crucial head per run identified by clean activation patching.
Run
Top head
Δ Loss
k=1
k=5
bert_jigsaw
L9H5
− 0.562
+50.5 pp
+84.9 pp
roberta_jigsaw
L6H1
− 0.721
+70.4 pp
+76.7 pp
bert_toxigen
L9H5
− 0.069
+23.6 pp
+34.1 pp
roberta_toxigen
L6H1
− 0.174
+23.5 pp
+38.9 pp
Table 4: Top vulnerable head and adversarial accuracy recovery ( Δ from near-zero baseline) at k =1 and k =5 heads suppressed.
Figure 3: Adversarial accuracy recovery as k heads are suppressed. Jigsaw curves (left) rise steeply at k =1; ToxiGen curves (right) are shallower across all k .
Run
Heads
Base
Supp.
Δ
bert_jigsaw
L9H5, L8H8
0.17%
67.22%
+67.1 pp
roberta_jigsaw
L6H1, L5H2
0.09%
56.79%
+56.7 pp
bert_toxigen
L9H5, L11H3
0.16%
25.66%
+25.5 pp
roberta_toxigen
L6H1, L4H10
0.33%
27.46%
+27.1 pp
Table 5: Overall adversarial accuracy before and after top-2 head suppression on the demographic adversarial pool.
Figure 4: Per-group adversarial accuracy after top-2 head suppression. Groups sorted by suppressed accuracy; ∗ = n<30 .
Figure 5: Adversarial (solid) and clean (hatched) accuracy for all robustness methods. Head suppression wins on Jigsaw; augmentation wins on ToxiGen.
Run
Method
Clean
Adv
bert_jig.
Baseline
94.88%
0.00%
Head supp.
94.71%
66.00%
FGM
94.96%
42.74%
Augmentation
88.74%
65.81%
rob._jig.
Baseline
94.45%
0.00%
Head supp.
93.89%
73.14%
Table 6: Clean and adversarial accuracy for all robustness methods. Bold = best adv. accuracy per run. Clean accuracy on unbalanced natural test.
Run
Full (k=1)
Held-out B
Diff
bert_jigsaw
50.54%
50.90%
+0.36 pp
roberta_jigsaw
70.64%
71.61%
+0.97 pp
bert_toxigen
23.79%
23.10%
− 0.69 pp
roberta_toxigen
23.62%
24.02%
+0.40 pp
Table 7: Split-sample validation: adversarial accuracy at k =1 on full set vs. held-out Set B. All differences ≤ 1 pp.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: PGD BERT-Attack pipeline: gradients are computed w.r.t. input token embeddings and used to guide substitution via a BERT MLM, iterating within a norm-bounded perturbation ball.
Figure 7: Adversarial-input activation patching: per-head Δ Loss on successfully-attacked examples. Red cells = heads whose ablation helps recover adversarial accuracy (negative Δ Loss). L9H5 (BERT) and L6H1 (RoBERTa) are the dominant bottleneck heads for Jigsaw; ToxiGen shows a more distributed pattern.
Figure 8: FGM adversarial training curves (10 epochs). Jigsaw models collapse catastrophically after the best epoch; ToxiGen models are stable throughout.
BERT Jigsaw
Ep.
Loss
Clean
Adv
1
0.122
94.98%
30.82%
2
0.118
95.10%
35.98%
3
0.104
94.96%
42.74%
4
0.092
94.63%
10.93%
5
0.082
94.01%
5.96%
Appendix
Table 8: FGM training curves: BERT Jigsaw (best ep. 3) and RoBERTa Jigsaw (best ep. 1). Bold = best epoch.
BERT ToxiGen
Ep.
Loss
Clean
Adv
1
0.354
82.57%
31.84%
2
0.328
82.93%
33.98%
3
0.276
83.10%
32.42%
4
0.228
83.16%
34.96%
5
0.184
83.04%
33.59%
Appendix
Table 9: FGM training curves: BERT ToxiGen (best ep. 4) and RoBERTa ToxiGen (best ep. 9). Bold = best epoch.
Run
k
Heads (Set A)
Held-out Adv Acc
bert_jigsaw
1
L9H5
50.90%
2
+L8H8
65.74%
3
+L11H6
76.60%
5
+L6H7, L10H11
84.68%
roberta_jigsaw
1
L6H1
71.61%
2
+L5H2
75.08%
Appendix
Table 10: Set B held-out cumulative adversarial accuracy (exp06).
Figure 9: Class-imbalance sweep: clean cost (left) and adversarial benefit (right) of top-2 suppression at five toxic ratios. Jigsaw clean cost rises monotonically (22 × spread for RoBERTa), confirming selective toxic-class detection.
Run
10%
30%
50%
70%
90%
bert_jigsaw
+0.88
+5.15
+9.62
+14.28
+18.63
roberta_jigsaw
+1.67
+10.87
+19.53
+27.93
+36.88
bert_toxigen
− 0.63
− 0.02
+0.90
+1.45
+1.82
roberta_toxigen
+2.35
+1.33
+0.45
− 0.22
− 1.03
Appendix
Table 11: Clean accuracy cost ( Δclean , pp) of top-2 head suppression at each toxic ratio.
Run
10%
30%
50%
70%
90%
bert_jigsaw
+66.51
+51.31
+36.38
+22.31
+7.28
roberta_jigsaw
+87.06
+67.59
+48.56
+29.58
+9.90
bert_toxigen
+31.72
+27.28
+23.92
+19.57
+15.52
roberta_toxigen
+27.30
+29.67
+32.67
+35.47
+38.35
Appendix
Table 12: Adversarial accuracy benefit ( Δadv , pp) of top-2 head suppression at each toxic ratio.
Source
Orig
Adv
Transfer
Toxic cls
bert_jigsaw
90.5%
88.3%
3.1%
26.3%
roberta_jigsaw
90.5%
90.0%
1.8%
18.9%
bert_toxigen
68.5%
60.6%
14.6%
42.2%
roberta_toxigen
68.5%
63.6%
8.3%
30.8%
Appendix
Table 13: Transfer attack results: BERT/RoBERTa adversaries vs. Llama Guard 2.
Figure 10: Llama Guard 2 activation patching heatmaps (32 layers × 32 heads). Left: clean-input Δ Loss; right: adversarial-input Δ Loss. The clean circuit concentrates in layers 7–14; the adversarial circuit concentrates in layers 0–1. The two circuits are fully decoupled.
Figure 11: Bottleneck layer depth (as a percentile of model depth) for the top vulnerable head across all five model runs. BERT: 75th percentile (L9/12); RoBERTa: 50th percentile (L6/12); Llama Guard 2: 3rd percentile (L1/32). The bottleneck migrates to earlier layers in larger, decoder-based models.
Production LLMs increasingly rely on toxicity-based moderation filters as a primary defense, assuming that harmful intent correlates with toxic surface wording. We show this assumption is fundamentally brittle: surface toxicity and adversarial intent can be decoupled by replacing as few as five tokens. We present OTTER (Obfuscated Toxicity-Evading Token Evolution for Rewriting), a black-box red-teaming framework requiring only standard API access, directly targeting the practical constraints of industry security audits. Evaluated on 457 AdvBench prompts across four GPT models, OTTER raises average ASR from 7.0% to 84.0%. We further provide the first quantitative analysis of the toxicity--bypass relationship and a per-category breakdown, translating our findings into actionable recommendations for classifier hardening in production deployments.
Guardrail Classifiers defend production language models against harmful behavior, but although results seem promising in testing, they provide no formal guarantees. Providing formal guarantees for such models is hard because "harmful behavior" has no natural specification in a discrete input space: and the standard epsilon-ball properties used in other domains do not carry semantic meaning. We close this gap by shifting verification from the discrete input space to the classifier's pre-activation space, where we define a harmful region as a convex shape enclosing the representations of known harmful prompts. Because the sigmoid classification head is monotonic, certifying the worst-case point is sufficient to certify the entire region, yielding a closed-form soundness proof without approximation in O(d) time. To formally evaluate these classifiers, we propose two constructions of such regions: SVD-aligned hyper-rectangles, which yield exact SAT/UNSAT certificates, and Gaussian Mixture Models, which yield probabilistic certificates over semantically coherent clusters. Applying this framework to three author-trained Guardrail Classifiers on the toxicity domain, every hyper-rectangle configuration returns SAT, exposing verifiable safety holes across all classifiers, despite seemingly high empirical metrics. Probabilistic GMM certificates also expose a divergent structural stability in how these models represent harm. While GPT-2 and Llama-3.1-8B maintain robust coverage of 90% and 80% across varying boundaries, BERT's safety guarantees prove uniquely volatile. This 'coverage collapse' to 55% at the optimal threshold reveals a sparsely populated safety margin in BERT, which only achieves full coverage by adopting an extremely conservative pessimistic threshold. These approaches combined, provide new insights on how effective Guardrail Classifiers really are, beyond traditional red-teaming.
Nikita Kezins, Urbas Ekka, Pascal Berrang +1
Delft University of Technology · University of Birmingham & Zeroth Research
Large language models (LLMs) frequently generate toxic content, posing significant risks for safe deployment. Current mitigation strategies often degrade generation quality or require costly human annotation. We propose CAUSALDETOX, a framework that identifies and intervenes on the specific attention heads causally responsible for toxic generation. Using the Probability of Necessity and Sufficiency (PNS), we isolate a minimal set of heads that are necessary and sufficient for toxicity. We utilize these components via two complementary strategies: (1) Local Inference-Time Intervention, which constructs dynamic, input-specific steering vectors for context-aware detoxification, and (2) PNS-Guided Fine-Tuning, which permanently unlearns toxic representations. We also introduce PARATOX, a novel benchmark of aligned toxic/non-toxic sentence pairs enabling controlled counterfactual evaluation. Experiments on ToxiGen, ImplicitHate, and ParaDetox show that CAUSALDETOX achieves up to 5.34% greater toxicity reduction compared to baselines while preserving linguistic fluency, and offers a 7x speedup in head selection.
Yian Wang, Yuen Chen, Agam Goyal +1
Department of Computer Science University of Illinois, Urbana-Champaign Champaign, IL 61801