cs.CLSep 16, 2025

Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content

Authors: Shaz Furniturewala, Arkaitz Zubiaga

Organizations: Center for Data Science, New York University · Queen Mary University of London

Abstract

Adversarial perturbations can reduce state-of-the-art toxicity classifiers to near-zero accuracy, yet existing defences treat models as black boxes. We apply mechanistic interpretability to toxicity classification for the first time, identifying the internal attention-head circuits responsible for both correct classification and adversarial vulnerability. Across a 2×\times2 factorial study (BERT ×\times RoBERTa) ×\times (Jigsaw ×\times ToxiGen), extended to Llama Guard~2 (8B), we show that zeroing a single attention head recovers up to 70.4 pp of adversarial accuracy for RoBERTa on Jigsaw and 37.3 pp for Llama Guard 2 on ToxiGen, at ≤\leq0.6 pp clean cost. Vulnerable heads generalise to held-out examples within ≤\leq1 pp, and a class-imbalance sweep confirms they act as selective toxic-class detectors. Head suppression matches or outperforms adversarial training on Jigsaw; data augmentation dominates on ToxiGen: a dataset-specific reversal explained by whether the classifier encodes a concentrated bottleneck or a distributed circuit. Demographic analysis across 20 Jigsaw and 13 ToxiGen minority groups reveals structurally unequal adversarial vulnerability, exposing mechanistically traceable fairness gaps in current toxicity classifiers.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. OTTER: A Red-Teaming System for Toxicity-Evading Jailbreak Prompt Optimization

    Jun 19, 2026Jerry Wang, Hsin-Ling Hsu, Yi-Cheng Lai +2ToxicityRed-Teaming

  2. Beyond Red-Teaming: Formal Guarantees of LLM Guardrail Classifiers

    May 11, 2026Nikita Kezins, Urbas Ekka, Pascal Berrang +1Large Language Model SafetyToxicity

  3. CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification

    Apr 16, 2026Yian Wang, Yuen Chen, Agam Goyal +1DetoxificationHarmful Language