cs.CLOct 5, 2026

From Abusive Language Classification to Sequence Labeling Identification

Authors: Nicolas Zampieri, Ignacio Lopez, Manon Girard, Jeremy Auguste

Organizations: Everdian / Marseille, France

Abstract

Industrial content moderation must process massive message streams under tight latency constraints, yet most abusive language (AL) detection systems rely on sentence-level classification (ALC), which neither localizes abusive spans nor identifies who is targeted. We define Abusive Language Identification (ALI) as a sequence-labeling task that jointly extracts AL spans and target mentions, and assess whether this approach can be used for text moderation. On a pilot corpus drawn from a production moderation pipeline, we compare ALI with ALC on cross-domain generalization and implicit abuse, and we also evaluate AL and target span detection. ALI remains competitive with ALC while providing localized outputs for moderators, with a modest and configuration-sensitive advantage on implicit abuse. Exact AL boundaries and target spans remain difficult to recover. We complement this comparison with a qualitative analysis and discuss perspectives on complete target--span linking and on structured benchmarks for ALI.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Multi-Stage Training for Abusive Comment Detection in Indic Languages

    May 21, 2026Pranshu Rastogi, Madhav Mathur, Ramaneswaran S +1Hate Speech DetectionContent Moderation

  2. MAAM: Anchor-Preserving Compression and Contextual Calibration for Chinese Discriminatory Language Detection

    Jun 8, 2026Yuxin Fu, Shijing SiSemantic AnchorData Compression Methods

  3. Algospeak, Hiding in the Open: The Trade-off Between Legible Meaning and Detection Avoidance

    May 7, 2026Jan Fillies, Ronald E. Robertson, Jeffrey HancockLinguisticsDeception