Industrial content moderation must process massive message streams under tight latency constraints, yet most abusive language (AL) detection systems rely on sentence-level classification (ALC), which neither localizes abusive spans nor identifies who is targeted. We define Abusive Language Identification (ALI) as a sequence-labeling task that jointly extracts AL spans and target mentions, and assess whether this approach can be used for text moderation. On a pilot corpus drawn from a production moderation pipeline, we compare ALI with ALC on cross-domain generalization and implicit abuse, and we also evaluate AL and target span detection. ALI remains competitive with ALC while providing localized outputs for moderators, with a modest and configuration-sensitive advantage on implicit abuse. Exact AL boundaries and target spans remain difficult to recover. We complement this comparison with a qualitative analysis and discuss perspectives on complete target--span linking and on structured benchmarks for ALI.
Figures & tables
Figure 1: Annotation examples for each class: Abusive terms , Offensive speech , Threat , and Hate speech . Text highlighted in blue designates the target, and other highlighted text designates AL (shown here to the left of the sentences). The arrow represents the association between AL and its target.
Models
Training sets
JL-Hate AL
JL-Hate Target
TRuST AL
F1-exact
F1-partial
F1-token
F1-exact
F1-partial
F1-token
F1-exact
F1-partial
F1-token
fBERT
Ours
18.20 ±0.40
51.85 ±0.34 †
67.72 ±0.55
–
–
–
13.62 ±1.04
61.23 ±0.80
48.53 ±0.36
Ours (with target)
17.49 ±0.90
51.75 ±0.64
67.66 ±0.43
24.08 ±0.92 †
38.86 ±0.37 †
66.17 ±0.64 †
13.70 ±1.25
59.45 ±0.71
48.99 ±0.89
TBO
26.52 ±0.35
51.61 ±0.88
71.40 ±0.55
–
–
–
27.50 ±0.88
60.09 ±0.57
49.31 ±0.78 †
TBO (with target)
25.93 ±0.48
50.85 ±0.44
71.03 ±0.48
18.49 ±1.13
29.82 ±0.78
61.05 ±0.15
27.58 ±1.66
62.15 ±0.70
48.25 ±0.30
HateBERT
Ours
19.07 ±0.83
54.09 ±1.08
67.98 ±0.82
–
–
–
12.81 ±0.32
62.74 ±0.25
48.01 ±0.21
Table 1: Median F1 ± std (%, 3 runs) for AL/Target spans on JL-Hate and AL spans on TRuST. Bold text marks the best score for each F1 type within a backbone, and the † symbol is used to indicate that the corresponding score is significantly better than the other scores of the same backbone.
Models
Training sets
HatEval
IHC
JL-Hate
TRuST
ALI
ALC
ALI
ALC
ALI
ALC
ALI
ALC
Arango et al. (2019)
–
54.1
–
–
–
–
–
–
Caselli et al. (2020)
–
59.1
–
–
–
–
–
–
Bose et al. (2021)
–
57.8
–
–
–
–
–
–
Yuan and Rizoiu (2025)
–
64.5
–
–
–
–
–
–
fBERT
Ours
54.64 ±0.80 ∗
53.28 ±0.11
63.21 ±0.83
62.44 ±0.26
69.82 ±0.81
69.01 ±0.67
71.97 ±0.43
74.22 ±0.07 †
Table 2: Median macro-F1 ± std (%, 3 runs) of ALI and ALC trained on single sources. Average per task is the mean of the nine medians (published HatEval scores excluded). Bold text marks the best score for each approach within a model and test set; the ∗ symbol is used to indicate that the corresponding score is significantly better in the ALI–ALC comparison for the same training setting, and the † symbol is used to indicate that it is significantly better than the other training settings of the same model and approach.
Figure 2: Example outputs for RoBERTa ALI and ALC models trained on different corpora. For ALC models, LIME explanations are shown: words highlighted in orange tend to predict AL, and those highlighted in blue the opposite (normal discourse). The more a word is highlighted, the more impact it has on the class predicted by the ALC model. For ALI models, ALI predictions are highlighted. Both examples shown are from the IHC test set.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Neutral
Implicit AL
Explicit AL
Sentences (n)
337
123
542
Statistics on sentences
Max. length (tokens)
101
80
90
Avg. length (tokens)
26.64
22.73
25.32
Statistics on AL tokens
Max. AL tokens per sentence
∅
29
61
Appendix
Table 3: Pilot corpus statistics by class. AL tokens are tokens tagged as part of an AL span.
Models
Training sets
Implicit AL
Explicit AL
ALI
ALC
ALI
ALC
fBERT
Ours
57.51 ±0.72
57.33 ±0.60
72.06 ±0.61 ∗
70.56 ±0.28
HX
62.41 ±0.30 ∗
58.02 ±1.06
80.35 ±0.43
78.95 ±0.20
TBO
54.02 ±0.06
49.34 ±0.72
65.87 ±0.65 ∗
46.63 ±1.68
Ours + HX
64.50 ±1.19 ∗
58.12 ±0.70
79.32 ±0.17
77.94 ±0.18
Ours + TBO
42.31 ±2.51
48.97 ±0.72
60.37 ±2.48 ∗
48.42 ±1.50
Appendix
Table 4: Median macro-F1 ± std (%, 3 runs) for ALI vs ALC on IHC subsets ( implicit +neutral and explicit +neutral). Average per task is the mean of the medians shown. Bold text marks the best score for each approach within an abuse type and model; the ∗ symbol is used to indicate that the corresponding score is significantly better in the ALI–ALC comparison for the same training setting, and the † symbol is used to indicate that it is significantly better than the other training settings of the same model and approach.
Models
Training sets
HatEval
IHC
JL-Hate
TRuST
ALI
ALC
ALI
ALC
ALI
ALC
ALI
ALC
fBERT
Ours
54.64 ±0.80 ∗
53.28 ±0.11
63.21 ±0.83
62.44 ±0.26
69.82 ±0.81
69.01 ±0.67
71.97 ±0.43
74.22 ±0.07
HX
67.15 ±0.08 †
67.87 ±0.30 †
68.00 ±0.21 ∗
65.70 ±0.86
67.44 ±0.57
71.14 ±3.98 ∗
68.34 ±0.53
70.62 ±1.98
TBO
50.06 ±0.38 ∗
48.44 ±1.64
59.90 ±3.21 ∗
42.55 ±1.95
69.88 ±1.34 ∗
43.62 ±0.55
68.28 ±0.21 ∗
43.35 ±1.03
Ours + HX
60.98 ±0.54
66.96 ±0.14 ∗
70.13 ±0.70 ∗†
65.82 ±0.31
74.84 ±0.49
76.37 ±0.54 †
70.86 ±0.03
73.32 ±0.39
Ours + TBO
53.09 ±0.31 ∗
48.14 ±2.38
44.62 ±3.21 ∗
43.00 ±1.29
66.89 ±0.49 ∗
43.39 ±0.55
61.54 ±1.24
42.23 ±1.25
Appendix
Table 5: Median macro-F1 ± std (%, 3 runs) for ALI vs ALC on HatEval, IHC, JL-Hate, and TRuST (all training combinations; the main text keeps Ours / HX / TBO only; TRuST: out of domain, N=982 ). Average per task is the mean over the 21 configurations shown. Bold text marks the best score for each approach within a model and test set; the ∗ symbol is used to indicate that the corresponding score is significantly better in the ALI–ALC comparison for the same training setting, and the † symbol is used to indicate that it is significantly better than the other training settings of the same model and approach.
Models
Training sets
JL-Hate AL
JL-Hate Target
TRuST AL
F1-exact
F1-partial
F1-token
F1-exact
F1-partial
F1-token
F1-exact
F1-partial
F1-token
fBERT
Ours
18.20 ±0.40
51.85 ±0.34
67.72 ±0.55
–
–
–
13.62 ±1.04
61.23 ±0.80
48.53 ±0.36
Ours (with target)
17.49 ±0.90
51.75 ±0.64
67.66 ±0.43
24.08 ±0.92 †
38.86 ±0.37 †
66.17 ±0.64 †
13.70 ±1.25
59.45 ±0.71
48.99 ±0.89
HX
13.12 ±0.68
36.31 ±0.12
61.93 ±0.16
–
–
–
17.68 ±0.90
52.46 ±0.43
46.66 ±0.25
TBO
26.52 ±0.35
51.61 ±0.88
71.40 ±0.55
–
–
–
27.50 ±0.88
60.09 ±0.57
49.31 ±0.78 †
TBO (with target)
25.93 ±0.48
50.85 ±0.44
71.03 ±0.48
18.49 ±1.13
29.82 ±0.78
61.05 ±0.15
27.58 ±1.66
62.15 ±0.70
48.25 ±0.30
Appendix
Table 6: Median F1 ± std (%, 3 runs) for AL and target spans on JL-Hate and AL spans on TRuST (all training settings). Bold text marks the best score for each F1 type within a backbone, and the † symbol is used to indicate that the corresponding score is significantly better than the other scores of the same backbone.
In recent years social media has become an increasingly popular tool for communication. People use it to share their ideas, exchange information, and discuss thoughts. Given its prevalence and widespread reach, social media must remain a safe space for people. Content generated on social media can be abusive and it has become increasingly important to detect such content. In this paper, we use a language-based preprocessing and an ensemble of several models and analyze their performance of abusive comment detection. Through extensive experimentation, we propose a pipeline that minimizes the false-positive rate (marking non-abusive as abusive) so that these systems can detect abusive comments without undermining the freedom of expression.
Pranshu Rastogi, Madhav Mathur, Ramaneswaran S +1
Department of ICE, NSUT Delhi · Department of CSE, JIIT Noida · Department of IT, VIT Vellore +1
Chinese discriminatory-language detection is challenging because harmful intent is often implicit and context-dependent. We propose MAAM (Myopia--Astigmatism Anchor Mechanism), a lightweight, model-agnostic framework inspired by functional visual blur: rather than preserving every token equally, MAAM retains discrimination-relevant semantic anchors and calibrates them with C--I--S contextual priors (Contextual Tone, Group Identity, and Stance Polarity). We also introduce ChLGBT, to our knowledge the first Chinese LGBT-focused discriminatory-language dataset, with 8,120 manually annotated samples and three ordinal labels: explicit bias, implicit bias, and emotional intensity. Across strong encoder baselines, MAAM improves all three prediction dimensions, with consistent gains in accuracy, F1, Brier score, and expected calibration error. Compared with frontier LLM baselines under zero-shot and few-shot prompting protocols, MAAM remains competitive while offering stronger compactness and stability. These results suggest that interpretable anchor preservation and contextual calibration provide a practical alternative to heavier model scaling for Chinese discriminatory-language assessment.
Yuxin Fu, Shijing Si
School of Economics and Finance, Shanghai International Studies University · Shanghai, China
As large language models (LLMs) increasingly mediate both content generation and moderation, linguistic evasion strategies known as Algospeak have intensified the coevolution between evaders and detectors. This research formalizes the underlying dynamics grounded in a joint action model: when Algospeak increases, detectability and understandability decrease. Further, the concept of Majority Understandable Modulation (MUM) is introduced and defined as the modulation level at which additional evasive alteration increases detector evasion but loses comprehension for the majority of recipients. To empirically probe this trade-off, we introduce a reproducible framework that can be used to create meaning-preserving, Algospeak-style variants, based on an existing taxonomy and with tunable modulation levels. Using COVID-19 disinformation as a first proof-by-example setting, we construct a reference dataset of 700 modulated items, drawn from twenty base sentences across five modulation levels and seven strategies. We then run two linked evaluations with seven different language models: one testing for interpretation through meaning recovery and one for disinformation detection through classification. Curve fitting over modulation levels yields an estimate of the Majority Understandable Modulation threshold and enables sensitivity analyses across strategies and models, see Figure 1. Results reveal the characteristic relationships between understandability and modulation. This study lays the groundwork for understanding the dynamics behind Algospeak and provides the framework, dataset, and experimental setups described.