cs.CYApr 14, 2026

The Enforcement and Feasibility of Hate Speech Moderation

Authors: Manuel TonneauDylan ThurgoodDiyi LiuNiyati MalhotraVictor Orozco-OlveraRalph SchroederScott A. HaleManoel Horta Ribeiro+2 more

Organizations: University of Oxford · University of Copenhagen · 2World Bank Group · 5Princeton University · 3New York University

Abstract

Online hate speech is associated with harms ranging from deteriorating mental health to violence, yet how consistently platforms moderate hate, and whether enforcement is feasible at scale, remain poorly understood. We audit hate speech moderation on Twitter (now X) using 540,000 tweets annotated by trained native speakers, representative of a full day on the platform. Five months after posting, 80% of hateful tweets, including violent ones, remained online. Removal was only marginally more likely than for non-hateful tweets, far below scams or adult content, and insensitive to severity and reach. Automated detection could not reliably classify hate but ranked it highly, enabling human triage. Simulating this workflow, current staffing curbed little exposure, yet substantial reductions proved financially feasible, far below applicable regulatory fines. Persistent hate reflects resource allocation, not technical limits.

Explore similar work

CardsList
  1. When Hate Meets Facts: LLMs-in-the-Loop for Check-worthiness Detection in Hate Speech

    Mar 26, 2026Nicolás Benjamín Ocampo, Tommaso Caselli, Davide CeolinHate SpeechHate