Lexicon-driven misogyny detectors cannot, by construction, distinguish a slur used against a woman from the same slur mentioned in counter-speech ("don't call her that") -- yet exactly this distinction governs whether moderation protects or silences the people discussing abuse. We study this problem in code-mixed Hinglish and make three contributions. First, we diagnose two evaluation artifacts on a publicly available redacted corpus: category-encoding anonymization placeholders leak the label (a no-learning rule scores 1.000), and even after they are neutralized misogynistic and benign comments occupy lexically disjoint registers, so bag-of-words reaches macro-F1 approximately 1.00 under random cross-validation but collapses under template-disjoint evaluation. Second, we release Hinglish-MGY-Diag, a deterministic generator and a 416-item / 163-minimal-pair contrast-set diagnostic across five linguistically motivated categories in which slur presence and gendered register are decorrelated from the label by construction. Third, we introduce a strict pair-consistency metric that credits a model only when both members of a minimal pair are correctly labelled. Five from-scratch classical baselines evaluated under construction-disjoint five-fold cross-validation reveal that the strongest model reaches 0.93 accuracy on the cleanest use-mention subset but only 0.82 consistency -- it still mislabels roughly one counter-speech pair in five. A frontier LLM used as an author-model ceiling attains 1.000 on all metrics, doubling as independent label validation and confirming the benchmark is a capability gradient rather than an adversarial wall. We release all code, data, the generator, and an arms-length LLM harness for reproducing every number.
Hate moderation is often evaluated as classification on clean English inputs, but deployed systems must route content to actions such as ALLOW, FLAG, or REVIEW. We study how this workflow changes under code-mixed inputs using a paired evaluation setting where the same underlying content is expressed as clean English and Tamil-English code-mix. Under thresholds tuned on clean English development data, code-mixed inputs produce substantial action instability, with a paired clean- to-code-mix decision flip rate of 0.265. The main workflow effects are increased review burden and increased false-flagging of non-hateful content: review rate rises from 0.138 to 0.297 and non-hate false-flag rate rises from 0.069 to 0.104. Tamil-only inputs show stronger degradation overall, suggesting a broader language-coverage limitation rather than the same code-mixed instability pattern. A simple disagreement-based deferral rule reduces automatic errors on stressed inputs, but only by increasing review load. These results show that workflow-level evaluation reveals moderation failures that standard classification summaries can miss.
Online hate speech and misinformation frequently overlap, yet NLP research has mainly treated them in isolation. While LLMs represent a scalable solution for assisting humans in the generation of counterspeech for both threats, zero-shot models frequently generate repetitive and vague responses, underscoring the need for high-quality examples to steer model generation. However, existing counterspeech datasets against the overlap of hate and misinformation are scarce and limited to single-turn English dialogues, while real-life interactions span across multiple turns and languages. To bridge this gap, we introduce the first large-scale, expert-curated, multilingual dataset of dialogues tackling the intersection of hate and misinformation. To ensure factual grounding, the dialogues are also anchored in verified external knowledge (i.e., fact-checking articles and NGO reports) and include document- and chunk-level span annotations, making it directly applicable for RAG systems. Covering five languages and targeting hate directed at seven marginalized groups, this novel resource enables the training and evaluation of more persuasive, factually grounded counterspeech models.
Helena Bonaldi, Genoveffa Martone, Marco Guerini
Fondazione Bruno Kessler, Italy · Università Cattolica del Sacro Cuore, Italy
Hidden-state hallucination probing - training a linear classifier on an LLM's internal activations to detect whether a generated answer is faithful to the input - is an active area of 2026 research, with recent work reporting 0.90-1.00 AUROC across several benchmarks and languages. However, none of this work has tested probes on code-mixed input, despite the fact that a huge population of chatbot users write in Hindi-English code-mixed text ("Hinglish"). We address this gap directly: does a hallucination probe trained on clean-language hidden states transfer to Hinglish, or does the signal degrade under code-mixing? We construct a 5,674-item Hindi/English/Hinglish QA benchmark, generate and label 17,022 model responses across three open-weight 7-8B LLMs (Qwen2.5-7B, Mistral-7B, Llama-3.1-8B), extract per-layer hidden states at two token positions, and train linear and MLP probes for in-distribution detection and cross-lingual transfer. We find that the hallucination signal survives code-mixing well: transfer AUROC ranges from 0.88 to 0.99, with gaps of mostly under 0.05 AUROC relative to in-distribution performance, and that Hindi-trained probes transfer to Hinglish more reliably than English-trained probes. As an independent, practically motivated finding, all three models hallucinate substantially more on Hindi and Hinglish than on English for matched facts. We release our code and synthetic Hinglish QA dataset to support further work on code-mixed hallucination detection.