cs.CLAug 25, 2026

Does the Truthfulness Signal Survive Code-Mixing? Probing Hidden States for Hallucination Detection in Hinglish

Authors: Tanveer Singh

Organizations: Plaksha University

Abstract

Hidden-state hallucination probing - training a linear classifier on an LLM's internal activations to detect whether a generated answer is faithful to the input - is an active area of 2026 research, with recent work reporting 0.90-1.00 AUROC across several benchmarks and languages. However, none of this work has tested probes on code-mixed input, despite the fact that a huge population of chatbot users write in Hindi-English code-mixed text ("Hinglish"). We address this gap directly: does a hallucination probe trained on clean-language hidden states transfer to Hinglish, or does the signal degrade under code-mixing? We construct a 5,674-item Hindi/English/Hinglish QA benchmark, generate and label 17,022 model responses across three open-weight 7-8B LLMs (Qwen2.5-7B, Mistral-7B, Llama-3.1-8B), extract per-layer hidden states at two token positions, and train linear and MLP probes for in-distribution detection and cross-lingual transfer. We find that the hallucination signal survives code-mixing well: transfer AUROC ranges from 0.88 to 0.99, with gaps of mostly under 0.05 AUROC relative to in-distribution performance, and that Hindi-trained probes transfer to Hinglish more reliably than English-trained probes. As an independent, practically motivated finding, all three models hallucinate substantially more on Hindi and Hinglish than on English for matched facts. We release our code and synthetic Hinglish QA dataset to support further work on code-mixed hallucination detection.

Explore similar work

Jul 4, 2026cs.CL

CrossHallu: Do Hallucination Signals Generalize Across Languages and Domains in Large Language Model's Internals?

Recent hallucination detection techniques in large language models (LLMs) focus on directly extracting features from a model's internal representations and training a classifier on these features to detect hallucinations, demonstrating promising results. Notwithstanding this advancement, most internal-state hallucination detection techniques have been explored predominantly in English, raising the question of whether such internal signals generalize across different languages and domains. To address this gap, we present CrossHallu, the first study to evaluate the cross-lingual and cross-domain generalization of hallucination detection using internal representations from six LLMs on the generative question-answering task. We conduct a systematic Arabic <-> English evaluation using TruthfulQA, an Arabic translated version of TruthfulQA, and HalluScore. This evaluation encompasses monolingual training and testing, cross-lingual transfer, cross-domain transfer, and combined cross-lingual and cross-domain transfer. The results reveal that internal-state hallucination signals in LLMs transfer across languages and domains for most models, with cross-lingual performance highly dependent on both class separability and language alignment in the feature space, whereas cross-domain transfer within Arabic varies depending on the training and testing datasets used for the hallucination detector. The code is publicly available at https://github.com/aishaalansari57/CrossHal.
Aug 28, 2026cs.CL

The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice

Hidden-state probes effectively detect LLM hallucinations, but the geometry of the signal remains poorly characterized, driving increasingly complex probe architectures. Across three 7B-scale models and three datasets in a paired-example paradigm, we find the signal overwhelmingly dominated by a single mean-shift component, and removing this direction collapses detection to chance. Shrinkage linear discriminant analysis closes about 73% of the gap between 1D and full-dimensional classifiers, so apparent architectural complexity largely reflects high-dimensional covariance estimation difficulty rather than exploitable non-linearity. A simple L2-regularized logistic regression (0.952 AUROC) bounds or outperforms twelve controlled architectural alternatives, and our multi-layer aggregation exceeds CLAP cross-layer attention probing under matched paradigm. Because the signal spans a contiguous layer band, LayerMix aggregates it to match oracle-layer performance without oracle access. Our claims characterize the geometry within the controlled paired-example paradigm. Our code is available at https://github.com/js-lee-AI/LayerMix.
May 30, 2026cs.LG

Hallucination Is Linearly Decodable from Mid-Layer Hidden States in Quantized LLMs

We investigate whether open-source LLMs encode a linearly separable truthfulness signal in their hidden states, and at which network depth this signal is strongest. Across three 77B--88B instruction-tuned models (Llama-3.1-8B, Mistral-7B, Qwen2.5-7B) loaded in 44-bit NF4 quantization, we extract per-layer hidden states on four hallucination benchmarks (TruthfulQA, HaluEval-QA, FEVER, and a controlled synthetic set) and compare four detection approaches: linear and MLP probes, INSIDE EigenScore, self-consistency, and attention entropy. A linear probe on a single mid-network layer achieves 0.9040.904--1.0001.000 AUROC on held-out splits, while sampling-based detectors do not exceed 0.5410.541 AUROC under the same protocol. The truthfulness signal is approximately linear: MLP probes rarely surpass linear probes by more than 0.010.01 AUROC. Peak probing layers fall in a consistent band across model families on natural-language benchmarks -- blocks~1313--1818 of~3232 for Llama and Mistral, and blocks~1919--2525 of~2828 for Qwen. First-block attention entropy provides a complementary signal in knowledge-grounded settings (0.8660.866--0.9410.941 AUROC on HaluEval-QA) at no additional inference cost. The low discriminability of sampling methods under this protocol reflects a structural mismatch between paired-label evaluation and the information these methods access, rather than an inherent limitation of those methods. Code and data are released for full reproducibility on a single 88,GB GPU.