cs.CLAug 28, 2026

The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice

Authors: Jungseob Lee, Jaehyung Seo, Heuiseok Lim

Organizations: Korea University · Konkuk University

Abstract

Hidden-state probes effectively detect LLM hallucinations, but the geometry of the signal remains poorly characterized, driving increasingly complex probe architectures. Across three 7B-scale models and three datasets in a paired-example paradigm, we find the signal overwhelmingly dominated by a single mean-shift component, and removing this direction collapses detection to chance. Shrinkage linear discriminant analysis closes about 73% of the gap between 1D and full-dimensional classifiers, so apparent architectural complexity largely reflects high-dimensional covariance estimation difficulty rather than exploitable non-linearity. A simple L2-regularized logistic regression (0.952 AUROC) bounds or outperforms twelve controlled architectural alternatives, and our multi-layer aggregation exceeds CLAP cross-layer attention probing under matched paradigm. Because the signal spans a contiguous layer band, LayerMix aggregates it to match oracle-layer performance without oracle access. Our claims characterize the geometry within the controlled paired-example paradigm. Our code is available at https://github.com/js-lee-AI/LayerMix.

Explore similar work

CardsList
  1. Automatic Layer Selection for Hallucination Detection

    May 25, 2026Xinpeng Wang, William X. Cao, Andrew Gordon Wilson +1Hallucination in Language ModelsLLM Hallucination Detection

  2. FLaG: Fine-Grained Latent Grouping for Hallucination Detection

    May 29, 2026Wentao Ye, Liyao Li, Zhiqing Xiao +6Latent Variable ModelsLLM Hallucination Detection