External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing
Authors: Kingshuk Gupta, Davide Buscaldi
Organizations: LIX, École Polytechnique · School of Electrical and Electronic Engineering, Nanyang Technological University · LIPN, Sorbonne Paris Nord
As Large Language Models (LLMs) increasingly serve as foundational reasoning engines, their tendency to hallucinate remains a critical vulnerability. While recent internal state probes offer a promising alternative to slow external retrieval systems, they largely reduce hallucination detection to a token-wise binary classification task, failing to capture the structured, sequential boundaries of semantic drift. Here, we introduce an internal hidden state framework for fine-grained, span-level hallucination detection. By inspecting layer-wise activation patterns, we attempt to detect the exact hallucination onset and continuation tokens in an LLM generation. Our experiments show that this approach successfully isolates hallucination onsets, achieving substantial improvements in Precision-Recall AUC over random baselines despite extreme class imbalance. Ultimately, we propose a novel cross-model detection framework in which one model observes the internal representations elicited by another model's generation. We find that an external observer can match or exceed a generator's self-detection of its own hallucination onsets, including when the observer is the smaller model, suggesting that self-detection is not the ceiling for onset localisation.
Figures & tables
Figure 1: Token-Level Hallucination Span Detection Framework. Phase 1 identifies critical layers offline to supply the extraction targets for Phase 2, while Phase 2 processes this continuous feature sequence through the trained classifier to output structurally valid span tags
Figure 2: Critical layer analysis of the SmolLM2 (1.7B) architecture across its Attention and MLP blocks. The bars represent the cosine similarity between normal and hallucinated feature vectors at each layer depth. Layers highlighted in yellow denote the selected critical layers ( Ccritical ), where sharp geometric deviations reveal the onset of semantic drift.
Dataset (Model)
Safe (O)
B-HALL
I-HALL
Total
PsiloQA (SmolLM2)
7,797
265
2,923
10,985
PsiloQA (TinyLlama)
13,022
707
18,288
32,017
RAGTruth (Mistral)
11,241
57
616
11,914
Table 1: Token-level distribution of the BIO tags across the test sets of the three evaluated datasets. Each dataset was uniformly partitioned into standard 80/20 train/test splits prior to hidden state extraction
SmolLM2 (1.7B)
TinyLlama (1.1B)
Mistral (7B)
Architecture
P
R
F1
P
R
F1
P
R
F1
Onset Token ( B-HALL )
MLP Baseline
0.32
0.39
0.35
0.29
0.36
0.32
0.19
0.26
0.22
BiLSTM
0.29
0.40
0.33
0.28
0.38
0.32
0.25
0.25
0.25
BiLSTM-CRF
0.42
0.24
0.31
0.32
0.30
0.31
0.38
0.18
0.24
BiLSTM-Attn-CRF
0.42
0.30
0.35
0.41
0.28
0.34
0.41
0.21
0.28
Table 2: Intra-model span detection performance across three LLM architectures. Precision (P), Recall (R), and F1-scores are reported exclusively for the minority hallucination classes to highlight the detection of semantic drift. The Safe (O) majority class is excluded to prevent metric inflation. All experiments were conducted using the top-5 Attention and top-5 MLP critical layers.
Token Class
Model
AUROC
PR-AUC
Onset ( B-HALL )
MLP Baseline
0.900
0.330
Ogasa ( Ogasa and Arase, 2025 )
0.931
0.371
BERT-CRF
0.895
0.362
Random Baseline: 0.024
Continuation ( I-HALL )
MLP Baseline
0.716
0.460
BERT-CRF
0.735
0.446
Table 3: Threshold-independent onset diagnostics on the SmolLM2 test set, including the intra-model attention baseline of Ogasa and Arase (2025) . The Random Baseline equals the class prevalence. Single-seed.
Cross-Model Framework Evaluation
Architecture
QwenSmol
Gemma2Smol
Gemma4Smol
Gemma2Mist
SmolMist
P
R
F1
P
R
F1
P
R
F1
P
R
F1
P
R
F1
Onset Token ( B-HALL )
MLP Baseline
0.37
0.41
0.39
0.42
0.57
0.48
0.34
0.67
0.45
0.10
0.39
0.16
0.11
0.32
0.17
BiLSTM
0.39
0.42
0.41
0.45
0.49
0.47
0.39
0.54
0.45
0.29
0.46
0.35
0.17
0.47
0.25
BiLSTM-CRF
0.53
0.29
0.38
0.55
0.36
0.43
0.50
0.33
0.40
0.40
0.32
0.35
0.33
0.14
0.20
Table 4: Comprehensive cross-model hallucination span detection performance. The detection pairs QwenSmol utilizes Qwen (1.5B) to evaluate SmolLM2 (1.7B) generations, Gemma2Smol and Gemma4Smol utilize Gemma 2 (2B) and Gemma 4 (E2B) respectively to evaluate SmolLM2 (1.7B), Gemma2Mist utilizes Gemma 2 (2B) to evaluate Mistral (7B), and SmolMist utilizes the smaller SmolLM2 (1.7B) to evaluate Mistral (7B). Precision (P), Recall (R), and F1-scores are reported as before.
Architecture
Self
External
(Mistral)
(SmolLM2 → Mistral)
MLP
0.172
0.130
BiLSTM
0.205
0.193
BiLSTM-CRF
0.172
0.227
BiLSTM-Attn-CRF
0.221
0.242
BERT-CRF
0.043
0.079
Table 5: Onset ( B-HALL ) PR-AUC on RAGTruth: Mistral-7B self-detection vs. a smaller external observer (SmolLM2-1.7B), same test fold. Bold marks the higher value per row. Single-seed.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Component
Configuration
MLP
Hidden layers
din –1024–512–128–64–3
Activation
ReLU, Dropout 0.3
BiLSTM
Input dim
64 (projected)
LSTM hidden
32 / direction, 1 layer
Head
Linear 64→3
BiLSTM-CRF
Base
BiLSTM (above)
Appendix
Table 6: Layer configurations for each classifier, held fixed across datasets. din is the concatenated NAS input dimension, determined by the generator’s hidden size and critical-layer count. The BERT-CRF projector intermediate is widened to 4096 for TinyLlama.
Architecture
P
R
F1
PR-AUC
AUROC
MLP
0.360
0.558
0.438
0.380
0.854
BiLSTM
0.323
0.585
0.416
0.305
0.800
BiLSTM-CRF
0.492
0.340
0.402
0.252
0.784
BiLSTM-Attn-CRF
0.517
0.340
0.410
0.313
0.783
Appendix
Table 7: Onset ( B-HALL ) detection on SmolLM2 using a single fixed random layer set, not the critical layers used elsewhere in this paper. One draw, not an average over draws; single classifier seed. Corresponding critical-layer figures are given in the text below.
Hyperparameter
Value
Scope
Optimizer
Adam / AdamW
All architectures.
Learning rate
1×10−3
Classifier heads and projectors.
BERT encoder LR
2×10−5 – 5×10−5
After a 1-epoch warm-up.
B-HALL weight
20 – 300
Tuned per dataset by onset sparsity.
I-HALL weight
5 – 30
Tuned per dataset.
Training epochs
5 – 15
By architecture and token volume.
Appendix
Table 8: Representative self-evaluation training hyperparameters. Class weights and epochs vary by dataset to accommodate differing onset sparsity; exact per-model values are released with our code.
Cross-Model Hyperparameter
Value
Scope
Auxiliary B-HALL weight
100 – 150
RAGTruth-Mistral onset penalty.
Auxiliary multiplier ( λ )
1 – 3
Token penalty vs. sequence logic.
BERT unfreeze epoch
5
Delayed for projector alignment.
Post-warmup BERT LR
1×10−5
Stable end-to-end tuning.
Gradient accumulation
4 steps
Memory boundary.
Appendix
Table 9: Training hyperparameters specific to the cross-model evaluation framework.
Large language models (LLMs) hallucinate with confidence: their outputs can be fluent, authoritative, and simply wrong. In medical, legal, and scientific applications this failure causes direct harm, and detecting it from internal model states offers a path to safer deployment. A growing body of work reports that this problem is increasingly tractable, with recent methods achieving high detection performance on widely used benchmarks. We show, however, that much of this apparent progress does not survive scrutiny. Four of the six corpora embed the ground-truth answer directly in the input prompt. A naïve text-similarity baseline we call \textsc{TxTemb} exploits this to achieve near-perfect detection scores without any access to model internals. To measure what genuine detection capability remains once these artifacts are controlled, we conduct a large-scale evaluation spanning twenty-two detection methods, twelve open-source models spanning six architectural families, and six corpora. We further introduce \textbf{DRIFT}, a supervised probe over inter-layer hidden-state transitions, as a point of comparison for live-generation detection. Our findings suggest that the field's reported progress on hallucination detection is substantially explained by benchmark construction artifacts in widely used corpora, and that the majority of established baselines perform near chance under controlled conditions; the consistent exceptions are SAPLMA and DRIFT, both supervised probes on upper-layer hidden states.
Recent hallucination detection techniques in large language models (LLMs) focus on directly extracting features from a model's internal representations and training a classifier on these features to detect hallucinations, demonstrating promising results. Notwithstanding this advancement, most internal-state hallucination detection techniques have been explored predominantly in English, raising the question of whether such internal signals generalize across different languages and domains. To address this gap, we present CrossHallu, the first study to evaluate the cross-lingual and cross-domain generalization of hallucination detection using internal representations from six LLMs on the generative question-answering task. We conduct a systematic Arabic <-> English evaluation using TruthfulQA, an Arabic translated version of TruthfulQA, and HalluScore. This evaluation encompasses monolingual training and testing, cross-lingual transfer, cross-domain transfer, and combined cross-lingual and cross-domain transfer. The results reveal that internal-state hallucination signals in LLMs transfer across languages and domains for most models, with cross-lingual performance highly dependent on both class separability and language alignment in the feature space, whereas cross-domain transfer within Arabic varies depending on the training and testing datasets used for the hallucination detector. The code is publicly available at https://github.com/aishaalansari57/CrossHal.
Aisha Alansari, Malak Alkhorasani, Hamzah Luqman
King Fahd University of Petroleum and Minerals · Imam Abdulrahman bin Faisal University
Hallucination in large language models (LLMs), defined as the generation of factually incorrect or unsupported content, remains a critical barrier to reliable deployment. We present BEACON (Behavioral Entropy Aggregation for Cross-model hallucination detectiON), a black-box hallucination detection framework that operates purely on model outputs without requiring access to internal representations or external knowledge bases. BEACON extracts a 31-dimensional feature vector from structured multi-pass generation, integrating NLI-based semantic entropy, embedding geometry, chain-of-thought consistency, and paraphrase stability signals. A gradient-boosted classifier trained on 7,617 labeled examples across seven benchmarks achieves 0.8123 +/- 0.0102 AUROC (95% CI: 0.7632-0.8251), outperforming standalone semantic entropy (+0.2298) and SelfCheckGPT-style consistency baselines (+0.2457). Feature importance analysis shows that hallucination is inherently multi-dimensional, requiring combined uncertainty signals. An efficient 5-call variant achieves 0.7795 AUROC, enabling practical deployment across black-box LLM APIs.
Naveen Bera, Pulijala Sai Nikhila, Kondaguduru Abhiram +6