External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing
Authors: Kingshuk Gupta, Davide Buscaldi
Organizations: LIX, École Polytechnique · School of Electrical and Electronic Engineering, Nanyang Technological University · LIPN, Sorbonne Paris Nord
As Large Language Models (LLMs) increasingly serve as foundational reasoning engines, their tendency to hallucinate remains a critical vulnerability. While recent internal state probes offer a promising alternative to slow external retrieval systems, they largely reduce hallucination detection to a token-wise binary classification task, failing to capture the structured, sequential boundaries of semantic drift. Here, we introduce an internal hidden state framework for fine-grained, span-level hallucination detection. By inspecting layer-wise activation patterns, we attempt to detect the exact hallucination onset and continuation tokens in an LLM generation. Our experiments show that this approach successfully isolates hallucination onsets, achieving substantial improvements in Precision-Recall AUC over random baselines despite extreme class imbalance. Ultimately, we propose a novel cross-model detection framework in which one model observes the internal representations elicited by another model's generation. We find that an external observer can match or exceed a generator's self-detection of its own hallucination onsets, including when the observer is the smaller model, suggesting that self-detection is not the ceiling for onset localisation.
Figures & tables
Figure 1: Token-Level Hallucination Span Detection Framework. Phase 1 identifies critical layers offline to supply the extraction targets for Phase 2, while Phase 2 processes this continuous feature sequence through the trained classifier to output structurally valid span tags
Figure 2: Critical layer analysis of the SmolLM2 (1.7B) architecture across its Attention and MLP blocks. The bars represent the cosine similarity between normal and hallucinated feature vectors at each layer depth. Layers highlighted in yellow denote the selected critical layers ( Ccritical ), where sharp geometric deviations reveal the onset of semantic drift.
Dataset (Model)
Safe (O)
B-HALL
I-HALL
Total
PsiloQA (SmolLM2)
7,797
265
2,923
10,985
PsiloQA (TinyLlama)
13,022
707
18,288
32,017
RAGTruth (Mistral)
11,241
57
616
11,914
Table 1: Token-level distribution of the BIO tags across the test sets of the three evaluated datasets. Each dataset was uniformly partitioned into standard 80/20 train/test splits prior to hidden state extraction
SmolLM2 (1.7B)
TinyLlama (1.1B)
Mistral (7B)
Architecture
P
R
F1
P
R
F1
P
R
F1
Onset Token ( B-HALL )
MLP Baseline
0.32
0.39
0.35
0.29
0.36
0.32
0.19
0.26
0.22
BiLSTM
0.29
0.40
0.33
0.28
0.38
0.32
0.25
0.25
0.25
BiLSTM-CRF
0.42
0.24
0.31
0.32
0.30
0.31
0.38
0.18
0.24
BiLSTM-Attn-CRF
0.42
0.30
0.35
0.41
0.28
0.34
0.41
0.21
0.28
Table 2: Intra-model span detection performance across three LLM architectures. Precision (P), Recall (R), and F1-scores are reported exclusively for the minority hallucination classes to highlight the detection of semantic drift. The Safe (O) majority class is excluded to prevent metric inflation. All experiments were conducted using the top-5 Attention and top-5 MLP critical layers.
Token Class
Model
AUROC
PR-AUC
Onset ( B-HALL )
MLP Baseline
0.900
0.330
Ogasa ( Ogasa and Arase, 2025 )
0.931
0.371
BERT-CRF
0.895
0.362
Random Baseline: 0.024
Continuation ( I-HALL )
MLP Baseline
0.716
0.460
BERT-CRF
0.735
0.446
Table 3: Threshold-independent onset diagnostics on the SmolLM2 test set, including the intra-model attention baseline of Ogasa and Arase (2025) . The Random Baseline equals the class prevalence. Single-seed.
Cross-Model Framework Evaluation
Architecture
QwenSmol
Gemma2Smol
Gemma4Smol
Gemma2Mist
SmolMist
P
R
F1
P
R
F1
P
R
F1
P
R
F1
P
R
F1
Onset Token ( B-HALL )
MLP Baseline
0.37
0.41
0.39
0.42
0.57
0.48
0.34
0.67
0.45
0.10
0.39
0.16
0.11
0.32
0.17
BiLSTM
0.39
0.42
0.41
0.45
0.49
0.47
0.39
0.54
0.45
0.29
0.46
0.35
0.17
0.47
0.25
BiLSTM-CRF
0.53
0.29
0.38
0.55
0.36
0.43
0.50
0.33
0.40
0.40
0.32
0.35
0.33
0.14
0.20
Table 4: Comprehensive cross-model hallucination span detection performance. The detection pairs QwenSmol utilizes Qwen (1.5B) to evaluate SmolLM2 (1.7B) generations, Gemma2Smol and Gemma4Smol utilize Gemma 2 (2B) and Gemma 4 (E2B) respectively to evaluate SmolLM2 (1.7B), Gemma2Mist utilizes Gemma 2 (2B) to evaluate Mistral (7B), and SmolMist utilizes the smaller SmolLM2 (1.7B) to evaluate Mistral (7B). Precision (P), Recall (R), and F1-scores are reported as before.
Architecture
Self
External
(Mistral)
(SmolLM2 → Mistral)
MLP
0.172
0.130
BiLSTM
0.205
0.193
BiLSTM-CRF
0.172
0.227
BiLSTM-Attn-CRF
0.221
0.242
BERT-CRF
0.043
0.079
Table 5: Onset ( B-HALL ) PR-AUC on RAGTruth: Mistral-7B self-detection vs. a smaller external observer (SmolLM2-1.7B), same test fold. Bold marks the higher value per row. Single-seed.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Component
Configuration
MLP
Hidden layers
din –1024–512–128–64–3
Activation
ReLU, Dropout 0.3
BiLSTM
Input dim
64 (projected)
LSTM hidden
32 / direction, 1 layer
Head
Linear 64→3
BiLSTM-CRF
Base
BiLSTM (above)
Appendix
Table 6: Layer configurations for each classifier, held fixed across datasets. din is the concatenated NAS input dimension, determined by the generator’s hidden size and critical-layer count. The BERT-CRF projector intermediate is widened to 4096 for TinyLlama.
Architecture
P
R
F1
PR-AUC
AUROC
MLP
0.360
0.558
0.438
0.380
0.854
BiLSTM
0.323
0.585
0.416
0.305
0.800
BiLSTM-CRF
0.492
0.340
0.402
0.252
0.784
BiLSTM-Attn-CRF
0.517
0.340
0.410
0.313
0.783
Appendix
Table 7: Onset ( B-HALL ) detection on SmolLM2 using a single fixed random layer set, not the critical layers used elsewhere in this paper. One draw, not an average over draws; single classifier seed. Corresponding critical-layer figures are given in the text below.
Hyperparameter
Value
Scope
Optimizer
Adam / AdamW
All architectures.
Learning rate
1×10−3
Classifier heads and projectors.
BERT encoder LR
2×10−5 – 5×10−5
After a 1-epoch warm-up.
B-HALL weight
20 – 300
Tuned per dataset by onset sparsity.
I-HALL weight
5 – 30
Tuned per dataset.
Training epochs
5 – 15
By architecture and token volume.
Appendix
Table 8: Representative self-evaluation training hyperparameters. Class weights and epochs vary by dataset to accommodate differing onset sparsity; exact per-model values are released with our code.
Cross-Model Hyperparameter
Value
Scope
Auxiliary B-HALL weight
100 – 150
RAGTruth-Mistral onset penalty.
Auxiliary multiplier ( λ )
1 – 3
Token penalty vs. sequence logic.
BERT unfreeze epoch
5
Delayed for projector alignment.
Post-warmup BERT LR
1×10−5
Stable end-to-end tuning.
Gradient accumulation
4 steps
Memory boundary.
Appendix
Table 9: Training hyperparameters specific to the cross-model evaluation framework.