Organizations: Faculty of Computer Engineering University of Isfahan Isfahan, Iran · School of Electrical and Computer Engineering University of Tehran Tehran, Iran · School of Computer Science University of Windsor Windsor, Canada · Center for Brain Health University of Texas at Dallas TX, US
Hallucination detection is particularly important for medical language models, but repeated-sampling approaches are expensive and existing uncertainty-head resources do not directly transfer to a new backbone and language. We adapt the LLM Uncertainty Head (LUH) framework to Aya-Expanse-8B-based Persian medical models, using Gaokerena-V and Gaokerena-R as two previously developed backbones. We first examine response variability on a 168-question Iranian medical entrance examination and observe substantially lower five-run consistency for Gaokerena-V than for Aya-Expanse-8B, whereas Gaokerena-R is comparable to Aya-Expanse-8B. We then construct two paired claim-level hallucination datasets directly in Persian, containing 1,600 responses for each backbone, and train lightweight claim-level heads on frozen backbone attention maps and token probabilities. On held-out test splits, the heads obtain PR-AUCs of 0.4820 and 0.4652, corresponding to 2.30 and 2.66 times their respective random baselines, and ROC-AUCs of 0.7852 and 0.7810. The heads require neither retrieval nor repeated sampling at inference time. These results provide an initial study of single-pass claim-level uncertainty estimation for Persian medical language models; the test splits are small and the labels are automatically generated.
Figures & tables
Gao
Gao
Aya-
Med-
kerena-V
kerena-R
Expanse
Gemma
Accuracy
29.76
38.69
35.71
37.50
Questions with a
14
37
36
168
single option taken
Questions with all
10
4
6
0
options taken
TABLE I: Five-run consistency on the September 2023 IBMSEE. Accuracy is the ensemble-answer accuracy (%) with not-agreed questions counted as incorrect; entropy is the mean per-question Shannon entropy in bits over valid runs (see text).
LUH-V
LUH-R
Annotated responses
1,600
1,600
Extracted claims
60,434
54,209
Supported claims
48,991
43,305
Hallucinated claims
10,557
9,528
Undecided claims
886
1,376
Hallucination rate
17.73%
18.03%
TABLE II: Statistics of the two native Persian uncertainty datasets.
Training setting
Value
Maximum epochs
10
Learning rate
1×10−4
Warmup ratio
0.05
Weight decay
0.1
Per-device batch size
8
Gradient accumulation
1
TABLE III: Training configuration used for both uncertainty heads.
Metric
Gaokerena-V
Gaokerena-R
Accuracy
76.24
62.96
Precision
44.85
29.52
Recall
58.17
80.84
F1
50.65
43.24
ROC-AUC
78.52
78.10
PR-AUC
48.20
46.52
TABLE IV: Claim-level uncertainty-head performance on the held-out test splits (%).