Halluscoring 2026: The first shared task on llms hallucination detection and answer verification
Organizations: King Fahd University of Petroleum and Minerals, Saudi Arabia · Hamad Bin Khalifa University, Qatar · University of the Basque Country, Spain · Imam Abdulrahman bin Faisal University, Saudi Arabia · University of Biskra, Algeria · King Saud University, Saudi Arabia · University of British Columbia, Canada · Universiti Malaysia Kelantan, Malaysia
Abstract
We present HalluScoring 2026, a shared task for evaluating hallucination detection and factual verification in Arabic question answering under challenging generalization settings. Its four subtasks are organized into two tasks. Task~1 evaluates binary hallucination detection, considering generalization to unseen questions (Subtask 1.1) and responses generated by unseen LLMs (Subtask 1.2). Task 2 extends the evaluation beyond detection by requiring the systems to additionally identify the correct factual answer from six related candidates, covering Islamic knowledge (Subtask 2.1) and general knowledge (Subtask 2.2). The shared task is based on two Arabic datasets: HalluScore and HalluTruthQA. A total of 13 teams participated in the shared task, ten of which submitted system description papers. The results of Task 1 demonstrate that hallucination detection remains challenging. On the Task 1 test sets, the top-ranked systems achieved AUC-ROC scores of 0.7717 for Subtask 1.1 (REGLAT) and 0.7670 for Subtask 1.2 (NAMAA). Under assisted evaluation, the highest combined detection and answer-selection scores for Subtasks 2.1 and 2.2 were 0.8824 and 0.8565, respectively.
Figures & tables
| # LLMs | # Questions | # Instances | ||||||||
| Subtask | Dataset | Train | Dev. | Test | Train | Dev. | Test | Train | Dev. | Test |
| Subtask 1.1 Across Questions | HalluScore HalluTruthQA | 5 | 5 | 5 | 2,221 | 900 | 206 | 4,705 | 1,300 | 1,030 |
| Subtask 1.2 Across Models | HalluScore HalluTruthQA | 5 | 2 | 2 | 2,221 | 100 | 206 | 4,705 | 200 | 412 |
| Subtask 2.1 Islamic Knowledge | HalluTruthQA | 1 | 1 | 1 | 400 | 400 | 200 | 400 | 400 | 200 |
| Subtask 2.2 General Culture | HalluTruthQA | 1 | 1 | 1 | 1,600 | 800 | 1,600 | 1,600 | 800 | 1,600 |
| Team | Subtask | Description |
|---|---|---|
| Anhnamxtanh Nguyen and Takahashi (2026) | All | QLoRA-adapted ALLaM-7B with counterfactual cross-check recognition (C3R), using complementary TRUTH and MATCH roles, counterfactual training, option permutations, and constrained likelihood scoring. |
| AyahVerse Rashid (2026) | 1.1, 1.2, 2.2 | CAMeLBERT-based hallucination detection for Tasks 1.1–1.2 and a detection–verification pipeline combining CAMeLBERT with QLoRA-tuned AceGPT-v2-8B-Chat for Tasks 2.1–2.2. |
| DzairVerse Chenini and Belhadef (2026) | 2.1 | QLoRA-tuned Fanar-1-9B-Instruct with a leak-free two-stage pipeline: blind hallucination detection followed by candidate-based answer selection using separate prompts. |
| Hashtag AI Chaudhuri et al. (2026) | 2.1, 2.2 | QLoRA-adapted Gemma 4 12B with offline knowledge-graph teacher distillation and contrastive training examples, followed by hallucination detection and candidate answer selection. |
| HelsinkiTeam Bounab et al. (2026) | 2.2 | LoRA-adapted ALLaM-7B-Instruct trained with approximately augmented domain-specific and synthetic examples for joint hallucination detection and answer selection. |
| ShadowDz Kraimia and Eutamene (2026) | 2.1 | LoRA-tuned Fanar-1-9B-Instruct using blind hallucination detection followed by a separate candidate-selection stage conditioned on the predicted hallucination label. |
| Development | Test | ||||
|---|---|---|---|---|---|
| Rank | Team | AUC-ROC | F1-Macro | AUC-ROC | F1-Macro |
| 1 | REGLAT | 94.13 | 87.04 | 77.17 | 65.76 |
| 2 | Baseline_ARBERT | 92.05 | 85.07 | 76.39 | 62.54 |
| 3 | AyahVerse | 92.17 | 86.01 | 76.30 | 64.01 |
| 4 | NAMAA | 92.63 | 85.95 | 75.96 | 64.44 |
| 5 | Baseline_MARBERT | 91.44 | 83.75 | 74.66 | 62.48 |
| Development | Test | ||||
|---|---|---|---|---|---|
| Rank | Team | AUC-ROC | F1-Macro | AUC-ROC | F1-Macro |
| 1 | NAMAA | 75.80 | 66.11 | 76.70 | 66.34 |
| 2 | REGLAT | 79.21 | 69.34 | 74.50 | 65.07 |
| 3 | Baseline_MARBERT | 79.64 | 69.34 | 73.92 | 65.24 |
| 4 | AyahVerse | 74.92 | 63.24 | 72.99 | 61.35 |
| 5 | Baseline_ARBERT | 79.27 | 68.53 | 72.08 | 62.81 |
| Rank | Team | Assisted | Blind |
|---|---|---|---|
| Task 2.1: Islamic Knowledge | |||
| 1 | Anhnamxtanh | 88.2 | 85.9 |
| 2 | Hashtag AI | 81.2 | 82.6 |
| 3 | DzairVerse | 78.8 | 75.8 |
| 4 | IslamicAI | 73.1 | 65.4 |
| 5 | HelsinkiTeam | 72.0 | 66.6 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Subtask | Model Family | Approach | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Team | 1.1 | 1.2 | Encoder | Decoder | Ensemble | Reference | Distill. | Ext. Feat. | PEFT |
| AyahVerse | ✓ | ✓ | ✓ | ||||||
| NAMAA | ✓ | ✓ | ✓ | ||||||
| RaghadAlrasheed | ✓ | ✓ | ✓ | ✓ | ✓ | ||||
| REGLAT | ✓ | ✓ | ✓ | ✓ | |||||
| Scalar_NITK | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | |
| Subtask | Model Family | Approach | |||||||||
| Team | 2.1 | 2.2 | Encoder | Decoder | PEFT | Two-Stage | Cross-Check | Distill. | Data Aug. | Retrieval | Priv. Info. |
| Anhnamxtanh | ✓ | ✓ | ✓ | ✓ | ✓ | ||||||
| AyahVerse | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | |||||
| DzairVerse | ✓ | ✓ | ✓ | ✓ | |||||||
| Hashtag AI | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | |||
| HelsinkiTeam | ✓ | ✓ | ✓ | ✓ | |||||||