cs.CLSep 29, 2026

Halluscoring 2026: The first shared task on llms hallucination detection and answer verification

Authors: Aisha Alansari, Abdessalam Bouchekif, Ahmed Hasanaath, Salah Eddine Bekhouche, Malak Alkhorasani, Mohammed-En-Nadhir Zighem, Saad Ezzini, Hichem Telli, +4 more

Organizations: King Fahd University of Petroleum and Minerals, Saudi Arabia · Hamad Bin Khalifa University, Qatar · University of the Basque Country, Spain · Imam Abdulrahman bin Faisal University, Saudi Arabia · University of Biskra, Algeria · King Saud University, Saudi Arabia · University of British Columbia, Canada · Universiti Malaysia Kelantan, Malaysia

Abstract

We present HalluScoring 2026, a shared task for evaluating hallucination detection and factual verification in Arabic question answering under challenging generalization settings. Its four subtasks are organized into two tasks. Task~1 evaluates binary hallucination detection, considering generalization to unseen questions (Subtask 1.1) and responses generated by unseen LLMs (Subtask 1.2). Task 2 extends the evaluation beyond detection by requiring the systems to additionally identify the correct factual answer from six related candidates, covering Islamic knowledge (Subtask 2.1) and general knowledge (Subtask 2.2). The shared task is based on two Arabic datasets: HalluScore and HalluTruthQA. A total of 13 teams participated in the shared task, ten of which submitted system description papers. The results of Task 1 demonstrate that hallucination detection remains challenging. On the Task 1 test sets, the top-ranked systems achieved AUC-ROC scores of 0.7717 for Subtask 1.1 (REGLAT) and 0.7670 for Subtask 1.2 (NAMAA). Under assisted evaluation, the highest combined detection and answer-selection scores for Subtasks 2.1 and 2.2 were 0.8824 and 0.8565, respectively.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering

    Jul 22, 2026Abdessalam Bouchekif, Mohammed-En-Nadhir Zighem, Salah Eddine Bekhouche +9Hallucination DetectionArabic

  2. HalluScore: Large Language Model Hallucination Question Answering Benchmark

    May 16, 2026Aisha Alansari, Hamzah LuqmanArabic Natural Language ProcessingCultural Awareness