cs.CLSep 28, 2026

CARDAMOM: A Micro-Dialectal Arabic Speech Dataset for ASR

Authors: Bashar Talafha, Samar M. Magdy, Aisha Alansari, Alaa Alkhawaldeh, Abdurrahman Juma, Sharaf Makahleh, Nour Gamal, Omar Attia, +31 more

Organizations: The University of British Columbia · King Fahd University of Petroleum and Minerals · Al al-Bayt University · Birzeit University · Jordan University of Science and Technology · Badr University in Cairo · Taibah University · Northern Border University · Imam Abdulrahman Bin Faisal University · El Sewedy University of Technology · Imperial College London · Institut Supérieur du Numérique · Lebanese University · American University of Beirut · Arab Center for Research and Policy Studies · Hamad Bin Khalifa University

Abstract

We present Cardamom, a micro-dialectal Arabic speech dataset designed to support fine-grained evaluation and adaptation of automatic speech recognition (ASR) systems. Community-curated by native speakers familiar with the represented varieties, Cardamom contains approximately 40 hours of transcribed YouTube speech spanning 21 micro-dialects across Egypt, Jordan, Lebanon, Mauritania, Palestine, and Saudi Arabia. Each segment is annotated with one or more operational micro-dialect labels, code-switching information, and utterance-level perceived gender, enabling analysis of sub-country variation that is obscured by conventional country-level labels. We describe the collection and annotation process, motivate the micro-dialect inventory linguistically, and benchmark four multilingual ASR systems in zero-shot and adapted settings. The strongest zero-shot system obtains 43.47% aggregate WER, with particularly high error rates on Mauritanian and Lebanese varieties; adaptation on Cardamom reduces its WER to 35.21%. Audio-based identification experiments further show that the annotations provide a learnable prediction target, with a dedicated classifier reaching 85.57% accuracy on 21-way micro-dialect identification. Cardamom provides a resource for studying localized dialectal variation and developing Arabic speech systems with broader regional coverage.

Explore similar work

CardsList
  1. Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition

    Sep 28, 2026Omid Ghahroodi, Anas Madkoor, Dima Faris Al Saudi +37Arabic Natural Language ProcessingArabic

  2. NADI 2026: The Second Multidialectal Arabic Speech Processing Shared Task

    Sep 22, 2026Peter Sullivan, Bashar Talafha, Ahmed Ashraf +11Arabic Natural Language ProcessingArabic