CARDAMOM: A Micro-Dialectal Arabic Speech Dataset for ASR
Organizations: The University of British Columbia · King Fahd University of Petroleum and Minerals · Al al-Bayt University · Birzeit University · Jordan University of Science and Technology · Badr University in Cairo · Taibah University · Northern Border University · Imam Abdulrahman Bin Faisal University · El Sewedy University of Technology · Imperial College London · Institut Supérieur du Numérique · Lebanese University · American University of Beirut · Arab Center for Research and Policy Studies · Hamad Bin Khalifa University
Abstract
We present Cardamom, a micro-dialectal Arabic speech dataset designed to support fine-grained evaluation and adaptation of automatic speech recognition (ASR) systems. Community-curated by native speakers familiar with the represented varieties, Cardamom contains approximately 40 hours of transcribed YouTube speech spanning 21 micro-dialects across Egypt, Jordan, Lebanon, Mauritania, Palestine, and Saudi Arabia. Each segment is annotated with one or more operational micro-dialect labels, code-switching information, and utterance-level perceived gender, enabling analysis of sub-country variation that is obscured by conventional country-level labels. We describe the collection and annotation process, motivate the micro-dialect inventory linguistically, and benchmark four multilingual ASR systems in zero-shot and adapted settings. The strongest zero-shot system obtains 43.47% aggregate WER, with particularly high error rates on Mauritanian and Lebanese varieties; adaptation on Cardamom reduces its WER to 35.21%. Audio-based identification experiments further show that the annotations provide a learnable prediction target, with a dedicated classifier reaching 85.57% accuracy on 21-way micro-dialect identification. Cardamom provides a resource for studying localized dialectal variation and developing Arabic speech systems with broader regional coverage.