Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition
Authors: Omid Ghahroodi, Anas Madkoor, Dima Faris Al Saudi, Fagr Tahir, Malak Annan, Talha shahid javad allah rakha, Omar Al-Busaidi, Zineb El Kahla, +32 more
Organizations: QCRI, HBKU · Qatar University · UDST · Algo AI · UCL · University of Tripoli · AUC · KAUST · CMU-Q · KFUPM · Alfaisal University · Damascus University · Princeton University · ENSIAS, Mohammed V University · Sultan Qaboos University · DFKI · University of Waterloo · USTHB
Arabic speech technology has largely focused on Modern Standard Arabic, leaving the living dialects spoken by hundreds of millions under-served. We introduce ALMIEYAR, a culturally grounded ASR benchmark covering 17 Arabic dialects across six families, built entirely from newly recorded speech unseen by existing models. Dialect-community coordinators selected culturally relevant images across 10 topics, and native speakers described them through five structured scenarios, yielding approximately 50 minutes per dialect (13.7 hours total). We benchmark 12 state-of-the-art ASR systems zero-shot, including GPT-4o-transcribe, Voxtral-Mini-4B, Fanar-STT-LF, Whisper, SeamlessM4T-v2, and wav2vec2-based models. GPT-4o-transcribe achieves the lowest overall WER at 35.0%, followed by Voxtral-Mini-4B, Fanar-STT-LF, and Whisper-Large-v3 at 41.1%, 45.9%, and 49.5%, respectively, indicating substantial remaining errors across Arabic dialect communities. Performance varies considerably across dialect groups, with no model performing uniformly best across all groups. WER alone also obscures dialectal ASR behaviour: wav2vec2-based models show large WER/CER gaps, where character-level agreement remains much higher than word-level accuracy, motivating joint WER/CER reporting. ALMIEYAR provides a unified benchmark for culturally grounded Arabic ASR evaluation, including the first published benchmark for Ahwazi Arabic.
Figures & tables
Benchmark
Dialects
Hours
MGB-2 [ 4 ]
MSA-dom.
1,200
MGB-3 [ 6 ]
1 (EGY)
16
MGB-5 [ 5 ]
1 (MOR)
13
Common Voice [ 8 ]
Arabic varieties
varies
FLEURS [ 10 ]
Arabic varieties
varies
SADA [ 3 ]
4
668
Table 1: Comparison of Arabic ASR benchmarks. “MSA-dom.”: benchmark data are primarily MSA despite containing some dialectal speech.
Figure 1 : Almieyar benchmark construction pipeline. Dialect speakers describe culturally selected images across five structured scenarios using a Telegram-based collection interface. Automatic transcription is used only to assist speaker correction; final references are verified through speaker correction and coordinator review.
WER (%)
CER (%)
Model
Gulf
Lev.
Mag.
Iraqi
Egy.
Sud.
Ovr.
Gulf
Lev.
Mag.
Iraqi
Egy.
Sud.
Ovr.
GPT-4o-transcribe
36.4
33.6
34.9
28.3
49.7
30.6
34.9
19.3
14.5
15.7
9.4
33.9
10.2
16.5
Voxtral-Mini-4B
35.7
47.1
38.8
44.4
59.9
35.8
41.1
17.9
26.3
16.5
22.1
40.2
15.0
20.6
Fanar-STT-LF
42.2
45.3
50.4
46.6
48.4
41.8
45.9
20.2
20.1
23.0
20.1
27.4
16.1
21.0
Whisper-Large-v3
46.8
53.5
50.0
47.5
56.8
44.2
49.5
28.6
32.4
28.2
25.5
36.1
23.4
29.0
SeamlessM4T-v2
52.2
69.7
55.6
50.2
43.9
56.1
56.6
34.7
52.0
37.1
30.9
26.2
37.6
38.5
Table 2: Family-level WER (%) and CER (%) per model on Almieyar . Cells are colour-coded green (low) to red (high). Best per column in bold . Gulf : Bahraini, Omani, Qatari, Saudi, Yemeni; Lev. : Levantine (Jordanian, Lebanese, Palestinian, Syrian); Mag. : Maghrebi (Algerian, Libyan, Moroccan, Tunisian); Iraqi : Iraqi & Ahwazi; Egy. : Egyptian; Sud. : Sudanese; Ovr. : Overall average.
Model
Gulf
Lev.
Mag.
Irq.
Egy.
Sud.
Ovr.
Voxtral-Mini-4B
33.4
41.4
37.5
40.8
49.6
33.4
37.8
Fanar-STT-LF
41.8
44.4
49.9
46.4
45.3
41.3
45.3
Whisper-Large-v3
43.4
49.7
46.6
44.3
48.9
40.8
45.9
SeamlessM4T-v2
51.2
69.4
55.4
49.7
40.5
55.1
55.9
Whisper-Medium
52.1
55.3
54.3
53.8
59.0
52.0
53.9
Whisper-Small
61.5
71.9
67.4
66.1
72.2
58.1
66.2
Table 3 : MER (%) per model and dialect family on Almieyar . Cells are colour-coded green (low) to red (high). Best per column in bold . Family abbreviations match Table 2 .
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
WER (%)
CER (%)
Dialect
Vox
Fan
W-L
Sea
W-M
W-S
Moo
W53
XLS
XLv2
Sin
Vox
Fan
W-L
Sea
W-M
W-S
Moo
W53
XLS
XLv2
Sin
Gulf
Bahraini
43
41
61
54
70
76
73
78
80
82
85
23
17
44
30
48
58
50
27
30
30
40
Omani
30
41
51
76
64
67
67
78
81
83
89
11
22
28
65
38
44
42
29
34
32
50
Qatari
35
42
41
38
53
58
77
76
79
82
83
17
18
23
17
29
31
55
27
29
30
37
Saudi
25
38
36
42
45
59
52
73
77
82
83
11
19
20
24
26
33
34
28
30
32
37
Appendix
Table 4 : Per-dialect WER (%) and CER (%) on Almieyar for the four multi-dialect families, to the nearest whole percent. Model abbreviations, in the column order of Table 2 : Vox = Voxtral-Mini-4B, Fan = Fanar-STT-LF, W-L = Whisper-Large-v3, Sea = SeamlessM4T-v2, W-M = Whisper-Medium, W-S = Whisper-Small, Moo = Moonshine-AR, W53 = Wav2Vec2-53-AR, XLS = Wav2Vec2-XLSR-AR, XLv2 = Wav2Vec2-XLSR-v2, Sin = Sinai-STT.
We present Cardamom, a micro-dialectal Arabic speech dataset designed to support fine-grained evaluation and adaptation of automatic speech recognition (ASR) systems. Community-curated by native speakers familiar with the represented varieties, Cardamom contains approximately 40 hours of transcribed YouTube speech spanning 21 micro-dialects across Egypt, Jordan, Lebanon, Mauritania, Palestine, and Saudi Arabia. Each segment is annotated with one or more operational micro-dialect labels, code-switching information, and utterance-level perceived gender, enabling analysis of sub-country variation that is obscured by conventional country-level labels. We describe the collection and annotation process, motivate the micro-dialect inventory linguistically, and benchmark four multilingual ASR systems in zero-shot and adapted settings. The strongest zero-shot system obtains 43.47% aggregate WER, with particularly high error rates on Mauritanian and Lebanese varieties; adaptation on Cardamom reduces its WER to 35.21%. Audio-based identification experiments further show that the annotations provide a learnable prediction target, with a dedicated classifier reaching 85.57% accuracy on 21-way micro-dialect identification. Cardamom provides a resource for studying localized dialectal variation and developing Arabic speech systems with broader regional coverage.
Bashar Talafha, Samar M. Magdy, Aisha Alansari +36
The University of British Columbia · King Fahd University of Petroleum and Minerals · Al al-Bayt University +13
Dialectal Arabic (DA) remains under-resourced compared to Modern Standard Arabic (MSA), particularly for machine reading comprehension (MRC) and question answering (QA). Existing Arabic QA benchmarks primarily focus on formal written MSA or multiple-choice QA, with limited coverage of naturally spoken dialects. Here, we aim to bridge this gap. We introduce EDRAC, the first large-scale benchmark for dialectal Arabic machine reading comprehension (MRC) and generative QA, covering five major dialects: Egyptian, Moroccan, Emirati, Syrian, and Saudi Arabic. EDRAC contains 499 passages derived from naturally occurring spoken interactions and 4,977 corresponding QA pairs generated through a human--LLM collaborative pipeline combining iterative generation, LLM-as-a-judge evaluation, and human verification. We benchmark Arabic-centric and multilingual LLMs on EDRAC using lexical and semantic metrics. Our results reveal substantial gaps between semantic answer quality and dialectal fidelity, highlighting the limitations of existing evaluation metrics for dialectal Arabic generation. EDRAC provides a realistic and challenging MRC benchmark for future research on dialectal Arabic NLP.
Noor Abo Mokh, Kirill Chirkunov, Teresa Lynn +15
Mohamed bin Zayed University of Artificial Intelligence · New York University Abu Dhabi · IBM Research AI +1
NADI 2026 is the seventh edition of the Nuanced Arabic Dialect Identification (NADI) shared task series and the second dedicated to multidialectal Arabic speech processing. This edition comprises five tasks and eight subtasks spanning Automatic Speech Recognition (ASR), Spoken Dialect Identification (SDID), Text-to-Speech (TTS), Spoken Language Translation (SLT), and Spoken Language Understanding (SLU). NADI 2026 emphasizes realistic evaluation through low-bandwidth, mixed-dialect, code-switched, out-of-domain, and zero-shot settings, while introducing TTS, SLT, and SLU to the series for the first time. The shared task attracted 21 participating teams from at least 13 countries, with 48 test-phase submissions and 14 submitted system-description papers. Results show that out-of-domain generalization remains a major bottleneck and highlight the effectiveness of recent Arabic-specialized speech models, multimodal dialect identification approaches, and ensemble methods. Overall, NADI 2026 provides a broader and more challenging benchmark for robust Arabic dialect speech processing.
Peter Sullivan, Bashar Talafha, Ahmed Ashraf +11
The University of British Columbia · King Fahd University of Petroleum & Minerals · Avignon Université +4