Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition
Authors: Omid Ghahroodi, Anas Madkoor, Dima Faris Al Saudi, Fagr Tahir, Malak Annan, Talha shahid javad allah rakha, Omar Al-Busaidi, Zineb El Kahla, +32 more
Organizations: QCRI, HBKU · Qatar University · UDST · Algo AI · UCL · University of Tripoli · AUC · KAUST · CMU-Q · KFUPM · Alfaisal University · Damascus University · Princeton University · ENSIAS, Mohammed V University · Sultan Qaboos University · DFKI · University of Waterloo · USTHB
Arabic speech technology has largely focused on Modern Standard Arabic, leaving the living dialects spoken by hundreds of millions under-served. We introduce ALMIEYAR, a culturally grounded ASR benchmark covering 17 Arabic dialects across six families, built entirely from newly recorded speech unseen by existing models. Dialect-community coordinators selected culturally relevant images across 10 topics, and native speakers described them through five structured scenarios, yielding approximately 50 minutes per dialect (13.7 hours total). We benchmark 12 state-of-the-art ASR systems zero-shot, including GPT-4o-transcribe, Voxtral-Mini-4B, Fanar-STT-LF, Whisper, SeamlessM4T-v2, and wav2vec2-based models. GPT-4o-transcribe achieves the lowest overall WER at 35.0%, followed by Voxtral-Mini-4B, Fanar-STT-LF, and Whisper-Large-v3 at 41.1%, 45.9%, and 49.5%, respectively, indicating substantial remaining errors across Arabic dialect communities. Performance varies considerably across dialect groups, with no model performing uniformly best across all groups. WER alone also obscures dialectal ASR behaviour: wav2vec2-based models show large WER/CER gaps, where character-level agreement remains much higher than word-level accuracy, motivating joint WER/CER reporting. ALMIEYAR provides a unified benchmark for culturally grounded Arabic ASR evaluation, including the first published benchmark for Ahwazi Arabic.
Figures & tables
Benchmark
Dialects
Hours
MGB-2 [ 4 ]
MSA-dom.
1,200
MGB-3 [ 6 ]
1 (EGY)
16
MGB-5 [ 5 ]
1 (MOR)
13
Common Voice [ 8 ]
Arabic varieties
varies
FLEURS [ 10 ]
Arabic varieties
varies
SADA [ 3 ]
4
668
Table 1: Comparison of Arabic ASR benchmarks. “MSA-dom.”: benchmark data are primarily MSA despite containing some dialectal speech.
Figure 1 : Almieyar benchmark construction pipeline. Dialect speakers describe culturally selected images across five structured scenarios using a Telegram-based collection interface. Automatic transcription is used only to assist speaker correction; final references are verified through speaker correction and coordinator review.
WER (%)
CER (%)
Model
Gulf
Lev.
Mag.
Iraqi
Egy.
Sud.
Ovr.
Gulf
Lev.
Mag.
Iraqi
Egy.
Sud.
Ovr.
GPT-4o-transcribe
36.4
33.6
34.9
28.3
49.7
30.6
34.9
19.3
14.5
15.7
9.4
33.9
10.2
16.5
Voxtral-Mini-4B
35.7
47.1
38.8
44.4
59.9
35.8
41.1
17.9
26.3
16.5
22.1
40.2
15.0
20.6
Fanar-STT-LF
42.2
45.3
50.4
46.6
48.4
41.8
45.9
20.2
20.1
23.0
20.1
27.4
16.1
21.0
Whisper-Large-v3
46.8
53.5
50.0
47.5
56.8
44.2
49.5
28.6
32.4
28.2
25.5
36.1
23.4
29.0
SeamlessM4T-v2
52.2
69.7
55.6
50.2
43.9
56.1
56.6
34.7
52.0
37.1
30.9
26.2
37.6
38.5
Table 2: Family-level WER (%) and CER (%) per model on Almieyar . Cells are colour-coded green (low) to red (high). Best per column in bold . Gulf : Bahraini, Omani, Qatari, Saudi, Yemeni; Lev. : Levantine (Jordanian, Lebanese, Palestinian, Syrian); Mag. : Maghrebi (Algerian, Libyan, Moroccan, Tunisian); Iraqi : Iraqi & Ahwazi; Egy. : Egyptian; Sud. : Sudanese; Ovr. : Overall average.
Model
Gulf
Lev.
Mag.
Irq.
Egy.
Sud.
Ovr.
Voxtral-Mini-4B
33.4
41.4
37.5
40.8
49.6
33.4
37.8
Fanar-STT-LF
41.8
44.4
49.9
46.4
45.3
41.3
45.3
Whisper-Large-v3
43.4
49.7
46.6
44.3
48.9
40.8
45.9
SeamlessM4T-v2
51.2
69.4
55.4
49.7
40.5
55.1
55.9
Whisper-Medium
52.1
55.3
54.3
53.8
59.0
52.0
53.9
Whisper-Small
61.5
71.9
67.4
66.1
72.2
58.1
66.2
Table 3 : MER (%) per model and dialect family on Almieyar . Cells are colour-coded green (low) to red (high). Best per column in bold . Family abbreviations match Table 2 .
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
WER (%)
CER (%)
Dialect
Vox
Fan
W-L
Sea
W-M
W-S
Moo
W53
XLS
XLv2
Sin
Vox
Fan
W-L
Sea
W-M
W-S
Moo
W53
XLS
XLv2
Sin
Gulf
Bahraini
43
41
61
54
70
76
73
78
80
82
85
23
17
44
30
48
58
50
27
30
30
40
Omani
30
41
51
76
64
67
67
78
81
83
89
11
22
28
65
38
44
42
29
34
32
50
Qatari
35
42
41
38
53
58
77
76
79
82
83
17
18
23
17
29
31
55
27
29
30
37
Saudi
25
38
36
42
45
59
52
73
77
82
83
11
19
20
24
26
33
34
28
30
32
37
Appendix
Table 4 : Per-dialect WER (%) and CER (%) on Almieyar for the four multi-dialect families, to the nearest whole percent. Model abbreviations, in the column order of Table 2 : Vox = Voxtral-Mini-4B, Fan = Fanar-STT-LF, W-L = Whisper-Large-v3, Sea = SeamlessM4T-v2, W-M = Whisper-Medium, W-S = Whisper-Small, Moo = Moonshine-AR, W53 = Wav2Vec2-53-AR, XLS = Wav2Vec2-XLSR-AR, XLv2 = Wav2Vec2-XLSR-v2, Sin = Sinai-STT.