cs.CLSep 28, 2026

Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition

Authors: Omid Ghahroodi, Anas Madkoor, Dima Faris Al Saudi, Fagr Tahir, Malak Annan, Talha shahid javad allah rakha, Omar Al-Busaidi, Zineb El Kahla, +32 more

Organizations: QCRI, HBKU · Qatar University · UDST · Algo AI · UCL · University of Tripoli · AUC · KAUST · CMU-Q · KFUPM · Alfaisal University · Damascus University · Princeton University · ENSIAS, Mohammed V University · Sultan Qaboos University · DFKI · University of Waterloo · USTHB

Abstract

Arabic speech technology has largely focused on Modern Standard Arabic, leaving the living dialects spoken by hundreds of millions under-served. We introduce ALMIEYAR, a culturally grounded ASR benchmark covering 17 Arabic dialects across six families, built entirely from newly recorded speech unseen by existing models. Dialect-community coordinators selected culturally relevant images across 10 topics, and native speakers described them through five structured scenarios, yielding approximately 50 minutes per dialect (13.7 hours total). We benchmark 12 state-of-the-art ASR systems zero-shot, including GPT-4o-transcribe, Voxtral-Mini-4B, Fanar-STT-LF, Whisper, SeamlessM4T-v2, and wav2vec2-based models. GPT-4o-transcribe achieves the lowest overall WER at 35.0%, followed by Voxtral-Mini-4B, Fanar-STT-LF, and Whisper-Large-v3 at 41.1%, 45.9%, and 49.5%, respectively, indicating substantial remaining errors across Arabic dialect communities. Performance varies considerably across dialect groups, with no model performing uniformly best across all groups. WER alone also obscures dialectal ASR behaviour: wav2vec2-based models show large WER/CER gaps, where character-level agreement remains much higher than word-level accuracy, motivating joint WER/CER reporting. ALMIEYAR provides a unified benchmark for culturally grounded Arabic ASR evaluation, including the first published benchmark for Ahwazi Arabic.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CARDAMOM: A Micro-Dialectal Arabic Speech Dataset for ASR

    Sep 28, 2026Bashar Talafha, Samar M. Magdy, Aisha Alansari +36DialectsMultilingual Automatic Speech Recognition

  2. EDRAC: Benchmarking Arabic Dialect Reading Comprehension

    Sep 1, 2026Noor Abo Mokh, Kirill Chirkunov, Teresa Lynn +15Arabic Natural Language ProcessingArabic

  3. NADI 2026: The Second Multidialectal Arabic Speech Processing Shared Task

    Sep 22, 2026Peter Sullivan, Bashar Talafha, Ahmed Ashraf +11Arabic Natural Language ProcessingArabic