DialectSentEval 2026: Arabic Dialect Sentiment Analysis and Swapping Shared Task
Authors: Saad Ezzini, Shadi Abudalfa, Maram Alharbi, Salmane Chafik, Hind Alatawi, Mo El-Haj, Ahmed Abdelali, Osamah Alnahari, +1 more
Organizations: King Fahd University of Petroleum and Minerals, KSA · Onaizah Colleges, KSA · University College of Applied Sciences, Palestine · Lancaster University, UK · UM6P, Morocco · VinUniversity, Vietnam · El Technology, Qatar · University of Luxembourg, Luxembourg
Sentiment analysis is a fundamental problem in Natural Language Processing (NLP). Standard sentiment classification for the Arabic language remains challenging due to the high volume of dialectal Arabic. To advance research in this area, this paper proposes the Shared Task on Sentiment Analysis and Swapping in Arabic Dialects (DialectSentEval), hosted with the Arabic Natural Language Processing Conference (ArabicNLP 2026). This shared task consists of two subtasks: Subtask 1 focuses on multi-class and multi-dialect sentiment analysis, requiring models to identify sentiment polarity across various Arabic dialects. Subtask 2 introduces a generative task for Arabic sentiment swap, challenging models to invert sentiment polarity while preserving core semantics. In this overview paper, we present the motivation, dataset creation, and summarize the main findings from participating models.
Figures & tables
Figure 1: Data creation process for both subtasks.
#
Team Name
CodaBench Username
Macro F1
1
ALEXIS
momen_ibrahim-870387
0.8667
2
Axiom
ajwad_hossain-867052
0.8210
3
PTUK-NLP
tasneemduridi-857223
0.8174
4
NU-SentEval
habiba_ayman-867381
0.8140
5
Gradient Descenders
normalman159-857755
0.8010
6
LahjaSensei
ahmadalmoustafa-856400
0.7977
Table 1: Official evaluation results for Subtask 1: Arabic Dialect Sentiment Analysis.
#
Team Name
CodaBench Username
Sentiment Style Accuracy
BLEU
chrF
1
ALEXIS
momen_ibrahim-870527
0.9581
56.91
72.18
2
AKARROUCH
akarrouch-871040
0.7851
49.82
70.12
3
LahjaSensei
ahmadalmoustafa-865064
0.7789
41.48
64.45
4
NAMAA
mona_-871054
0.7597
51.58
71.02
5
–
melmoussaoui-870973
0.7530
49.03
69.53
6
Sabaa
mohamedsabaa-855812
0.7447
49.07
69.33
Table 2: Official evaluation results for Subtask 2: Arabic Sentiment Swap.
With the rapid growth of Arabic NLP, several models, datasets and benchmarks have been reported. This paper asks whether approaches developed for majority languages like English can be adapted to Arabic tasks. We adapt an English aspect-based sentiment analysis framework to Arabic classification tasks and present the adaptation as BARRAC: Brainstorming Alignment and Replaced Representation learning for ArabiC tasks. BARRAC replaces consumer-review attribute pools with Arabic linguistic devices and markers for dialectal sentiment, sarcasm, and dialect identification, and replaces noisy self-training with two-stage training. Evaluated on five Arabic dialect datasets, BARRAC achieves a mean macro-F1 of 63.93%, outperforming the best few-label SOTA by 3%, and outperforming GPT-4o on four out of five tasks. Error analysis provides insights into remaining challenges. These results demonstrate that adapting task-specific approaches is a promising direction for Arabic NLP alongside adapting models, datasets and benchmarks.
Ali Almutairi, Gelareh Mohammadi, Imran Razzak +1
University of New South Wales, Australia · MBZUAI, UAE
NADI 2026 is the seventh edition of the Nuanced Arabic Dialect Identification (NADI) shared task series and the second dedicated to multidialectal Arabic speech processing. This edition comprises five tasks and eight subtasks spanning Automatic Speech Recognition (ASR), Spoken Dialect Identification (SDID), Text-to-Speech (TTS), Spoken Language Translation (SLT), and Spoken Language Understanding (SLU). NADI 2026 emphasizes realistic evaluation through low-bandwidth, mixed-dialect, code-switched, out-of-domain, and zero-shot settings, while introducing TTS, SLT, and SLU to the series for the first time. The shared task attracted 21 participating teams from at least 13 countries, with 48 test-phase submissions and 14 submitted system-description papers. Results show that out-of-domain generalization remains a major bottleneck and highlight the effectiveness of recent Arabic-specialized speech models, multimodal dialect identification approaches, and ensemble methods. Overall, NADI 2026 provides a broader and more challenging benchmark for robust Arabic dialect speech processing.
Peter Sullivan, Bashar Talafha, Ahmed Ashraf +11
The University of British Columbia · King Fahd University of Petroleum & Minerals · Avignon Université +4
Dialectal Arabic (DA) remains under-resourced compared to Modern Standard Arabic (MSA), particularly for machine reading comprehension (MRC) and question answering (QA). Existing Arabic QA benchmarks primarily focus on formal written MSA or multiple-choice QA, with limited coverage of naturally spoken dialects. Here, we aim to bridge this gap. We introduce EDRAC, the first large-scale benchmark for dialectal Arabic machine reading comprehension (MRC) and generative QA, covering five major dialects: Egyptian, Moroccan, Emirati, Syrian, and Saudi Arabic. EDRAC contains 499 passages derived from naturally occurring spoken interactions and 4,977 corresponding QA pairs generated through a human--LLM collaborative pipeline combining iterative generation, LLM-as-a-judge evaluation, and human verification. We benchmark Arabic-centric and multilingual LLMs on EDRAC using lexical and semantic metrics. Our results reveal substantial gaps between semantic answer quality and dialectal fidelity, highlighting the limitations of existing evaluation metrics for dialectal Arabic generation. EDRAC provides a realistic and challenging MRC benchmark for future research on dialectal Arabic NLP.
Noor Abo Mokh, Kirill Chirkunov, Teresa Lynn +15
Mohamed bin Zayed University of Artificial Intelligence · New York University Abu Dhabi · IBM Research AI +1