cs.CLFeb 3, 2026

SalamahBench: Dialect and Category Level Safety Evaluation of Arabic Language Models

Authors: Omar Abdelnasser, Fatemah Alharbi, Khaled Khasawneh, Ihsen Alouani, Mohammed E. Fouda

Organizations: Compumacy for Artificial Intelligence Solutions, Cairo, Egypt · Computer Science Department, College of Computer Science and Engineering, Taibah University, Yanbu 46522, KSA · Electrical and Computer Engineering Department, George Mason University, VA, USA · CSIT, Queen’s University Belfast, UK · Upper Bound Ltd, Belfast, UK

Abstract

While different stakeholders are trying to leverage Arabic Language Models (ALMs), safety alignment in ALMs remains largely underexplored, hindering their mainstream adoption. Existing safety benchmarks are predominantly English-centric and evaluate Arabic only in its standardized form, obscuring fine-grained safety vulnerabilities in Arabic NLP systems. This paper introduces SalamahBench, a unified benchmark of 8{,}270 human-verified harmful prompts across ML Commons hazard categories, each rendered in Modern Standard Arabic (MSA) and five regional Arabic varieties, namely Egyptian, Syrian, Saudi, Lebanese, and Moroccan, for a total of 49{,}620 paired instances. To analyze the resulting data, we introduce two complementary metrics, namely Dialect Shift, which measures a model's aggregate change in safety under dialectal reformulation, and Category-Specific Dialect Deviation, which isolates harm categories whose change departs from that aggregate trend. Evaluating models such as Fanar 2, ALLaM 2, and Karnak 1 under multiple safeguard configurations, we find that cross-variety robustness is strongly model dependent, and that aggregate scores can conceal category-level divergence. Our findings highlight the necessity of evaluating Arabic model safety jointly across linguistic varieties and harm domains rather than relying on aggregate scores or MSA alone.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification

    Aug 2, 2026Wajdi Zaghouani, Md. Rafiul Biswas, Kholoud Khalil Aldous +1Arabic Natural Language ProcessingArabic

  2. Evaluation of Small Language Models for Arabic Language Processing

    Jun 19, 2026Jumana Alsubhi, Ahmed Alhusayni, Abdulrahman Gharawi +5Arabic Natural Language ProcessingSmall Large Language Models