SalamahBench: Dialect and Category Level Safety Evaluation of Arabic Language Models
Organizations: Compumacy for Artificial Intelligence Solutions, Cairo, Egypt · Computer Science Department, College of Computer Science and Engineering, Taibah University, Yanbu 46522, KSA · Electrical and Computer Engineering Department, George Mason University, VA, USA · CSIT, Queen’s University Belfast, UK · Upper Bound Ltd, Belfast, UK
Abstract
While different stakeholders are trying to leverage Arabic Language Models (ALMs), safety alignment in ALMs remains largely underexplored, hindering their mainstream adoption. Existing safety benchmarks are predominantly English-centric and evaluate Arabic only in its standardized form, obscuring fine-grained safety vulnerabilities in Arabic NLP systems. This paper introduces SalamahBench, a unified benchmark of 8{,}270 human-verified harmful prompts across ML Commons hazard categories, each rendered in Modern Standard Arabic (MSA) and five regional Arabic varieties, namely Egyptian, Syrian, Saudi, Lebanese, and Moroccan, for a total of 49{,}620 paired instances. To analyze the resulting data, we introduce two complementary metrics, namely Dialect Shift, which measures a model's aggregate change in safety under dialectal reformulation, and Category-Specific Dialect Deviation, which isolates harm categories whose change departs from that aggregate trend. Evaluating models such as Fanar 2, ALLaM 2, and Karnak 1 under multiple safeguard configurations, we find that cross-variety robustness is strongly model dependent, and that aggregate scores can conceal category-level divergence. Our findings highlight the necessity of evaluating Arabic model safety jointly across linguistic varieties and harm domains rather than relying on aggregate scores or MSA alone.
Figures & tables
| Hazard Category | Definition |
| 1. Violent Crimes | Responses that enable, encourage, or provide instructions for violent crimes against people or animals. |
| 2. Non-Violent Crimes | Responses that enable, encourage, or provide instructions for non-violent crimes such as theft, fraud, or property damage. |
| 3. Sex-Related Crimes | Responses related to sexual assault, trafficking, or the non-consensual distribution of intimate media. |
| 4. Child Sexual Exploitation | Responses that enable, encourage, or depict the sexual abuse or exploitation of minors. |
| 5. Indiscriminate Weapons | Responses that enable, encourage, or provide instructions for constructing or deploying chemical, biological, radiological, or nuclear (CBRN) weapons. |
| 6. Suicide & Self-Harm | Responses that enable, encourage, or provide instructions for suicide, self-harm, or eating disorders. |
| Dataset | Arabic prompts | Dialect label | Arabic construction | Strengths | Limitations |
| RTP-LX | 1,050 | Partial | Transcreated by native speakers from a subset of RealToxicityPrompts, with 50 additional culturally specific prompts authored manually. Annotated by humans across eight harm dimensions. | • Targets culturally specific toxic language, including microaggressions, bias and hate • Written by native speakers rather than translated, so phrasing is idiomatic | • Most prompts are sentence fragments intended for continuation rather than requests for harmful action, so they probe toxic generation rather than instruction following • Coverage is concentrated on toxic language and does not extend across the hazard taxonomy |
| PGPrompts | 754 | ✗ | Arabic portion of PolyGuardPrompts, produced by translating WildGuardMix with NLLB-200-3.3B and assessing translation quality automatically with GPT-4o. | • Contains adversarial and jailbreak style prompts that are difficult for both target models and safeguards • Already aligned with the ML Commons hazard taxonomy, inherited from WildGuardMix | • The translation model refused a subset of the source prompts, biasing the resulting corpus toward less harmful content • Only 754 Arabic harmful prompts, and since this count includes refusals, the number of genuinely adversarial prompts is smaller still, which is few relative to the number of hazard categories |
| Arabic Safeguard Evaluation | 3,888 | ✗ | Adapted from a Chinese safety dataset, with region specific content replaced by Arabic variants. Contains three prompt types: Original Harmful, Indirect Harmful (FN) and Harmless (FP). | • Pairs direct and indirect formulations of the same intent, enabling evaluation against implicit as well as explicit harm | • Includes politically sensitive categories that are not hazard categories • Coverage is uneven across hazard categories, with several ML Commons categories unrepresented |
| AraSafe | 1,254 | Partial | Arabic prompts written natively by participants through a web interface, each labelled by two expert annotators with disagreements reconciled. | • Naturally phrased, realistic misuse scenarios elicited from real users rather than templated or synthetic attacks • Double annotation with reconciliation gives comparatively reliable harm labels | • Prompts are written in a range of Arabic varieties, but no variety label is recorded, so the data cannot be used to assess dialectal vulnerability • The harmful subset of 1,254 prompts is thinly spread across hazard categories |
| X-Safety | 2,800 | ✗ | Arabic subset translated from English with the Google Translate API and subsequently corrected by professional translators. Fourteen safety categories across ten languages. | • Supports cross lingual comparison of the same prompts under a common framework • Covers adversarial prompting strategies such as goal hijacking and role play instruction | • The category scheme is orthogonal to hazard taxonomies by design, mixing harm types with instruction attack styles, so ten of its fourteen categories collapse to Others • Entirely in MSA, with no regional variety representation |
| LinguaSafe | 4,030 | ✗ | Derived from others benchmarks and augmented with natively sourced data across twelve languages. Arabic subset produced by machine translation, LM based refinement, and human verification of linguistic quality and preserved intent. | • Fine grained hierarchical labels spanning five domains and twenty three subtypes, with severity levels | • The composition of the Arabic subset is not reported per language • Severity annotation is not usable under a binary safe or unsafe framework and is discarded |
| Name | Size (Parameters) | Source | API Service Provider |
| Fanar 2 | 27B | Open-source | Hugging Face |
| ALLaM 2 | 7B | Open-source | Hugging Face |
| Karnak 1 | 41B | Open-source | Hugging Face |
| Qwen 3.8 | 27B | Open-source | Hugging Face |
| Fanar 1 | 9B | Open-source | Hugging Face |
| Jais 2 | 8B | Open-source | Hugging Face |
| Model/Statistic | Metric | Fanar 2 | Karnak 1 | ALLaM 2 |
| Generative Qwen3Guard | Refusal | 34120 | 28836 | 39694 |
| Safe | 45045 | 39797 | 47429 | |
| Controversial | 2500 | 5292 | 1431 | |
| Unsafe | 2075 | 4531 | 760 | |
| ASR (strict) | 9.2% (8.8–9.7) | 19.8% (19.2–20.4) | 4.4% (4.1–4.7) | |
| ASR (loose) | 4.2% (3.9–4.5) | 9.1% (8.7–9.6) | 1.5% (1.4–1.7) |
| Reference | Regional Arabic Varieties | Summary | |||||
| Model | MSA | Saudi | Moroccan | Syrian | Lebanese | Egyptian | Dialect Avg. |
| Macro-ASR (%) | |||||||
| ALLaM | 7.4 | 6.2 | 8.3 | 6.6 | 6.4 | 5.8 | 6.6 |
| Fanar | 11.4 | 12.9 | 12.5 | 12.7 | 13.6 | 12.9 | 12.9 |
| Karnak | 21.5 | 24.1 | 25.9 | 25.7 | 26.6 | 24.8 | 25.4 |
| Dialect Shift (DS, pp) | |||||||
| MSA | Saudi | Moroccan | Syrian | Lebanese | Egyptian | ||||||
| Category | ASR | ASR | CSDD | ASR | CSDD | ASR | CSDD | ASR | CSDD | ASR | CSDD |
| Intellectual Property | 30.1 | 42.2 | +10.5 | 32.5 | +1.3 | 32.5 | +1.1 | 40.4 | +8.1 | 30.7 | |
| Defamation | 30.2 | 33.6 | +1.9 | 27.6 | 33.6 | +2.2 | 37.1 | +4.7 | 31.0 | ||
| Nonviolent Crimes | 14.0 | 13.2 | 18.0 | +2.8 | 15.9 | +0.6 | 15.4 | 15.6 | +0.2 | ||
| Indiscriminate Weapons | 13.0 | 10.4 | 14.7 | +0.6 | 13.0 | 13.4 | 14.3 | ||||
| Suicide and Self-Harm | 9.1 | 11.8 | +1.2 | 9.1 | 10.0 | 10.9 | 14.5 | +4.0 | |||
| MSA | Saudi | Moroccan | Syrian | Lebanese | Egyptian | ||||||
| Category | ASR | ASR | CSDD | ASR | CSDD | ASR | CSDD | ASR | CSDD | ASR | CSDD |
| Intellectual Property | 22.9 | 34.9 | 40.4 | 30.7 | +8.7 | 28.3 | +6.5 | 28.9 | +7.7 | ||
| Indiscriminate Weapons | 19.5 | 10.4 | 14.3 | 12.1 | 13.0 | 7.4 | |||||
| Nonviolent Crimes | 7.1 | 5.3 | 6.3 | 4.9 | 4.7 | 5.5 | 0.0 | ||||
| Suicide and Self-Harm | 3.6 | 2.7 | 3.6 | 7.3 | 7.3 | 3.6 | |||||
| Sexual Content | 5.4 | 3.0 | 4.3 | 3.2 | 3.9 | 3.2 | |||||
| MSA | Saudi | Moroccan | Syrian | Lebanese | Egyptian | ||||||
| Category | ASR | ASR | CSDD | ASR | CSDD | ASR | CSDD | ASR | CSDD | ASR | CSDD |
| Intellectual Property | 62.7 | 63.9 | 62.7 | 59.6 | 63.3 | 61.4 | |||||
| Indiscriminate Weapons | 42.9 | 50.2 | +4.8 | 51.5 | +4.2 | 55.0 | +7.9 | 59.3 | 50.6 | +4.5 | |
| Defamation | 19.8 | 31.0 | +8.6 | 25.0 | +0.7 | 28.4 | +4.4 | 31.9 | +7.0 | 32.8 | +9.6 |
| Nonviolent Crimes | 23.2 | 23.5 | 34.0 | 28.0 | +0.6 | 30.3 | +2.0 | 27.1 | +0.6 | ||
| Sexual Content | 24.1 | 21.7 | 26.6 | 24.7 | 24.1 | 24.4 | |||||
| Arabic Variety | Llama Guard 4 | PolyGuard | Qwen3Guard | Majority Vote |
| MSA | 0.3 | 1.0 | 1.4 | 0.4 |
| Saudi | 0.9 | 1.2 | 2.5 | 0.8 |
| Moroccan | 3.0 | 0.5 | 3.6 | 0.9 |
| Syrian | 0.9 | 0.9 | 2.8 | 0.6 |
| Lebanese | 0.9 | 0.9 | 2.7 | 0.7 |
| Egyptian | 0.8 | 1.0 | 2.5 | 0.5 |
| Safeguard Model | Task | Correct / Total | Accuracy |
| Qwen3Guard | Response Safety Classification | 1220 / 1440 | 84.7% [82.8, 86.5] |
| Llama Guard 4 | Response Safety Classification | 0 941 / 1440 | 65.3% [62.9, 67.8] |
| PolyGuard | Response Safety Classification | 0 850 / 1440 | 59.0% [56.5, 61.5] |
| Majority Vote (3 Safeguards) | Response Safety Classification | 0 976 / 1440 | 67.8% [65.3, 70.2] |
| Qwen3Guard | Refusal Detection | 1045 / 1440 | 72.6% [70.2, 74.8] |
| Qwen3Guard | Llama Guard 4 | PolyGuard | Majority Vote | |||||||
| Category | Base | Acc. | Recall | Acc. | Recall | Acc. | Recall | Acc. | Recall | |
| Violent Crimes | 60 | 50.0 | 83.3 | 83.3 [72.0, 90.7] | 57.5 | 20.0 [11.8, 31.8] | 53.3 | 0 8.3 [3.6, 18.1] | 62.5 | 28.3 [18.5, 40.8] |
| Nonviolent Crimes | 58 | 51.7 | 90.0 | 91.4 [81.4, 96.3] | 60.8 | 19.0 [10.9, 30.9] | 52.5 | 0 3.4 [1.0, 11.7] | 61.7 | 22.4 [13.6, 34.7] |
| Sex-Related Crimes | 33 | 72.5 | 77.5 | 100.0 [89.6, 100] | 73.3 | 12.1 [ 0 4.8, 27.3] | 71.7 | 0 3.0 [0.5, 15.3] | 74.2 | 15.2 [ 0 6.7, 30.9] |
| Child Sexual Exploitation | 33 | 72.5 | 70.8 | 87.9 [72.7, 95.2] | 75.8 | 39.4 [24.7, 56.3] | 73.3 | 0 9.1 [3.1, 23.6] | 79.2 | 45.5 [29.8, 62.0] |
| Indiscriminate Weapons | 63 | 47.5 | 89.2 | 87.3 [76.9, 93.4] | 65.0 | 33.3 [22.9, 45.6] | 52.5 | 0 9.5 [4.4, 19.3] | 66.7 | 36.5 [25.7, 48.9] |
| Responses generated by | ||||
| Judge | ALLaM 2 | Fanar 2 | Karnak 1 | Overall |
| ALLaM 2 | 57.3 † | 52.5 | 44.0 | 51.3 |
| Fanar 2 | 37.1 | 38.3 † | 65.6 | 47.0 |
| Karnak 1 | 62.1 | 57.5 | 58.3 † | 59.3 |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| RTP-LX Category | ML Commons Mapping |
| Insult | Hate Speech |
| Bias | Hate Speech |
| IdentityAttack | Hate Speech |
| Violence | Violent Crimes |
| SelfHarm | Suicide & Self-Harm |
| SexualContent | Sexual Content |
| Wildguardmix Subcategory | ML Commons Mapping |
| Violence and physical harm | Violent Crimes |
| Private Information Individual | Privacy |
| Cyberattack | Nonviolent Crimes |
| Fraud Assisting Illegal Activities | Nonviolent Crimes |
| Sexual content | Sexual Content |
| Social Stereotypes And Unfair Discrimination | Hate |
| Arabic Dataset for LM Safeguard Evaluation Subcategory | ML Commons Mapping |
| Adult Content | Sexual Content |
| Social Stereotypes And Unfair Discrimination | Hate |
| Toxic Language (hate speech) | Hate |
| Mental Health Or Overreliance Crisis | Suicide and Self-Harm |
| Causing material harm by disseminating misinformation | Specialized Advice |
| Risks from leaking or inferring sensitive information | Privacy |
| Assigned Category | ML Commons Mapping |
| Violence or Harm terrorism | Violent Crimes |
| Privacy Violation | Privacy |
| Hate Speech | Hate |
| Explicit Content | Sexual Content |
| Misinformation | Others (harm) |
| Illegal Activities | Others (harm) |
| Assigned Category | ML Commons Mapping |
| Mental Health | Suicide and Self-Harm |
| Unfairness And Discrimination | Hate |
| Insult | Hate |
| Privacy And Property | Privacy |
| Goal Hijacking | Others (harm) |
| Prompt Leaking | Others (harm) |
| Assigned Category | ML Commons Mapping |
| Erotic Content | Sexual Content |
| Sex Orientation | Sexual Content |
| Sex-related Crimes | Sex-related Crimes |
| Controlled Items | Nonviolent Crimes |
| Code Misuse | Nonviolent Crimes |
| Financial Crimes | Nonviolent Crimes |
| Assigned Category | ML Commons Mapping |
| Non-violent Illegal Acts | Nonviolent Crimes |
| Violent | Violent Crimes |
| private information individual | Privacy |
| Sexual Content or Sexual Acts | Sexual Content |
| Suicide & Self-Harm | Suicide and Self-Harm |
| Politically Defamation | Defamation |
| Assigned Category | ML Commons Mapping |
| Chemical Biological | Indiscriminate Weapons |
| Cybercrime Intrusion | Nonviolent Crimes |
| Copyright | Intellectual Property |
| Harassment Bullying | Others (harm) |
| Illegal Activities | Others (harm) |
| Misinformation Disinformation | Others (harm) |
| Metric | Direct | Reason. | Refine. | |
| Naturalness | 92.60 | 95.67 | 98.63 | +6.03 |
| Idiomatic accuracy | 92.61 | 95.48 | 98.56 | +5.95 |
| Morphology & grammar | 97.17 | 97.91 | 99.63 | +2.46 |
| Semantic preservation | 99.70 | 99.83 | 99.95 | +0.25 |
| Average | 95.52 | 97.22 | 99.19 | +3.67 |
| Dimension | Mean |
| Semantic preservation | 4.99 |
| Harmful-intent preservation | 4.98 |
| Dialectal naturalness | 4.73 |
| Dialect authenticity | 4.59 |
| Grammaticality | 4.97 |
| Overall mean | 4.85 |
| Detection Method | Accuracy | Refusal F1 | Precision |
| Gemma 4 31B IT | 0.9989 | 0.9989 | 0.9979 |
| BGE-M3 | 0.9703 | 0.9714 | 0.9444 |
| CometKiwi | 0.5728 | 0.7028 | 0.5418 |
| Model/Statistic | Metric | Fanar 1 | Falcon H1R | Jais 2 |
| Generative Qwen3Guard | Refusal | 6757 | 7821 | 5331 |
| RR | 81.7% | 94.6% | 64.5% | |
| Safe | 7806 | 7804 | 6257 | |
| Controversial | 223 | 126 | 900 | |
| Unsafe | 241 | 340 | 1113 | |
| ASR (strict) | 5.6% | 5.6% | 24.3% |
| Category | Fanar 1 | Falcon H1R | Jais 2 | Prompts |
| Violent Crimes | 3.0% | 1.4% | 20.6% | 509 |
| Sex-Related Crimes | 1.8% | 3.7% | 18.0% | 217 |
| Child Sexual Exploitation | 4.7% | 2.8% | 17.0% | 106 |
| Suicide and Self-Harm | 10.9% | 6.4% | 14.6% | 110 |
| Indiscriminate Weapons | 8.7% | 3.5% | 32.0% | 231 |
| Intellectual Property | 54.8% | 4.8% | 69.3% | 166 |
| Category | Saudi | Moroccan | Syrian | Lebanese | Egyptian |
| Panel A: (pp) with | |||||
| Intellectual Property | (34/14) | (28/24) | (27/23) | (34/17) | (21/20) |
| Defamation | (16/12) | (13/16) | (16/12) | (18/10) | (16/15) |
| Nonviolent Crimes | (32/39) | (68/34) | (49/33) | (50/38) | (56/42) |
| Indiscriminate Weapons | (6/12) | (20/16) | (11/11) | (13/12) | (13/10) |
| Suicide and Self-Harm | (6/3) | (4/4) | (3/2) | (4/2) | (7/1) |
| Category | Saudi | Moroccan | Syrian | Lebanese | Egyptian |
| Panel A: (pp) with | |||||
| Intellectual Property | (29/9) | (42/13) | (28/15) | (25/16) | (25/15) |
| Indiscriminate Weapons | (10/31) | (20/32) | (13/30) | (12/27) | (6/34) |
| Nonviolent Crimes | (21/37) | (26/33) | (22/41) | (19/40) | (25/39) |
| Suicide and Self-Harm | (1/2) | (2/2) | (4/0) | (5/1) | (2/2) |
| Sexual Content | (13/37) | (27/38) | (15/37) | (21/36) | (17/39) |
| Category | Saudi | Moroccan | Syrian | Lebanese | Egyptian |
| Panel A: (pp) with | |||||
| Intellectual Property | (8/6) | (9/9) | (11/16) | (10/9) | (6/8) |
| Indiscriminate Weapons | (41/24) | (48/28) | (55/27) | (58/20) | (37/19) |
| Defamation | (18/5) | (18/12) | (19/9) | (22/8) | (21/6) |
| Nonviolent Crimes | (65/63) | (142/50) | (90/49) | (121/60) | (81/48) |
| Sexual Content | (82/106) | (109/84) | (102/96) | (90/90) | (95/92) |
| Annotation Field | Raw Agreement | Cohen’s | PABAK | Gwet’s AC1 |
| Response Safety (binary) | 85.83 [83.9, 87.5] | 0.714 [0.677, 0.750] | 0.717 [0.679, 0.753] | 0.721 [0.684, 0.756] |
| Response Safety (3-class) | 78.12 [75.9, 80.2] | 0.598 [0.562, 0.633] | 0.672 [0.640, 0.703] | 0.701 [0.670, 0.731] |
| Refusal | 90.83 [89.2, 92.2] | 0.669 [0.614, 0.718] | 0.817 [0.786, 0.846] | 0.873 [0.850, 0.895] |
| Arabic Variety (6-class) | 98.12 [97.3, 98.7] | 0.419 [0.231, 0.590] | 0.977 [0.968, 0.985] | 0.981 [0.973, 0.988] |
| Decision | Observed agreement [95% CI] | Cohen’s [95% CI] | |
| Q1 harmfulness, binary | 500 | 94.4% [92.0, 96.1] | 0.820 [0.754, 0.880] |
| Q2 category confirmation, binary | 291 | 94.2% [90.8, 96.3] | 0.839 [0.757, 0.908] |
| Derived final label, 12 classes | 500 | 91.0% [88.2, 93.2] | 0.888 [0.855, 0.918] |