Large language models (LLMs) often default to Modern Standard Arabic (MSA) when generating Arabic, even when prompted with dialectal Arabic. A natural explanation is that their internal representations are dominated by MSA. We test this hypothesis by adapting the language-dominance framework of Shani and Basirat (2025) (https://doi.org/10.18653/v1/2025.blackboxnlp-1.7) to 26 Arabic varieties. Across layers and model families, we find no evidence that MSA acts as a dominant internal representation for Arabic dialects. Instead, dialect representations form a dense and highly overlapping space: normalized mutual information drops sharply compared to patterns reported for more distinct languages. Moreover, the strongest separability effects are not confined to intermediate layers, but can shift toward later layers depending on the architecture. These findings challenge a common interpretation of MSA-biased generation: output preference does not necessarily reveal internal representational dominance. Analyses of multilingual and dialectal LLMs should therefore distinguish generation bias from the geometry of internal representations.
Figures & tables
Figure 1: Layer-wise NMI values for cross-lingual vs. cross-dialectal representations for mGPT (left) and BLOOM (right). Dialectal NMI is lower and remains low over a wider middle-layer range. Cross-lingual representations were computed using Parallel Universal Dependencies (PUD) treebanks ( Zeman et al., 2017 ) (following Shani and Basirat (2025) ), cross-dialectal representations were computed using the MADAR corpus ( Bouamor et al., 2018 ) .
Figure 2: Confusion matrices for dialect clustering in mGPT (left) and BLOOM (right) grouped by country. Clear diagonals show that dialect identity is preserved despite low separability.
Figure 3: Token-level dominance statistics. Percentage of tokens with Λ<1 . A mostly flatline percentage of tokens are being assigned to competing dialects.
Figure 4: Layer-wise NMI values for Meta-Llama-3.1-8B-Instruct. The model maintains relatively stable dialect separability across middle and later layers.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Supervised linear probe accuracy across layers for BLOOM. Accuracy ranges from 0.19 to 0.25.
Figure 8: Supervised linear probe accuracy across layers for mGPT. Accuracy ranges from 0.14 to 0.21.
Figure 9: Dialect dominance scores across layers for mGPT (top) and BLOOM (bottom). Dominance remains weak across layers.
Figure 10: Complete dialectal confusion matrices for mGPT (top) and BLOOM (bottom).
Figure 11: Replicated NMI curves for mGPT (top) and BLOOM (bottom).
Figure 12: Replicated confusion matrices for mGPT (left) and BLOOM (right).
Figure 13: Replicated layer-wise dominance scores for mGPT (top) and BLOOM (bottom).
Figure 14: Replicated token-level dominance statistics for mGPT (left) and BLOOM (right).
Figure 15: Replicated token-level Λ histograms for mGPT (left) and BLOOM (right).