cs.CLSep 28, 2026

AraDynFact: Dynamic Evaluation of Factual Knowledge in Arabic

Authors: Ignacio Iacobacci, Faroq Altam, Zhaozhi Qian, Muhammad Alqurishi

Organizations: Elm Company

Abstract

As Large Language Models (LLMs) continue to scale both in size and capabilities, their proficiency in the Arabic Language has seen significant advancement. However, a critical gap remains: the extent of their factual knowledge and cultural sensitivity to the diverse Arabic-speaking world remains largely underexplored. Current evaluation metrics often focus on translation or generic reasoning, failing to capture the rich historical, social, and regional nuances inherent to Arabic culture. In addition, most benchmarks rely on heavy work, with human intervention in some steps, making the evaluation of knowledge coverage expensive and slow. To address this deficiency, we introduce AraDynFact, a novel dynamic evaluation framework designed to rigorously assess the factual Arabic knowledge embedded in LLMs. Unlike static benchmarks, AraDynFact employs a dynamic approach to extract factual information and generate rich and answerable questions in a fast and automatic way. We apply AraDynFact to Arabic Wikipedia and audit the performance of several state-of-the-art models, ranging from Arabic-centric specialized LLMs to high-resource general purpose LLMs. In addition we found a high degree of correlation with existing, hand-crafted Arabic-centric benchmarks, confirming the potential of our dynamic approach.

Figures & tables

Explore similar work

CardsList
  1. ARAFA: An LLM-Generated Arabic Fact-Checking Dataset

    Sep 22, 2026Christophe Khalil, Shady Elbassuoni, Rida AssafArabicArabic Natural Language Processing

  2. Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues

    Apr 30, 2026Muhammad Dehan Al Kautsar, Saeed Almheiri, Momina Ahsan +13Arabic Natural Language ProcessingArabic