cs.CLMay 26, 2026

PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions

Authors: Ruhallah NiaziFaeze GhorbanpourAlexander Fraser

Organizations: School of Computation, Information and Technology, TU Munich · Munich Center for Machine Learning (MCML)

Abstract

Despite impressive multilingual capabilities, large language models (LLMs) remain poorly evaluated on literary knowledge in non-English languages. We introduce PersLitEval, a benchmark of 4,514 Persian literature multiple-choice questions across eight fine-grained categories spanning spelling, literary devices, grammar, vocabulary, word formation, and conceptual understanding, sourced from materials for the Konkur university entrance examination. We evaluate six LLMs across ten prompting strategies, revealing striking category-level disparities across three tiers of task difficulty: models reach higher accuracy on conceptual similarity tasks but struggle with formal linguistic analysis, with spelling and word formation proving the hardest across all models. Prompting strategy has a significant impact on performance, with explained few-shot examples yielding the best results, particularly on formal linguistic categories. An error analysis identifies three failure modes: semantic comprehension gaps, formal linguistic knowledge gaps, and counting/enumeration errors, suggesting that different categories require different improvement strategies.

Explore similar work

CardsList
  1. CLIN: an Objective Framework for Evaluating Creativity in Short Persian Literary Text

    Aug 31, 2026Mohammad Reza Modarres, Armin Tourajmehr, Yadollah Yaghoobzadeh +1CreativityDiversity

  2. PARSA-Bench: A Comprehensive Persian Audio-Language Model Benchmark

    Mar 15, 2026Mohammad Javad Ranjbar Kalahroodi, Mohammad Amini, Parmis Bathayan +2PersianLarge Audio Language Models