cs.CLSep 28, 2026

Pass or Fail? Evaluating LLMs on Two Greek Examination Benchmarks

Authors: Panagiota Kyriazi, Eleni Kasoura, Prokopis Prokopidis

Organizations: Institute for Language and Speech Processing / Athena RC

Abstract

The rapid advancement of Large Language Models (LLMs) imposes a thorough evaluation of their linguistic and analytical capabilities as well as constraints, particularly for a language with limited benchmark coverage such as Greek. To address the limited availability of comprehensive benchmarks in this domain, we introduce Prot-Ex and Pan-Ex, two benchmarks consisting of questions from entrance exams for Greek Model and Experimental schools as well as the Panhellenic exams (the Greek national university entrance examinations). These benchmarks are employed to assess the performance of text-only LLMs-including the Greek-adapted KriKri-8B-Instruct, Llama-3.1-8B, Gemma-4-26B, and Qwen-3-32B-across diverse academic disciplines (Modern Greek, Mathematics, Physics, etc.) and task formats (closed, structured, and open-ended), including textualized visual context (i.e., image descriptions). Our findings indicate the localized KriKri-8B significantly outperforms its base model, successfully rivalling much larger LLMs in linguistically demanding humanities tasks. By leveraging an LLM-as-a-Judge methodology, we expose the inadequacy of traditional lexical metrics for evaluating complex reasoning. Crucially, we uncover a few-shot prompting paradox: while synthetic examples improve accuracy in closed-ended questions, they severely overload the context window of 8B models in structured tasks, causing significant performance degradation. Ultimately, this study suggests targeted linguistic adaptation offsets lower parameter counts in specialized domains, despite the fragility of smaller models to prompt verbosity.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Maistros: A Greek Large Language Model Adapted Through Knowledge Distillation From Large Reasoning Models

    May 3, 2026Nikolaos Giarelis, Charalampos Mastrokostas, Nikos KaracapilidisGreekLarge Language Model Training

  2. KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding

    Jul 19, 2026Timur Turatali, Aida Turdubaeva, Rustem Izmailov +2Multilingual Benchmark

  3. BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams

    Jun 21, 2026João Guilherme Alves Santos, Giovana Kerche Bonás, Thiago Laitz +2Open-Ended ResponsesPortuguese