LLM-as-a-Judge for Low-Resource Languages: Adapting Ragas and Comparative Ranking for Romanian
Organizations: Interdisciplinary School of Doctoral Studies, University of Bucharest, Romania · HLT Research Center, University of Bucharest, Romania · Faculty of Mathematics and Computer Science, University of Bucharest, Romania
Abstract
Evaluating Retrieval-Augmented Generation (RAG) systems remains a challenge for Low-Resource Languages (LRLs), where standard reference-based metrics fall short. This paper investigates the viability of the "LLM-as-a-Judge" paradigm for Romanian by adapting the Ragas framework using next-generation models (Gemini 2.5 and Gemini 3). We introduce AdminRo-Eval, a curated dataset of Romanian administrative documents annotated by native speakers, to serve as a ground truth for benchmarking automated evaluators. We compare three evaluation methodologies - direct scoring, comparative ranking, and granular decomposition - across metrics for Faithfulness, Answer Relevance, and Context Relevance. Our findings reveal that evaluation strategies must be metric-specific: granular decomposition achieves the highest human alignment for Faithfulness (96% with Gemini 2.5 Pro), while comparative ranking outperforms in Answer Relevance (90%). Furthermore, we demonstrate that while lightweight models struggle with complex reasoning in LRLs, the Gemini 2.5 Pro architecture establishes a robust, transferable baseline for automated Romanian RAG evaluation.
Figures & tables
| Model | Faithfulness | Answer Relevancy | Context Relevancy |
|---|---|---|---|
| Romanian Dataset Evaluation | |||
| Gemini 3 Pro | Ragas(84%) Baseline(82%) Ranking(81%) | Ragas(76%) Baseline(83%) Ranking(80%) | Ragas(93%) Baseline(0%) Ranking(8%) |
| Gemini 3 Flash | Ragas(87%) Baseline(76%) Ranking(76%) | Ragas(76%) Baseline(76%) Ranking(77%) | Ragas(0%) Baseline(2%) Ranking(5%) |
| Gemini 2.5 Pro | Ragas(96%) Baseline(86%) Ranking(85%) | Ragas(85%) Baseline(83%) Ranking(90%) | Ragas(79%) Baseline(70%) Ranking(65%) |
| English Dataset Evaluation | |||
| Gemini 3 Pro | Ragas(94%) Baseline(90%) Ranking(92%) | Ragas(84%) Baseline(89%) Ranking(92%) | Ragas(100%) Baseline(90%) Ranking(90%) |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.