cs.CLOct 7, 2026

Large Language Models for Machine Translation Quality Annotation: Humans and Models Are Both Challenged

Authors: Hala Almaghout, Christian Federmann, Qin Gao

Organizations: Apple

Abstract

Large Language Models (LLMs) are considered to be a more efficient and cost-effective alternative to human judgment for Machine Translation (MT) evaluation. With MT evaluation spanning a large number of language pairs, domains and levels of annotation granularity, LLMs must be thoroughly evaluated across these dimensions before being reliably used as alternatives to human evaluation. In this paper, we evaluate the performance of LLMs for two prominent MT quality evaluation schemes: Multidimensional Quality Metrics (MQM) and Error Span Annotation (ESA) by comparing their agreement with human annotators. We present results on a long-context test set of 70 language pairs and the publicly available WMT23 and WMT25 data, investigating both score and error span annotation agreement across a variety of language pairs and domains. Our results show that while LLM agreement with human annotators exceeds agreement between human annotators for some evaluation tasks, both vary substantially across annotation schemes, language pairs and domains and remain unreliable for most settings. Furthermore, we identify challenges facing both human and LLM annotators: humans are particularly challenged by fine-grained MQM annotations and low-resource language pairs, while LLMs struggle with minor errors, wrong language variants and error span annotation. Our results highlight the potential for improvement for both human and LLM annotation performance, possibly through human-LLM collaborative annotation pipelines that address the reliability issues identified in this work.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LQM: Linguistically Motivated Multidimensional Quality Metrics for Machine Translation

    Apr 20, 2026Samar M. Magdy, Fakhraddin Alwajih, Abdellah El Mekki +2Arabic NLPMachine Translation

  2. Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation

    May 24, 2026Klaudia-Doris Thellmann, Bernhard Stadler, Michael Färber +1Multilingual Language Model EvaluationMT Evaluation

  3. CompactQE: Interpretable Translation Quality Estimation via Small Open-Weight LLMs

    May 15, 2026Kamil Guttmann, Zofia Fraś, Artur Nowakowski +1Small Language ModelsMT Evaluation