cs.CLJan 29, 2026

MGSM-Pro: A Simple Strategy for Robust Multilingual Mathematical Reasoning Evaluation

Authors: Tianyi Xu, Kosei Uemura, Alfred Malengo Kondoro, Tadesse Destaw Belay, Catherine Nana Nyaah Essuman, Ifeoma Okoh, Ganiyat Afolabi, Ayodele Awokoya, +1 more

Organizations: McGill University · Mila-Quebec AI Institute · University of Toronto · Hanyang University, Rep. of Korea · Masakhane · Instituto Politécnico Nacional, Mexico · Umbaji · University of Ibadan, Nigeria · McPherson University, Nigeria · Canada CIFAR AI Chair

Abstract

Large language models have made substantial progress in mathematical reasoning. However, benchmark development for multilingual evaluation has lagged behind English in both difficulty and recency. Recently, GSM-Symbolic showed a strong evidence of high variance when models are evaluated on different instantiations of the same question; however, the evaluation was conducted only in English. In this paper, we introduce MGSM-Pro, an extension of MGSM dataset with GSM-Symbolic approach. Our dataset provides five instantiations per MGSM question by varying names, digits and irrelevant context. Evaluations across nine languages reveal that many low-resource languages suffer large performance drops when tested on digit instantiations different from those in the original test set. We further find that models robustness in HRL setting do not necessarily translate to LRL. Moreover, proprietary models, such as Gemini 2.5 Flash and GPT-4.1 are less robust to digit, whereas Gemini 3.0 Pro is more robust. Among open models, GPT-OSS 120B and DeepSeek v3 show stronger robustness. Based on these findings, we recommend evaluating each problem using at least five digit-varying instantiations to obtain a more robust and realistic assessment of math reasoning.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

    Jul 7, 2026Daryna Dementieva, Nikolay Babakov, Kathy Hämmerl +14Multilingual BenchmarkTable Reasoning Datasets

  2. GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models

    Jul 14, 2026Bidyarthi Paul, Nahida Jannat Mayouree, Md. Asif Karim +2BanglaLarge Language Model Benchmarks