Organizations: McGill University · Mila-Quebec AI Institute · University of Toronto · Hanyang University, Rep. of Korea · Masakhane · Instituto Politécnico Nacional, Mexico · Umbaji · University of Ibadan, Nigeria · McPherson University, Nigeria · Canada CIFAR AI Chair
Large language models have made substantial progress in mathematical reasoning. However, benchmark development for multilingual evaluation has lagged behind English in both difficulty and recency. Recently, GSM-Symbolic showed a strong evidence of high variance when models are evaluated on different instantiations of the same question; however, the evaluation was conducted only in English. In this paper, we introduce MGSM-Pro, an extension of MGSM dataset with GSM-Symbolic approach. Our dataset provides five instantiations per MGSM question by varying names, digits and irrelevant context. Evaluations across nine languages reveal that many low-resource languages suffer large performance drops when tested on digit instantiations different from those in the original test set. We further find that models robustness in HRL setting do not necessarily translate to LRL. Moreover, proprietary models, such as Gemini 2.5 Flash and GPT-4.1 are less robust to digit, whereas Gemini 3.0 Pro is more robust. Among open models, GPT-OSS 120B and DeepSeek v3 show stronger robustness. Based on these findings, we recommend evaluating each problem using at least five digit-varying instantiations to obtain a more robust and realistic assessment of math reasoning.
Figures & tables
Figure 1: Relative decrease in accuracy from the original dataset to five instances of changing both names and numbers and adding irrelevant context, averaged over all nine languages.
Figure 2: MGSM-Pro creation diagram, illustrating both template creation and multilingual data construction steps on a sample language.
Gemini 2.5 Flash
Gemini 3.0 Pro
Claude 4 Sonnet
GPT-4.1
GPT-5
Δ Ave.
Δ Med.
Language
DO
IC_N
SYM_#
IC_#
DO
IC_N
SYM_#
IC_#
DO
IC_N
SYM_#
IC_#
DO
IC_N
SYM_#
IC_#
DO
IC_N
SYM_#
IC_#
–
–
English
96.8
94.6
83.1
81.0
98.0
96.2
94.4
93.2
98.0
96.0
91.5
90.4
96.4
91.7
81.5
79.6
96.8
93.3
92.8
88.0
−10.7
−8.8
Chinese
89.9
89.0
78.0
79.0
93.5
93.1
93.2
93.5
93.1
91.9
88.7
88.5
89.9
90.6
78.6
76.8
92.3
91.9
89.3
88.4
−6.5
−4.7
French
91.5
87.1
77.5
73.8
90.7
89.7
88.6
87.2
91.5
90.4
86.4
85.3
88.7
85.5
75.9
74.5
89.9
88.4
85.4
84.2
−9.5
−6.2
Japanese
86.7
83.9
74.9
73.1
90.7
89.4
87.6
88.0
89.9
84.8
83.1
81.5
87.1
83.6
74.8
74.1
90.7
84.1
83.6
82.6
−9.2
−8.5
Swahili
91.5
89.9
80.3
78.7
97.6
93.9
93.0
92.4
91.9
90.9
85.1
84.4
91.5
89.0
79.8
77.5
90.7
92.4
87.7
89.3
−8.2
−7.5
Table 1: Different closed models’ accuracy across dataset variations ( SYM_# , IC_N , IC_# ) and original ( DO ). Cells are shaded by how far each variant falls below its row’s DO within the same model group; values at or above DO are left white. We report the Average (Ave.) and Median (Med.) of Δ(IC_#−DO) taken across the five models.
Gemma 3 27B
Qwen 3 32B
Qwen 3.5 27B
DeepSeek V3
GPT-OSS 120B
Δ Ave.
Δ Med.
Language
DO
IC_N
SYM_#
IC_#
DO
IC_N
SYM_#
IC_#
DO
IC_N
SYM_#
IC_#
DO
IC_N
SYM_#
IC_#
DO
IC_N
SYM_#
IC_#
–
–
English
95.6
93.3
81.6
77.6
89.1
87.3
87.7
86.5
98.4
94.2
93.8
90.7
98.4
94.3
92.7
89.7
96.4
93.9
93.5
92.1
−8.3
−7.7
Chinese
87.9
85.5
75.2
72.3
89.9
88.8
87.8
88.1
89.9
90.1
87.3
83.9
92.3
91.4
89.9
88.1
91.5
90.6
90.2
89.7
−5.9
−4.2
French
89.1
84.5
73.4
69.9
89.5
87.6
86.3
84.0
91.1
87.2
87.2
84.1
90.7
88.7
87.8
85.7
90.7
88.7
86.5
86.5
−8.2
−5.5
Japanese
85.1
79.4
71.1
64.6
88.7
85.2
84.8
83.3
87.5
80.1
76.0
72.1
88.7
83.0
79.8
79.8
89.1
85.2
85.2
83.9
−11.1
−8.9
Swahili
89.1
85.4
72.2
71.2
78.2
71.0
73.6
67.3
91.9
88.5
88.1
84.9
90.7
89.6
86.3
84.0
84.7
82.2
81.5
79.6
−9.5
−7.0
Table 2: Different open models’ accuracy across dataset variations ( SYM_# , IC_N , IC_# ) and original ( DO ). Cells are shaded by how far each variant falls below its row’s DO within the same model group; values at or above DO are left white. We report the Average (Ave.) and Median (Med.) of Δ(IC_#−DO) taken across the five models.
Figure 3: Comparison of relative accuracy decrease from HRL and LRL Do , averaged across six variants of MGSM-Pro. Models are sorted in descending order of average accuracy drop.
Figure 4: Relative Accuracy Drop Across Model Families. The figures illustrate the relative decline in accuracy for (a) the Gemma-3 family and (b) the GPT-OSS family. The drop is measured from the original dataset to two configurations: IC_N# and SYM_N#, averaged over nine languages.
All Language
High-Resource
Low-Resource
Model
DO
Avg-3
Avg-5
Avg-10
Avg-5
Avg-5
±
±
±
±
±
Gemini 3.0 Pro
1
90.2
1
–
84.2
2.7
1
–
84.1
1.5
1
–
84.2
0.9
1
89.6
1.3
1
79.7
1.6
Gemini 3.0 Flash
2
89.2
2
–
82.6
2.3
2
–
82.3
1.5
2
–
82.4
0.8
2
89.5
1.5
2
76.5
1.5
Gemini 2.5 Flash
3
85.4
6
70.3
3.3
7
70.1
1.8
7
–
69.6
1.1
12
75.4
1.6
3
65.8
2.0
Claude 4
4
83.9
3
74.3
3.2
3
–
74.2
1.5
3
–
73.9
1.0
4
85.9
1.4
4
64.8
1.6
Table 3: Model ranking and average accuracy on IC_N# under various resource levels, with 3, 5, and 10 instances per problem. Sub-columns include: rank, rank change vs. the preceding metric ( rise, fall, – unchanged), accuracy, ± half-width of the 95% confidence interval. Rows are sorted by DO rank.
Languages
DeepSeek V3
Gemini 2.5 Flash
English Solve
Native Solve
English Solve
Native Solve
DO
Sym_#
Δ
DO
Sym_#
Δ
DO
Sym_#
Δ
DO
Sym_#
Δ
English
98.4
92.7
-5.7
97.6
91.8
-5.8
96.8
83.1
-13.7
95.2
83.8
-11.4
Chinese
92.3
89.9
-2.4
93.2
90.3
-2.9
89.9
78.0
-11.9
90.0
79.3
-10.7
French
90.7
87.8
-2.9
90.8
86.0
-4.8
91.5
77.5
-14.0
90.8
79.0
-11.8
Japanese
88.7
79.8
-9.0
89.6
82.8
-6.8
86.7
74.9
-11.8
88.0
74.4
-13.6
Table 4: Comparison of native and English solve on 5-instances of SYM_# with Deepseek V3 and Gemini 2.5 . We report Δ ( SYM_# −Do )
DeepSeek V3
Gemini 2.5 Flash
Language
English
1.2
5.2
4.0
0.8
4.8
14.8
Chinese
4.0
8.4
2.4
5.2
8.4
15.2
French
3.5
4.8
3.6
3.2
8.8
16.0
Japanese
9.2
13.6
4.8
6.8
7.2
13.2
Swahili
7.2
12.8
3.6
4.8
7.6
10.0
Table 5: Error analysis of LLM models: DeepSeek V3 and Gemini 2.5 Flash across high- and low-resource languages via GPT-5.4 as a judge. Errors rates (%) are categorized into Linguistic ( ), Logic ( ), and Arithmetic ( ).
Figure 5: GPT-OSS 20B average accuracy over all languages on DO , IC_N , SYM_# , and IC_# under 0-shot and 8-shot prompting.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Language
Code
Language Family
Joshi Class
English
eng_Latn
Indo-European
Class 5
Chinese
zho_Hans
Sino-Tibetan
Class 5
French
fra_Latn
Indo-European
Class 5
Japanese
jpn_Jpan
Japonic
Class 5
Swahili
swh_Latn
Niger-Congo
Class 2
Amharic
amh_Ethi
Afro-Asiatic
Class 2
Appendix
Table 6: Selected languages categorized by ISO code, linguistic family, and resource availability (Joshi Class).
Domain
Name Types
People
Male name, Female name, Family name
Places
City name, Mountain name
Pet
Dragon name, Dinosaur name, Cat name
Appendix
Table 7: Grouped name variables categorized by domain
Language
Template Correction Rate (%)
English
–
Chinese
4.4
French
7.6
Japanese
9.2
Swahili
72.4
Amharic
67.2
Appendix
Table 8: Percentage of templates requiring correction by native annotators, per language. English is the source language and therefore not applicable.
Misjudg.
Lang.
Logic
Arith.
Human 1
98.25
96.49
96.49
100.0
Human 2
100.0
96.49
98.25
100.0
Appendix
Table 9: Human–model agreement rates (%) on error classification for Chinese. Misjudg. = Misjudgment, Lang. = Language Understanding, Logic = Logical Reasoning, Arith. = Arithmetic Calculation.
Misjudg.
Lang.
Logic
Arith.
Human 1
98.68
96.05
92.11
98.68
Human 2
100.0
100.0
90.67
98.67
Appendix
Table 10: Human–model agreement rates (%) on error classification for Yoruba. Column abbreviations as in Table 9 .
Lang.
Do
Sym#
ICN
IC#
0-S
8-S
0-S
8-S
0-S
8-S
0-S
8-S
English
95.6
97.6
88.9
91.1
87.7
93.0
80.6
86.8
Chinese
89.1
88.7
84.4
86.6
86.6
88.1
81.9
84.0
French
87.9
87.5
84.1
84.6
85.6
87.3
81.9
82.7
Japanese
85.1
85.5
81.3
80.8
79.0
82.0
77.2
79.8
Swahili
74.6
73.0
66.2
68.7
61.5
68.5
53.9
61.5
Appendix
Table 11: 0-Shot(0-S) and 8-Shot(8-S) performance across Do, SYM_# , IC_N , and IC_# for various languages for GPT-OSS-20B
Gemini 2.0 Flash
Gemini 2.5 Flash
Gemini 3 Flash
Language
DO
SYM_N
IC_N
SYM_#
IC_#
SYM_N#
IC_N#
DO
SYM_N
IC_N
SYM_#
IC_#
SYM_N#
IC_N#
DO
SYM_N
IC_N
SYM_#
IC_#
SYM_N#
IC_N#
English
96.0
94.1
94.0
85.2
84.3
84.5
81.4
96.8
95.0
94.6
83.1
81.0
80.6
80.1
98.8
97.3
95.9
95.3
94.0
94.9
93.3
Chinese
86.7
89.4
86.5
80.0
77.9
80.7
76.5
89.9
90.7
89.0
78.0
79.0
78.3
76.2
93.1
92.3
91.5
91.7
91.4
91.3
90.2
French
90.7
86.7
87.2
77.2
76.5
78.5
76.7
91.5
87.3
87.1
77.5
73.8
75.8
73.5
92.3
89.7
89.2
88.1
87.6
87.1
86.5
Japanese
84.3
84.4
82.1
75.9
72.4
74.4
70.4
86.7
85.0
83.9
74.9
73.1
73.2
71.7
90.7
89.5
88.5
88.9
88.4
88.6
88.0
Swahili
91.9
90.8
87.6
79.0
77.9
78.1
77.3
91.5
90.7
89.9
80.3
78.7
78.1
77.7
96.8
94.0
93.0
91.6
90.8
90.5
89.4
Appendix
Table 12: Different models’ accuracy across different dataset variations ( DO , SYM_N , IC_N , SYM_# , IC_# , SYM_N# , IC_N# ) for each language
Gemini 3.0 Pro
GPT 4.1
GPT 5
Language
DO
SYM_N
IC_N
SYM_#
IC_#
SYM_N#
IC_N#
DO
SYM_N
IC_N
SYM_#
IC_#
SYM_N#
IC_N#
DO
SYM_N
IC_N
SYM_#
IC_#
SYM_N#
IC_N#
English
98.0
96.3
96.2
94.4
93.2
93.5
92.8
96.4
94.7
91.7
81.5
79.6
81.1
79.9
96.8
-
93.3
92.8
-
-
87.6
Chinese
93.5
93.5
93.1
93.2
93.5
92.5
92.3
89.9
91.4
90.6
78.6
76.8
78.3
76.4
92.3
-
91.9
89.3
-
-
87.1
French
90.7
89.3
89.7
88.6
87.2
88.5
86.0
88.7
86.7
85.5
75.9
74.5
76.0
74.0
89.9
-
88.4
85.4
-
-
83.0
Japanese
90.7
89.0
89.4
87.6
88.0
88.0
87.2
87.1
85.8
83.6
74.8
74.1
73.5
73.2
90.7
-
83.6
83.6
-
-
81.1
Swahili
97.6
94.5
93.9
93.0
92.4
91.3
91.2
91.5
90.0
89.0
79.8
77.5
77.5
77.3
90.7
-
87.7
87.7
-
-
87.9
Appendix
Table 13: Different models’ accuracy across different dataset variations ( DO , SYM_N , IC_N , SYM_# , IC_# , SYM_N# , IC_N# ) for each language
GPT-OSS 20 B
GPT-OSS 120B
DeepSeek V3
Language
DO
SYM_N
IC_N
SYM_#
IC_#
SYM_N#
IC_N#
DO
SYM_N
IC_N
SYM_#
IC_#
SYM_N#
IC_N#
DO
SYM_N
IC_N
SYM_#
IC_#
SYM_N#
IC_N#
English
95.6
94.9
87.7
88.9
80.6
88.1
78.8
96.4
96.2
93.9
93.5
92.1
92.0
90.6
98.4
96.6
94.3
92.7
89.7
91.9
89.5
Chinese
89.1
89.0
86.6
84.4
81.9
84.0
81.3
91.5
92.3
90.6
90.2
89.7
90.4
88.1
92.3
93.2
91.4
89.9
88.1
89.1
87.2
French
87.9
87.3
85.6
84.1
81.9
83.3
79.9
90.7
88.6
88.7
86.5
86.5
86.5
83.9
90.7
89.9
88.7
87.8
85.7
86.6
85.7
Japanese
85.1
83.2
79.0
81.3
77.2
81.0
75.7
89.1
86.6
85.2
85.2
83.9
84.5
82.9
88.7
83.9
83.0
79.8
79.8
80.8
79.8
Swahili
74.6
74.9
61.5
66.2
53.9
64.1
54.7
84.7
84.7
82.2
81.5
79.6
81.4
79.5
90.7
90.5
89.6
86.3
84.0
85.9
85.2
Appendix
Table 14: Different models’ accuracy across different dataset variations ( DO , SYM_N , IC_N , SYM_# , IC_# , SYM_N# , IC_N# ) for each language
Gemma 3 4B
Gemma 3 12B
Gemma 3 27B
Language
DO
SYM_N
IC_N
SYM_#
IC_#
SYM_N#
IC_N#
DO
SYM_N
IC_N
SYM_#
IC_#
SYM_N#
IC_N#
DO
SYM_N
IC_N
SYM_#
IC_#
SYM_N#
IC_N#
English
85.5
85.6
80.0
68.1
64.0
66.5
61.3
92.7
92.8
92.7
76.9
77.3
75.8
76.5
95.6
94.4
93.3
81.6
77.6
80.2
77.6
Chinese
76.2
78.6
67.8
63.5
54.0
60.2
53.1
86.7
88.0
84.9
70.9
65.6
70.6
66.0
87.9
89.7
85.5
75.2
72.3
76.4
72.1
French
79.0
78.2
69.3
59.0
53.8
60.3
52.7
87.9
85.9
81.8
71.7
66.2
70.2
65.4
89.1
87.4
84.5
73.4
69.9
74.4
69.0
Japanese
70.2
67.3
58.1
51.9
45.2
52.1
43.4
83.5
81.2
79.4
66.1
60.2
65.3
59.8
85.1
83.1
79.4
71.1
64.6
70.5
64.8
Swahili
59.7
64.4
55.6
48.7
40.6
46.6
39.0
81.5
86.5
81.4
65.6
61.1
65.6
61.7
89.1
87.7
85.4
72.2
71.2
72.3
70.1
Appendix
Table 15: Different models’ accuracy across different dataset variations ( DO , SYM_N , IC_N , SYM_# , IC_# , SYM_N# , IC_N# ) for each language
Claude 4
Llama 3 70B
Gemma 2 27B
Language
DO
SYM_N
IC_N
SYM_#
IC_#
SYM_N#
IC_N#
DO
SYM_N
IC_N
SYM_#
IC_#
SYM_N#
IC_N#
DO
SYM_N
IC_N
SYM_#
IC_#
SYM_N#
IC_N#
English
98.0
96.4
96.0
91.5
90.4
90.8
90.5
95.2
93.4
88.4
68.9
65.0
68.7
64.4
90.3
80.5
78.5
56.0
53.0
54.2
51.8
Chinese
93.1
93.5
91.9
88.7
88.5
88.7
87.6
69.8
73.9
66.9
51.5
41.5
51.6
43.5
81.0
86.8
80.5
55.7
52.3
58.1
53.3
French
91.5
90.4
90.4
86.4
85.3
84.9
83.2
75.0
73.9
64.8
47.6
41.2
49.5
41.7
84.7
68.1
58.7
51.5
39.8
49.7
39.9
Japanese
89.9
86.3
84.8
83.1
81.5
81.2
82.1
69.8
65.9
54.5
43.0
34.8
42.3
35.2
79.0
76.9
68.1
52.2
44.0
52.1
44.5
Swahili
91.9
92.6
90.9
85.1
84.4
85.1
83.9
63.7
64.4
50.8
41.6
32.4
39.8
31.2
86.7
82.4
74.0
54.7
47.6
53.5
48.1
Appendix
Table 16: Different models’ accuracy across different dataset variations ( DO , SYM_N , IC_N , SYM_# , IC_# , SYM_N# , IC_N# ) for each language
Qwen 3 32B
Qwen 3.5 27B
Gemma 2 9B
Language
DO
SYM_N
IC_N
SYM_#
IC_#
SYM_N#
IC_N#
DO
SYM_N
IC_N
SYM_#
IC_#
SYM_N#
IC_N#
DO
SYM_N
IC_N
SYM_#
IC_#
SYM_N#
IC_N#
English
89.1
84.7
87.3
87.7
86.5
86.0
86.0
98.4
97.2
94.2
93.8
90.7
92.0
88.9
75.9
77.9
76.7
45.3
47.3
44.7
48.5
Chinese
89.9
90.9
88.8
87.8
88.1
88.2
86.1
89.9
91.5
90.1
87.3
83.9
86.4
85.1
78.4
83.3
76.8
49.0
46.1
49.2
44.2
French
89.5
86.9
87.6
86.3
84.0
85.1
83.5
91.1
89.8
87.2
87.2
84.1
86.9
83.6
79.6
80.1
75.2
48.6
46.1
48.5
45.6
Japanese
88.7
86.2
85.2
84.8
83.3
84.3
82.9
87.5
83.7
80.1
76.0
72.1
76.1
68.5
75.1
70.0
64.4
45.1
39.3
43.2
39.9
Swahili
78.2
79.2
71.0
73.6
67.3
71.3
64.6
91.9
90.9
88.5
88.1
84.9
88.2
84.0
69.8
71.3
68.6
44.4
42.4
43.2
43.8
Appendix
Table 17: Different models’ accuracy across different dataset variations ( DO , SYM_N , IC_N , SYM_# , IC_# , SYM_N# , IC_N# ) for each language
Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. The recently released PolyMath (Wang et al., 2025) dataset represents a significant step forward, yet its coverage is still limited to 18 only high-resource languages. To address this gap, we introduce PluraMath, an extension of PolyMath to 18 additional {underrepresented languages spanning 6 language families -- ranging from mid-resource to extreme low-resource settings. We constructed the dataset through a human-curated pipeline, where native speakers thoroughly validated pre-computed translations. Using PluraMath, we then benchmark 27 reasoning LLMs across four model scales -- small, mid-size, large, and closed-source ensembles -- probing the multilingual mathematical reasoning capabilities of state-of-the-art models under diverse linguistic conditions. Our fine-grained analysis confirms a persistent gap in mathematical reasoning performance between high-resource and underrepresented languages, with stronger results largely associated with better instruction-following ability. We fully open-source our dataset, data acquisition pipeline, and evaluation framework, with the goal of lowering the barrier to multilingual benchmark development for underrepresented communities.
Daryna Dementieva, Nikolay Babakov, Kathy Hämmerl +14
Technical University of Munich (TUM) · Munich Center for Machine Learning (MCML) · Independent Researcher +7
Large reasoning models (LRMs) achieve strong mathematical reasoning performance in English, but remain much less reliable in many low- and medium-resource languages. This gap is often explained as a failure to understand non-English problem statements. We show that this view is incomplete: even when the problem is given in English, controlling the model's reasoning language can substantially reduce accuracy, suggesting that language also affects reasoning execution itself. To study this effect, we introduce DATG, a Directed Acyclic Trace Graph framework that maps reasoning traces to language-independent mathematical anchors and dependencies. This allows us to align target-language traces with reference DAGs and measure whether they cover required mathematical nodes, respect dependency edges, and avoid harmful mathematical actions. Experiments on the Qwen3 series across 12 languages show that non-English reasoning often suffers from reduced anchor coverage and weaker dependency fidelity, especially in low-resource languages. Motivated by this diagnosis, we propose Loop-Retry and Formula-Retry, two simple test-time controls targeting DATG-exposed failure modes, and show that they consistently improve target-language reasoning performance in low-resource languages.
The evaluation of mathematical reasoning in large language models (LLMs) has predominantly focused on high-resource languages like English. This has created a significant barrier to the equitable development and deployment of AI in linguistically diverse regions such as Bangladesh, where over 230 million people speak Bengali. Despite this global significance, there has been minimal prior work on mathematical reasoning in Bengali and no existing research that systematically benchmarks a perturbated Bengali mathematical dataset, leaving a critical void in assessing model robustness and true comprehension beyond pattern recognition. This study addresses this gap by introducing GSM-Plus-BN, a novel perturbated Bengali mathematical dataset derived from the English GSM-Plus benchmark and verified by human translators. We evaluate six open-source LLMs Qwen3-32B, Llama-3.1-8B-Instant, Llama-3.3-70B-Versatile, Llama-4-Scout-17B-16E-Instruct, GPT-OSS-120B, and GPT-OSS-20B using a benchmark of 9,000 evaluation samples comprising 1,000 seed questions and 8,000 perturbed variants under both Standard Prompting and Chain-of-Thought (CoT) Prompting. Experimental results show that GPT-OSS-20B achieves the highest seed question accuracy of 96.08% under Standard Prompting, while larger models such as Llama-3.3-70B and GPT-OSS-120B demonstrate superior robustness across perturbation types. Furthermore, CoT prompting substantially improves reasoning for most models compared to Standard Prompting, yet a notable performance gap persists across all models relative to their English benchmarks, underscoring the inherent difficulty of perturbed Bengali text. This research makes a foundational contribution by providing GSM-PLUS-BN as a new resource and baseline for future Bengali mathematical reasoning research.
Computer Science and Engineering Southeast University Dhaka, Bangladesh · Computer Science and Engineering Ahsanullah University of Science and Technology Dhaka, Bangladesh