Multilingual GSM-Symbolic: What determines capability transfer across languages?
Authors: Kenneth Enevoldsen, Riley Herchert, Sofie Mosegaard, Dan Saattrup Smart, Simon Enni, Isaac Chung, Sofie Bruun, Ayush Sunil Munot, +17 more
Organizations: Aarhus University · Danish Foundation Models · University of Alabama · Alexandra Institute · Indian Institute of Technology Kharagpur · The University of Tokyo · IT University of Copenhagen · Massachusetts General Hospital · Bocconi University · University of Southern Denmark · University of Iceland · University of the Faroe Islands · Zendesk · University of Copenhagen · Indian Institute of Technology Madras · National Library of Sweden
We understand little about how capabilities acquired in one language carry over to another, or what governs this transfer: evaluations rely on incomparable, saturation-prone datasets and rarely examine its determinants jointly. Identifying what predicts transfer would let us avoid exhaustive evaluation across all language pairs and let developers target the factors that limit performance in low-resource languages. To evaluate cross-lingual capability transfer, we introduce Multilingual GSM-Symbolic, an extensible multilingual mathematical dataset covering 30,000 item-matched question-answer pairs and spanning 15 languages. It utilises symbolic templates to prevent overfitting and ensure generalisation by allowing generation of millions of high-quality variations from a single sample. Using Multilingual GSM-Symbolic, we quantify the largest determinants of capability as model size (β=1.77), language resource level (β=0.77), reasoning (β=0.67) and typological distance (β=−0.25). This joint estimation allows these determinants to be expressed in terms of one another: a 32B model evaluated in Marathi performs like a 10B model in English. Our findings have important implications for model developers, showing that model size and reasoning narrow the performance gap between low- and high-resource languages (β=−0.27 and β=−0.20, respectively), while similar levers have little or no effect on typologically distant languages. Overall, our analysis framework explains 92% of between-language variation, but only 23% of the model-by-language variation, and predicts a model's performance on an unseen language within 6.0pp (r=.96). Incorporating measurements from just 10 templates in the target language reduces this to 4.19pp, enabling reasonable estimates of performance with little or no downstream dataset.
Figures & tables
Figure 1: Overview of Multilingual GSM-Symbolic : Consisting of 100 samples in 15 languages, multilingual GSM-Symbolic consists of 1500 verified and localized templates. We sample 20 per template for a total of 30,000 and show an example evaluation of Qwen2.5-7B across English, Danish and Marathi. Showing both the language gap between Marathi and English and synthetic gap between the original samples and their template samples distributions.
Validation
Dataset
Languages
Symbolic
Code
Extensible
Error Analysis
Translation
Localization
Validity
Multilingual GSM-symbolic (ours)
15 (100)
✓
✓
✓
✓
✓
✓
✓
Previous work
MGSM ( Shi et al., 2023 )
11
✓
(✓)
GSM-Symbolic ( Mirzadeh et al., 2025 )
1
✓
GSM8K-Platinum ( Vendrow et al., 2025 )
1
✓
Table 1: Comparison of GSM-style mathematical reasoning benchmarks . The parenthesis refer to languages that have not been human validated yet. For more see A.2 Our dataset combines multilingual templates, controllable variation and extensibility with extensive validation. Error analysis refer to the process of running a model and correcting ambiguous or incorrect questions. Extensible as defined by Enevoldsen et al. (2026) , indicate that there is a process for adding to, extending or correcting the dataset. Validity denote automated tests that ensures the quality of the generated examples. While MGSM and MGSM-Pro localizes languages they do not localize names, currency and metrics even though imperial units are rarely used in a non-English contexts.
Figure 3
Figure 4: Levers : Model prediction for four interaction terms denoting two levers, showing the effect of reasoning and size across resource level and typological distance. Depending on the language you are working with, the method for improving capability transfer differs.
Figure 5
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6
Figure 9: How much data is needed . Left: the standard error of a single (model, language) accuracy at a given template budget, which falls as 1/k and is still 2.29 points at the full 100 templates. Right: the rank correlation between the model ordering at that budget and the ordering on all 100 templates, which reaches 0.96 by 10 symbolic templates. Rankings settle long before the scores do, and the symbolic split is more precise than the original at every budget.
Rank correlation
SE (accuracy points)
Templates
Original
Symbolic
Original
Symbolic
10
0.914
0.964
10.60
7.26
20
0.951
0.977
7.49
5.13
50
0.983
0.991
4.74
3.24
75
0.993
0.997
3.87
2.65
100
1.000
1.000
3.35
2.29
Appendix
Table 2: Effect of the template budget on model ranking and on the precision of a single (model, language) score. Rank correlation is against the full 100-template ordering and reaches 1 there by construction; the standard errors do not.
β^
SE
z
p
Language features
Resource level
0.773
0.074
10.41
<2×10−16
Typological distance
−0.253
0.074
−3.41
0.0006
Relative fertility (within)
−0.041
0.032
−1.30
0.195
Model features
Model size
1.768
0.133
13.34
<2×10−16
Appendix
Table 3: Fitted estimates of Equation 1 . Fixed effects in log-odds per standard deviation of the predictor; random effects ordered by magnitude, with the number of levels of each grouping factor.
Term
Raw quantity
Mean
SD
Resource level
log10 Common Crawl pages
7.396
0.862
Typological distance
URIEL syntax_knn cosine distance
0.219
0.139
Relative fertility
within-language deviation
0.000
0.472
Model size
log2 parameters (billions)
2.791
1.865
Appendix
Table 4: Standardisation constants
Figure 10: Symbolic penalty : Mean penalty across languages along with a 95% CI.
Figure 11: Original set accuracy vs synthetic set accuracy : Each dot denote a (model, language) pair. Accuracy is the mean accuracy on either the synthetic or original set.
Figure 12: Transfer gap with reasoning enabled vs disabled across the Qwen models.
Figure 13: Reasoning on/off at a given compute budget : We include only models that allow enabling or disabling reasoning.
Figure 14: Forecasting model performance on a given language, showing (model, language) pairs grouped by language.
k
Model in language
Language mean
0
5.98[5.22,6.86]
3.17[2.32,4.06]
1
8.38[7.15,9.69]
6.66[5.05,8.41]
2
7.02[6.26,7.94]
4.35[3.26,5.38]
5
5.63[4.88,6.48]
3.29[2.26,4.33]
10
4.19[3.75,4.69]
2.38[1.78,3.08]
20
2.73[2.50,2.96]
1.44[1.11,1.78]
Appendix
Table 5: Effect of including a few templates when predicting on an otherwise held-out language. Mean absolute error in accuracy points over three draws per language, with 95% bootstrap intervals over the 15 languages.
Figure 15: What does it cost to be a low-resource language in parameters?
Figure 16: Feature design space : We cover both high and low resource languages at various typological distances from English.
Operand reads
Intermediates
Final answer
Resource level
0.299∗∗∗
0.420∗∗∗
0.712∗∗∗
Typological distance
−0.091∗
−0.172∗∗∗
−0.288∗∗∗
Relative fertility (within)
−0.035
−0.067∗∗∗
−0.071∗
Model size
0.493∗∗∗
0.934∗∗∗
1.881∗∗∗
Resource × size
−0.070∗∗∗
−0.120∗∗∗
−0.246∗∗∗
sd(language)
0.136
0.163
0.272
Appendix
Table 6: Fixed effects in log-odds per SD and variance components at each stage. ∗p<0.05 , ∗∗p<0.01 , ∗∗∗p<0.001 .
Figure 17: Fitted effects at each stage.
Full
English <95%
English <90%
Cells ≤90%
Models
46
29
20
45
Cells above 90%
171
12
0
0
Resource × size
−0.265∗∗∗
−0.205∗∗∗
−0.188∗∗∗
−0.245∗∗∗
Distance × size
0.024
0.005
−0.012
0.017
Resource level
0.773∗∗∗
0.941∗∗∗
1.030∗∗∗
0.853∗∗∗
Typological distance
−0.253∗∗∗
−0.277∗∗∗
−0.311∗∗∗
−0.259∗∗
Appendix
Table 7: The size interaction under progressively stricter removal of saturated cells. ∗p<0.05 , ∗∗p<0.01 , ∗∗∗p<0.001 .
Model ID
Rev
Model ID
Rev
Qwen2.5 Instruct
Qwen/Qwen2.5-0.5B-Instruct
7ae5576
Qwen/Qwen2.5-1.5B-Instruct
989aa79
Qwen/Qwen2.5-3B-Instruct
aa8e725
Qwen/Qwen2.5-7B-Instruct
a09a354
Qwen/Qwen2.5-14B-Instruct
cf98f3b
Qwen/Qwen2.5-32B-Instruct
5ede1c9
Qwen/Qwen2.5-72B-Instruct
495f393
Qwen3
Appendix
Table 8: Models repository identifiers on Hugging Face and revision commit hashes used in the evaluation.
Language
Human verified
EU
Ethno- logue
Language
Human verified
EU
Ethno- logue
Afrikaans
✓
Lithuanian
✓
Amharic
✓
Magahi
✓
Arabic
✓
Maithili
✓
Assamese
✓
Malayalam
✓
Bajjika
✓
Maltese
✓
Bamanankan
✓
Marathi
✓
Appendix
Table 9: Overview of included languages: We include a total of 102 languages. Besides the human verified languages we fill in languages from the Ethnologue top 200, prioritizing one main language in cases of a family of multiple languages and full coverage of the official European languages. For example, we select only standard Arabic and not Egyptian Arabic or Algerian Arabic. Languages with check marks in parenthesis are in the process of validation and will to added to the analysis later.
Language
Human review
Comp. validated
Error analysis
Initial translation model
Source lang.
ara
a native speaker
✓
✓
anthropic/claude-opus-4-8
eng
dan
three native speakers
✓
✓
gpt-5.4
eng
deu
two native speakers
✓
✓
gpt-5.4
dan
eng
a native speaker
✓
✓
est
a native speaker
✓
✓
gpt-5.4-nano
eng_metric
fra
a native speaker
✓
✓
anthropic/claude-opus-4-8
dan
Appendix
Table 10: Human-validated languages and reviews in progress, alongside the English originals. Computational validation is enforced by CI. Error analysis is marked only when complete for all active templates. Note that the numbers does not match with the number of translators as some translated more than one language (bilinguals) and some are in the process evaluation.
Figure 18: Distribution of model performances on the machine-translated and verified Icelandic subsets.
Figure 19: Distribution of model performances on the English and metric-converted English subsets.
Model
Reas.
Params
Orig.
Symb.
English
Non-Eng.
Gap
Worst
Qwen3.5-27B
on
27.0
96.3
94.9
98.0
94.7
3.3
90.2 (mar)
Qwen3.5-27B
off
27.0
95.4
93.7
97.9
93.5
4.4
86.9 (mar)
Qwen3-32B
on
32.0
93.9
93.1
97.5
92.8
4.7
89.5 (isl)
Qwen3-14B
on
14.0
93.0
91.7
96.2
91.4
4.8
88.4 (isl)
Qwen3.5-9B
on
9.0
93.8
91.7
98.0
91.2
6.7
73.6 (mar)
Qwen2.5-72B-Instruct
off
72.0
93.6
91.2
97.5
90.7
6.8
80.2 (mar)
Appendix
Table 11: Exact-answer accuracy on Multilingual GSM-Symbolic, averaged over the 15 benchmark languages and ordered by symbolic accuracy. Gap is English minus the mean of the other 14 languages. Worst is the lowest-scoring language on the symbolic split. Per-language scores carry roughly ±2.3 points of sampling error on the symbolic split and ±3.4 on the original; the 15-language averages about ±1.4 (Appendix A.4 ). ∗ Reported here but excluded from the model in Equation 1 .
Model
Language
Original Accuracy
Synthetic Accuracy
Model
Language
Original Accuracy
Synthetic Accuracy
Qwen2.5-0.5B-Instruct
Chinese
16.0% [10.1-24.4]
8.3% [7.2-9.6]
Qwen3.5-4B
Chinese
95.0% [88.8-97.8]
94.6% [93.5-95.5]
Hindi
1.0% [0.2-5.4]
0.6% [0.3-1.0]
Hindi
85.0% [76.7-90.7]
83.4% [81.7-84.9]
English
42.0% [32.8-51.8]
38.6% [36.5-40.8]
English
98.0% [93.0-99.4]
97.5% [96.7-98.1]
English metric
41.0% [31.9-50.8]
38.7% [36.6-40.9]
English metric
97.0% [91.5-99.0]
96.7% [95.8-97.4]
Arabic
3.0% [1.0-8.5]
1.1% [0.7-1.7]
Arabic
92.0% [85.0-95.9]
92.4% [91.2-93.5]
Japanese
2.2% [1.2-4.2]
1.7% [1.4-2.0]
Japanese
95.0% [91.0-97.3]
91.5% [90.6-92.3]
Appendix
Table 12: Original and synthetic exact-answer accuracy by model and language for all 48 models. Bracketed values are 95% confidence intervals (CI), computed using Wilson score intervals over scored examples within each split.
Cross-lingual transfer is a model's ability to generalize capabilities from well-represented source languages to under-represented target languages. Existing measures of a model's transfer strength conflate improvements in transfer with general improvements to accuracy in the source language. We advocate for an alternate metric that reliably captures transfer strength called Hardness Adjusted Transfer (HAT) Score, and use it to derive multiple insights on factors influencing transfer strength. Our analysis across twenty diverse language models and three popular mainstream multilingual benchmarks argues that 1) transfer in small models is not broken, 2) we are making slower than expected progress in cross-lingual transfer with model size, and 3) we have made clear progress over time.
Prasoon Bajpai, Eleftheria Briakou, Colin Cherry +2
Google DeepMind · Indian Institute of Technology Bombay
Large language models exhibit substantial performance variation across languages, even when solving semantically equivalent tasks. Existing analyses often treat this phenomenon as an observational disparity caused by differences in pretraining data, tokenization, or benchmark coverage. We study a complementary hypothesis: high-resource languages (HRLs) may more reliably elicit latent computations useful for task-specific (i.e. mathematical) reasoning, while lower-resource languages (LRLs) may under-activate those computations despite expressing the same task. To test this hypothesis, we introduce a mechanistic intervention framework for identifying and transferring task-relevant sparse latent features across languages. Using sparse autoencoders over residual-stream activations, we isolate features enriched in successful HRL task-specific reasoning while filtering out source-language and generic-generation features. We then construct steering directions from these features and inject them during LRL inference. The resulting interventions test whether the selected features are functionally involved in the observed reasoning gap: suppressing them should impair source-language reasoning, while activating them should partially recover target-language reasoning beyond random and non-task controls. Our framework reframes some cross-lingual reasoning gaps as failures of mechanism elicitation rather than capability absence, and offers a causally testable route to feature-mediated transfer without translation, fine-tuning, or changing the user-facing language.
Minju Song, Hyeon Hwang, Junhyun Lee +1
Korea University · Hankuk University of Foreign Studies · Noah’s Farm +1
Large language models have made substantial progress in mathematical reasoning. However, benchmark development for multilingual evaluation has lagged behind English in both difficulty and recency. Recently, GSM-Symbolic showed a strong evidence of high variance when models are evaluated on different instantiations of the same question; however, the evaluation was conducted only in English. In this paper, we introduce MGSM-Pro, an extension of MGSM dataset with GSM-Symbolic approach. Our dataset provides five instantiations per MGSM question by varying names, digits and irrelevant context. Evaluations across nine languages reveal that many low-resource languages suffer large performance drops when tested on digit instantiations different from those in the original test set. We further find that models robustness in HRL setting do not necessarily translate to LRL. Moreover, proprietary models, such as Gemini 2.5 Flash and GPT-4.1 are less robust to digit, whereas Gemini 3.0 Pro is more robust. Among open models, GPT-OSS 120B and DeepSeek v3 show stronger robustness. Based on these findings, we recommend evaluating each problem using at least five digit-varying instantiations to obtain a more robust and realistic assessment of math reasoning.
Tianyi Xu, Kosei Uemura, Alfred Malengo Kondoro +6
McGill University · Mila-Quebec AI Institute · University of Toronto +7