Multilingual GSM-Symbolic: What determines capability transfer across languages?
Authors: Kenneth Enevoldsen, Riley Herchert, Sofie Mosegaard, Dan Saattrup Smart, Simon Enni, Isaac Chung, Sofie Bruun, Ayush Sunil Munot, +17 more
Organizations: Aarhus University · Danish Foundation Models · University of Alabama · Alexandra Institute · Indian Institute of Technology Kharagpur · The University of Tokyo · IT University of Copenhagen · Massachusetts General Hospital · Bocconi University · University of Southern Denmark · University of Iceland · University of the Faroe Islands · Zendesk · University of Copenhagen · Indian Institute of Technology Madras · National Library of Sweden
We understand little about how capabilities acquired in one language carry over to another, or what governs this transfer: evaluations rely on incomparable, saturation-prone datasets and rarely examine its determinants jointly. Identifying what predicts transfer would let us avoid exhaustive evaluation across all language pairs and let developers target the factors that limit performance in low-resource languages. To evaluate cross-lingual capability transfer, we introduce Multilingual GSM-Symbolic, an extensible multilingual mathematical dataset covering 30,000 item-matched question-answer pairs and spanning 15 languages. It utilises symbolic templates to prevent overfitting and ensure generalisation by allowing generation of millions of high-quality variations from a single sample. Using Multilingual GSM-Symbolic, we quantify the largest determinants of capability as model size (β=1.77), language resource level (β=0.77), reasoning (β=0.67) and typological distance (β=−0.25). This joint estimation allows these determinants to be expressed in terms of one another: a 32B model evaluated in Marathi performs like a 10B model in English. Our findings have important implications for model developers, showing that model size and reasoning narrow the performance gap between low- and high-resource languages (β=−0.27 and β=−0.20, respectively), while similar levers have little or no effect on typologically distant languages. Overall, our analysis framework explains 92% of between-language variation, but only 23% of the model-by-language variation, and predicts a model's performance on an unseen language within 6.0pp (r=.96). Incorporating measurements from just 10 templates in the target language reduces this to 4.19pp, enabling reasonable estimates of performance with little or no downstream dataset.
Figures & tables
Figure 1: Overview of Multilingual GSM-Symbolic : Consisting of 100 samples in 15 languages, multilingual GSM-Symbolic consists of 1500 verified and localized templates. We sample 20 per template for a total of 30,000 and show an example evaluation of Qwen2.5-7B across English, Danish and Marathi. Showing both the language gap between Marathi and English and synthetic gap between the original samples and their template samples distributions.
Validation
Dataset
Languages
Symbolic
Code
Extensible
Error Analysis
Translation
Localization
Validity
Multilingual GSM-symbolic (ours)
15 (100)
✓
✓
✓
✓
✓
✓
✓
Previous work
MGSM ( Shi et al., 2023 )
11
✓
(✓)
GSM-Symbolic ( Mirzadeh et al., 2025 )
1
✓
GSM8K-Platinum ( Vendrow et al., 2025 )
1
✓
Table 1: Comparison of GSM-style mathematical reasoning benchmarks . The parenthesis refer to languages that have not been human validated yet. For more see A.2 Our dataset combines multilingual templates, controllable variation and extensibility with extensive validation. Error analysis refer to the process of running a model and correcting ambiguous or incorrect questions. Extensible as defined by Enevoldsen et al. (2026) , indicate that there is a process for adding to, extending or correcting the dataset. Validity denote automated tests that ensures the quality of the generated examples. While MGSM and MGSM-Pro localizes languages they do not localize names, currency and metrics even though imperial units are rarely used in a non-English contexts.
Figure 3
Figure 4: Levers : Model prediction for four interaction terms denoting two levers, showing the effect of reasoning and size across resource level and typological distance. Depending on the language you are working with, the method for improving capability transfer differs.
Figure 5
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6
Figure 9: How much data is needed . Left: the standard error of a single (model, language) accuracy at a given template budget, which falls as 1/k and is still 2.29 points at the full 100 templates. Right: the rank correlation between the model ordering at that budget and the ordering on all 100 templates, which reaches 0.96 by 10 symbolic templates. Rankings settle long before the scores do, and the symbolic split is more precise than the original at every budget.
Rank correlation
SE (accuracy points)
Templates
Original
Symbolic
Original
Symbolic
10
0.914
0.964
10.60
7.26
20
0.951
0.977
7.49
5.13
50
0.983
0.991
4.74
3.24
75
0.993
0.997
3.87
2.65
100
1.000
1.000
3.35
2.29
Appendix
Table 2: Effect of the template budget on model ranking and on the precision of a single (model, language) score. Rank correlation is against the full 100-template ordering and reaches 1 there by construction; the standard errors do not.
β^
SE
z
p
Language features
Resource level
0.773
0.074
10.41
<2×10−16
Typological distance
−0.253
0.074
−3.41
0.0006
Relative fertility (within)
−0.041
0.032
−1.30
0.195
Model features
Model size
1.768
0.133
13.34
<2×10−16
Appendix
Table 3: Fitted estimates of Equation 1 . Fixed effects in log-odds per standard deviation of the predictor; random effects ordered by magnitude, with the number of levels of each grouping factor.
Term
Raw quantity
Mean
SD
Resource level
log10 Common Crawl pages
7.396
0.862
Typological distance
URIEL syntax_knn cosine distance
0.219
0.139
Relative fertility
within-language deviation
0.000
0.472
Model size
log2 parameters (billions)
2.791
1.865
Appendix
Table 4: Standardisation constants
Figure 10: Symbolic penalty : Mean penalty across languages along with a 95% CI.
Figure 11: Original set accuracy vs synthetic set accuracy : Each dot denote a (model, language) pair. Accuracy is the mean accuracy on either the synthetic or original set.
Figure 12: Transfer gap with reasoning enabled vs disabled across the Qwen models.
Figure 13: Reasoning on/off at a given compute budget : We include only models that allow enabling or disabling reasoning.
Figure 14: Forecasting model performance on a given language, showing (model, language) pairs grouped by language.
k
Model in language
Language mean
0
5.98[5.22,6.86]
3.17[2.32,4.06]
1
8.38[7.15,9.69]
6.66[5.05,8.41]
2
7.02[6.26,7.94]
4.35[3.26,5.38]
5
5.63[4.88,6.48]
3.29[2.26,4.33]
10
4.19[3.75,4.69]
2.38[1.78,3.08]
20
2.73[2.50,2.96]
1.44[1.11,1.78]
Appendix
Table 5: Effect of including a few templates when predicting on an otherwise held-out language. Mean absolute error in accuracy points over three draws per language, with 95% bootstrap intervals over the 15 languages.
Figure 15: What does it cost to be a low-resource language in parameters?
Figure 16: Feature design space : We cover both high and low resource languages at various typological distances from English.
Operand reads
Intermediates
Final answer
Resource level
0.299∗∗∗
0.420∗∗∗
0.712∗∗∗
Typological distance
−0.091∗
−0.172∗∗∗
−0.288∗∗∗
Relative fertility (within)
−0.035
−0.067∗∗∗
−0.071∗
Model size
0.493∗∗∗
0.934∗∗∗
1.881∗∗∗
Resource × size
−0.070∗∗∗
−0.120∗∗∗
−0.246∗∗∗
sd(language)
0.136
0.163
0.272
Appendix
Table 6: Fixed effects in log-odds per SD and variance components at each stage. ∗p<0.05 , ∗∗p<0.01 , ∗∗∗p<0.001 .
Figure 17: Fitted effects at each stage.
Full
English <95%
English <90%
Cells ≤90%
Models
46
29
20
45
Cells above 90%
171
12
0
0
Resource × size
−0.265∗∗∗
−0.205∗∗∗
−0.188∗∗∗
−0.245∗∗∗
Distance × size
0.024
0.005
−0.012
0.017
Resource level
0.773∗∗∗
0.941∗∗∗
1.030∗∗∗
0.853∗∗∗
Typological distance
−0.253∗∗∗
−0.277∗∗∗
−0.311∗∗∗
−0.259∗∗
Appendix
Table 7: The size interaction under progressively stricter removal of saturated cells. ∗p<0.05 , ∗∗p<0.01 , ∗∗∗p<0.001 .
Model ID
Rev
Model ID
Rev
Qwen2.5 Instruct
Qwen/Qwen2.5-0.5B-Instruct
7ae5576
Qwen/Qwen2.5-1.5B-Instruct
989aa79
Qwen/Qwen2.5-3B-Instruct
aa8e725
Qwen/Qwen2.5-7B-Instruct
a09a354
Qwen/Qwen2.5-14B-Instruct
cf98f3b
Qwen/Qwen2.5-32B-Instruct
5ede1c9
Qwen/Qwen2.5-72B-Instruct
495f393
Qwen3
Appendix
Table 8: Models repository identifiers on Hugging Face and revision commit hashes used in the evaluation.
Language
Human verified
EU
Ethno- logue
Language
Human verified
EU
Ethno- logue
Afrikaans
✓
Lithuanian
✓
Amharic
✓
Magahi
✓
Arabic
✓
Maithili
✓
Assamese
✓
Malayalam
✓
Bajjika
✓
Maltese
✓
Bamanankan
✓
Marathi
✓
Appendix
Table 9: Overview of included languages: We include a total of 102 languages. Besides the human verified languages we fill in languages from the Ethnologue top 200, prioritizing one main language in cases of a family of multiple languages and full coverage of the official European languages. For example, we select only standard Arabic and not Egyptian Arabic or Algerian Arabic. Languages with check marks in parenthesis are in the process of validation and will to added to the analysis later.
Language
Human review
Comp. validated
Error analysis
Initial translation model
Source lang.
ara
a native speaker
✓
✓
anthropic/claude-opus-4-8
eng
dan
three native speakers
✓
✓
gpt-5.4
eng
deu
two native speakers
✓
✓
gpt-5.4
dan
eng
a native speaker
✓
✓
est
a native speaker
✓
✓
gpt-5.4-nano
eng_metric
fra
a native speaker
✓
✓
anthropic/claude-opus-4-8
dan
Appendix
Table 10: Human-validated languages and reviews in progress, alongside the English originals. Computational validation is enforced by CI. Error analysis is marked only when complete for all active templates. Note that the numbers does not match with the number of translators as some translated more than one language (bilinguals) and some are in the process evaluation.
Figure 18: Distribution of model performances on the machine-translated and verified Icelandic subsets.
Figure 19: Distribution of model performances on the English and metric-converted English subsets.
Model
Reas.
Params
Orig.
Symb.
English
Non-Eng.
Gap
Worst
Qwen3.5-27B
on
27.0
96.3
94.9
98.0
94.7
3.3
90.2 (mar)
Qwen3.5-27B
off
27.0
95.4
93.7
97.9
93.5
4.4
86.9 (mar)
Qwen3-32B
on
32.0
93.9
93.1
97.5
92.8
4.7
89.5 (isl)
Qwen3-14B
on
14.0
93.0
91.7
96.2
91.4
4.8
88.4 (isl)
Qwen3.5-9B
on
9.0
93.8
91.7
98.0
91.2
6.7
73.6 (mar)
Qwen2.5-72B-Instruct
off
72.0
93.6
91.2
97.5
90.7
6.8
80.2 (mar)
Appendix
Table 11: Exact-answer accuracy on Multilingual GSM-Symbolic, averaged over the 15 benchmark languages and ordered by symbolic accuracy. Gap is English minus the mean of the other 14 languages. Worst is the lowest-scoring language on the symbolic split. Per-language scores carry roughly ±2.3 points of sampling error on the symbolic split and ±3.4 on the original; the 15-language averages about ±1.4 (Appendix A.4 ). ∗ Reported here but excluded from the model in Equation 1 .
Model
Language
Original Accuracy
Synthetic Accuracy
Model
Language
Original Accuracy
Synthetic Accuracy
Qwen2.5-0.5B-Instruct
Chinese
16.0% [10.1-24.4]
8.3% [7.2-9.6]
Qwen3.5-4B
Chinese
95.0% [88.8-97.8]
94.6% [93.5-95.5]
Hindi
1.0% [0.2-5.4]
0.6% [0.3-1.0]
Hindi
85.0% [76.7-90.7]
83.4% [81.7-84.9]
English
42.0% [32.8-51.8]
38.6% [36.5-40.8]
English
98.0% [93.0-99.4]
97.5% [96.7-98.1]
English metric
41.0% [31.9-50.8]
38.7% [36.6-40.9]
English metric
97.0% [91.5-99.0]
96.7% [95.8-97.4]
Arabic
3.0% [1.0-8.5]
1.1% [0.7-1.7]
Arabic
92.0% [85.0-95.9]
92.4% [91.2-93.5]
Japanese
2.2% [1.2-4.2]
1.7% [1.4-2.0]
Japanese
95.0% [91.0-97.3]
91.5% [90.6-92.3]
Appendix
Table 12: Original and synthetic exact-answer accuracy by model and language for all 48 models. Bracketed values are 95% confidence intervals (CI), computed using Wilson score intervals over scored examples within each split.