Skeleton-based reasoning prompting is a promising training-free approach for structuring LLM reasoning, but prior work largely assumes an English-centric setting. We propose the Language-Aware Skeleton Exploration Framework (LASEF) to study skeleton-language choice in multilingual mathematical reasoning. Across math benchmarks, model scales, and languages, we show that English skeletons yield a small positive tendency on average, most visible for smaller models and low-resource languages. However, few language-level gains remain significant after correction, and English is not universally optimal. Combining greedy decoding, multi-rollout evaluation, translation ablation, and cross-benchmark validation, we further find three patterns of skeleton-language effects: directionally consistent, evaluation- and benchmark-dependent, and asymmetric negative. These effects cannot be fully explained by generation quality alone. Overall, skeleton language is a context-dependent design variable that requires multi-level exploration. All resources are released at https://github.com/lhsstn/LASEF.
Figures & tables
Figure 1: Overview of the Skeleton-Guided Reasoning Framework. (A) The process begins with an input query q(ℓq) . (B) A structured skeleton s(ℓs) is generated to abstract the problem logic. (C) The final reasoning and answer a^ are produced in the answer language ℓa , conditioned on the query and skeleton. (D) We define the language configuration space as (ℓq,ℓs,ℓa) and compare four primary strategies: standard CoT and skeleton-guided approaches, applied in both the target language ( ℓt ) and English ( en ). (E) Effective configurations are identified through training-free benchmark analysis.
Method
Language
MGSM
MATH-500
PolyMath
( ℓq,ℓs,ℓa )
zh
es
ko
th
sw
AVG.
zh
es
ko
th
sw
AVG.
zh
es
ko
th
sw
AVG.
Qwen2.5-7B-Instruct
CoT- ℓt
ℓt,−,ℓt
79.9
80.7
69.7
80.2
16.0
65.3
56.4
57.5
47.8
49.8
14.2
45.1
30.8
31.0
26.4
26.1
5.1
23.9
+ Skeleton
ℓt,en,ℓt
80.7
83.8
71.0
79.0
22.9
67.5
55.4
57.7
49.3
51.9
20.4
46.9
30.6
32.1
28.0
27.6
7.4
25.1
CoT- en†
en,−,ℓt
82.3
84.3
74.5
78.1
63.1
76.5
56.0
57.4
51.6
49.5
24.4
47.8
29.7
32.1
27.5
29.8
20.3
27.9
+ Skeleton
en,en,ℓt
82.3
85.9
73.2
78.9
71.7
78.4
57.4
58.2
52.3
53.8
29.8
50.3
32.0
34.4
27.7
30.4
19.9
28.9
Table 1: Results on MGSM, MATH-500, and PolyMath. (ℓq,ℓs,ℓa) denote the query, skeleton, and answer languages, respectively. CoT- en† indicates queries translated into English using Google Translate, and the skeleton language is fixed to English unless otherwise noted. ‘AVG.’ denotes the average over five languages (zh, es, ko, th, sw). Greedy decoding throughout.
Language
ℓt,en,ℓt
en†,en,ℓt
CoT- lt
+ Skeleton
CoT- en†
+ Skeleton
Qwen2.5-7B-Instruct
kk
40.08
51.42 ( +11.34 )
62.82
72.22 ( +9.40 )
ky
33.33
37.86 ( +4.53 )
61.92
65.69 ( +3.77 )
mn
27.16
33.33 ( +6.17 )
66.96
64.78 ( -2.18 )
ug
27.13
37.65 ( +10.52 )
57.14
60.37 ( +3.23 )
Table 2: Performance comparison on 19 low-resource languages (MGSM) using Qwen2.5-7B with English Skeleton. CoT- en† : English-translated queries via Google Translate. Parentheses show Δ gains.
Figure 2: Performance differences (Non-English − English) on MGSM under CoT- ℓt with Qwen2.5-7B . The English skeleton is replaced with five Non-English skeletons (ZH, ES, RU, KO, TH); the same skeleton set is used in both panels. (a) Greedy decoding (single rollout). (b) Multi-rollout sampling ( T=0.7 , N=5 ). The corresponding English-skeleton reference values (English skeleton vs. no skeleton per language) are reported in Tab. 2 .
Figure 3: Translation ablation results. Each point is a (target, skeleton) pair, with Δ from LLM-generated skeletons on the x-axis and translated skeletons on the y-axis. Points near the diagonal indicate effects robust to skeleton quality.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Table 3: The prompt instructs the model to produce a concise, high-level outline of the reasoning process, while omitting detailed computations and final answers.
Table 4: List of the 19 low-resource target languages included in the MGSM-LowResource dataset.
Target Lang
Acc (%)
Δ
Target Lang
Acc (%)
Δ
English (Baseline)
90.00
–
Quechua (qu)
89.60
-0.40
Tamil (ta)
93.60
+3.60
Guarani (gn)
90.00
0.00
Gujarati (gu)
91.60
+1.60
Basque (eu)
90.00
0.00
Kannada (kn)
91.20
+1.20
Amharic (am)
90.00
0.00
Khmer (km)
91.20
+1.20
Javanese (jv)
90.00
0.00
Sinhala (si)
91.20
+1.20
Uyghur (ug)
89.60
-0.40
Appendix
Table 5: Reasoning accuracy of Qwen2.5-7B on MGSM tasks using Round-trip Translation (English → Pivot → English). We used the target languages as pivots to validate the semantic consistency of the translated dataset. Δ denotes the performance difference relative to the English baseline.
Statistic
Value
Total Target Languages
34
Samples per Language
250
Total Samples
8,500
Gemini Selection Ratio
62.1%
GPT Selection Ratio
37.9%
Avg. BT Similarity Score
0.96
Appendix
Table 6: Final statistics of the MGSM-LowResource dataset.
Code
Language
Family
Morph.
Order
Target languages
kk
Kazakh
Turkic
Agg
SOV
ky
Kyrgyz
Turkic
Agg
SOV
mn
Mongolian
Mongolic
Agg
SOV
ug
Uyghur
Turkic
Agg
SOV
hy
Armenian
IE (Armenian)
Agg †
SOV
Appendix
Table 7: Typological features of the 19 target and six skeleton languages. Morph. : dominant morphological type (Ana: analytic/isolating; Agg: agglutinative; Fus: fusional). Order : basic constituent order. Scripts are listed in Tab. 4 . † Armenian exhibits agglutinative nominal morphology. These features derive surface affinity and structural alignment as defined above.
Figure 4: Average token length of skeletons across query languages.
Figure 5: Average generated token counts of the Qwen-series models.
Figure 6: Quality evaluation results for five models in the PolyMath
Figure 7: Skeleton quality heatmap across language pairs. English skeletons yield highest quality (7.57), while Thai performs worst (5.65).
Language
Method
Problem
Conceptual
Reasoning
Calculation
Output Gen.
kk
CoT- ℓt
100
61
129
33
116
+SKELETON
91
54
105
18
135
ug
CoT- ℓt
106
74
156
57
95
+SKELETON
100
60
134
33
135
mn
CoT- ℓt
107
78
157
47
88
+SKELETON
108
62
142
41
115
Appendix
Table 8: Comparison of CoT- ℓt and +SKELETON Results Across Different Reasoning Stages
Method
MGSM
MATH-500
PolyMath
zh
es
ko
th
sw
te
AVG.
zh
es
ko
th
sw
te
AVG.
zh
es
ko
th
sw
te
AVG.
Llama-3.1-8B-Instruct
CoT (ℓq=target)
71.5
74.7
58.6
65.5
61.7
56.2
64.7
33.2
37.7
29.8
31.0
31.9
25.1
31.5
23.6
24.9
26.2
19.3
27.6
16.2
23.0
+ Skeleton
68.3
75.5
61.0
65.1
63.8
57.4
65.2
34.7
35.6
34.0
32.0
34.2
25.9
32.7
24.1
26.3
24.4
21.1
29.1
16.8
23.6
CoT (ℓq=EN†)
69.3
78.3
64.9
69.1
68.3
63.9
69.0
33.8
41.6
29.7
33.3
32.5
26.3
32.9
22.4
27.6
24.4
21.6
32.4
17.3
24.3
+ Skeleton
69.8
73.9
63.7
70.3
66.7
62.7
67.9
37.7
37.3
32.9
31.6
38.8
29.6
34.7
26.3
28.0
27.3
22.4
33.9
17.3
25.9
Appendix
Table 9: Multilingual mathematical reasoning results for the Llama 3.1 series (8B, 70B). The experiments use the MGSM, MATH-500, and PolyMath benchmarks. ℓq=target indicates that questions are presented in the target language, while ℓq=EN† denotes the use of English-translated questions. The skeleton language is fixed to English in all cases ( ℓs=EN ). ‘AVG.’ denotes the average over the six languages.
Model
MGSM
MATH-500
PolyMath
AVG.
AVG. ∗
Ministral-3B
+1.34
+1.79
+0.74
+1.29
+0.38
Ministral-14B
−1.10
+0.35
+1.08
+0.11
−0.12
Appendix
Table 10: English-skeleton effect Δ (pp) for the Ministral family under the CoT- ℓt setting, averaged over the five languages of Tab. 1 (zh, es, ko, th, sw) per benchmark. ‘AVG.’ denotes the average over all five languages; AVG. ∗ excludes Swahili (sw), which retains relatively few valid samples.
Model
CoT- ℓt
CoT- en†
Base
+ Skeleton
Base
+ Skeleton
Qwen2.5-7B
4.75
4.91 ( +0.16 )
5.12
5.82 ( +0.70 )
Qwen2.5-14B
6.62
8.40 ( +1.78 )
7.13
7.56 ( +0.43 )
Qwen2.5-72B
9.27
9.97 ( +0.70 )
8.92
8.96 ( +0.04 )
Appendix
Table 11: Skeleton gains on the two hardest PolyMath subsets (Top + High; 250 of 500 problems), as Exact-Match accuracy (%) averaged over five languages (zh, es, ko, th, sw). Parentheses show the gain ( Δ ) from adding the English skeleton. Compared with the full-set gains (Tab. 1 ), improvements shrink on the harder subset, consistent with the diminishing effect of skeletons as difficulty increases.
Target
Skeleton
Δ (%)
p
Positive effects (Non-English skeleton > English skeleton)
Table 12: Pairs that reached nominal significance in McNemar’s test before correction for multiple comparisons ( p<0.05 ). Δ denotes the accuracy difference between the non-English skeleton and the English skeleton.
Target
Skel.
MGSM Δ
MSVAMP Δ
Outcome
hy
ko
+5.17
+2.75
Retained
lo
zh
+3.63
+2.63
Retained
lo
th
+2.81
+2.42
Retained
eu
ko
+6.32
0.00
Disappeared
mt
es
+6.61
+2.00
Attenuated
mt
th
+6.90
+0.43
Attenuated
Appendix
Table 13: Cross-benchmark transfer of skeleton-language effects from MGSM to MSVAMP ( Qwen2.5-7B , CoT- ℓt ). Δ : accuracy difference (pp) between the non-English and the English skeleton. Pairs were fixed using only the MGSM results before inspecting MSVAMP. Across all 35 target–skeleton cells, the sign is preserved in 23/35 (66%).
Method
MATH-500 (te)
PolyMath (te)
MGSM (te)
Qwen2.5-7B-Instruct
CoT- ℓt
29.8
12.2
30.40
+SKELETON
27.8
12.4
33.60
CoT- en†
35.7
17.3
60.40
+SKELETON
33.9
17.7
54.00
Qwen2.5-14B-Instruct
Appendix
Table 14: Telugu (te) Accuracy (%) on MATH-500, PolyMath, and MGSM
Figure 8: Performance differences (Non-English − English) under the CoT- ℓt setting using Qwen2.5-7B across all 34 target languages. While English generally serves as a stable default skeleton language, several target languages exhibit gains with specific Non-English skeletons, indicating that skeleton effectiveness is language-dependent.
Reinforcement learning has proven effective for enhancing multi-step reasoning in large language models (LLMs), yet its benefits have not fully translated to multilingual contexts. Existing methods struggle with a fundamental trade-off: prioritizing input-language consistency severely hampers reasoning quality, while prioritizing reasoning often leads to unintended language drift toward English. We address this challenge with LANG, a novel framework that leverages language-conditioned hints to guide exploration in non-English reasoning tasks. Our method incorporates two key mechanisms to prevent dependency on these hints: a progressive decay schedule that gradually withdraws scaffolding, and a language-adaptive switch that tailors learning horizons to specific language difficulties. Empirical results on challenging multilingual mathematical benchmarks reveal that LANG substantially enhances reasoning performance without compromising language consistency. Moreover, we show that our framework generalizes beyond mathematics, fostering more consistent language alignment across model layers
Yuchun Fan, Bei Li, Peiguang Li +9
NLP Lab, School of Computer Science and Engineering, Northeastern University, Shenyang, China · Meituan Inc. 3NiuTrans Research, Shenyang, China
Large reasoning models (LRMs) achieve strong mathematical reasoning performance in English, but remain much less reliable in many low- and medium-resource languages. This gap is often explained as a failure to understand non-English problem statements. We show that this view is incomplete: even when the problem is given in English, controlling the model's reasoning language can substantially reduce accuracy, suggesting that language also affects reasoning execution itself. To study this effect, we introduce DATG, a Directed Acyclic Trace Graph framework that maps reasoning traces to language-independent mathematical anchors and dependencies. This allows us to align target-language traces with reference DAGs and measure whether they cover required mathematical nodes, respect dependency edges, and avoid harmful mathematical actions. Experiments on the Qwen3 series across 12 languages show that non-English reasoning often suffers from reduced anchor coverage and weaker dependency fidelity, especially in low-resource languages. Motivated by this diagnosis, we propose Loop-Retry and Formula-Retry, two simple test-time controls targeting DATG-exposed failure modes, and show that they consistently improve target-language reasoning performance in low-resource languages.
Languages encode distinct abstractions and inductive priors, yet most large language models (LLMs) overlook this diversity by reasoning in a single dominant language. In this work, we introduce x1, a family of reasoning models that can adaptively reason in an advantageous language on a per-instance basis. To isolate the effect of reasoning-language choice, x1 is constructed without expanding the model's knowledge boundaries and is trained by contrasting linguistically distinct reasoning trajectories for the same input. Our extensive experiments demonstrate the benefits of adaptive multilingual reasoning across multilingual mathematical reasoning and culturally grounded tasks. Moreover, our results challenge a simplistic view of scaling laws: while scaling reduces cross-lingual disparities in procedural domains such as math reasoning, it does not eliminate the advantages of culture-associated languages in culturally grounded tasks, as we empirically show that such reasoning enables more efficient and accurate cultural knowledge recall. Overall, our findings establish language choice as a functional component of reasoning, with implications for building more generalist and globally competent reasoning models.
Yangfan Ye, Xiaocheng Feng, Xiachong Feng +8
Harbin Institute of Technology · Peng Cheng Laboratory · The University of Hong Kong +1