Pre-generation success probes estimate response correctness from a language model's hidden activations before decoding, enabling cost-aware routing. While prior work has demonstrated their utility primarily on English inputs, we study their reliability across languages along three dimensions: (1) whether they preserve the ranking of likely successes and failures (DISCRIMINATION); (2) whether they retain probabilities that match observed success frequencies (CALIBRATION); and (3) whether they produce scores comparable enough across candidate models for cost-aware multilingual routing (UTILITY). Using 3,000 MATH problems in 10 languages and 8 open-weight model configurations, we compare cross-lingual transfer from English-trained probes and equal-budget pooled multilingual probes. English-trained probes retain useful cross-lingual discrimination but become less well calibrated after transfer. Pooled multilingual supervision improves both properties and yields more reliable estimates of success. In routing experiments, the pooled router achieves a 0.7% higher test success rate while reducing modeled cost by 13.0% relative to always selecting the model with the highest average success. These results show that multilingual routing requires success estimates that remain well calibrated and comparable across languages and models.
Figures & tables
Figure 1: Multilingual routing with pre-generation success probes. Test rollout success versus modeled cost saving relative to always selecting gpt-oss-high , on the balanced ten-language test set. Each point corresponds to a different value of the routing trade-off coefficient λ . Curves compare English-trained, recalibrated English-trained, and equal-budget pooled multilingual probes used for routing.
Language
Qwen3-4B
English
0.716
Chinese
0.630
Italian
0.621
French
0.613
Russian
0.611
Arabic
0.537
Table 1: Empirical success rates of Qwen3-4B across languages. Each value is the average empirical success score over the test split, where each problem-level score is computed from five sampled generations.
Figure 2: Per-language reliability diagrams for Qwen3-4B at its shared reporting layer. The English-trained probe is transferred unchanged to each target language, whereas the pooled probe uses the same total training budget distributed equally across the ten languages. The diagonal denotes perfect calibration.
Evaluation regime
Rollout AUROC ↑
Rollout Brier ↓
ECE ↓
English
0.729
0.213
0.125
Pooled
0.741
0.190
0.043
Oracle
0.758
0.183
0.031
Table 2: Self-probe quality for Qwen3-4B at the shared reporting layer (L23). Values are macro-averages over the nine non-English test languages. AUROC treats the five rollouts as weighted binary outcomes.
Predictor
AUROC ↑
Brier ↓
NLL ↓
ECE ↓
English raw
0.729
0.213
0.624
0.125
English + cal.
0.729
0.208
0.614
0.099
Pooled raw
0.741
0.190
0.573
0.043
Table 3: Prediction quality of the Qwen3-4B ETD routing head at the shared encoder’s reporting layer (L23), macro-averaged over the nine non-English test languages. Bold marks the best value for each metric.
Router
Val. choice
λ
Test success [95% CI]
Saving [95% CI]
English raw
Max. success
0.00
80.8 [78.8,82.8]
33.7 [31.8,35.5]
Anchor match
–
Not attained on validation
English + cal.
Max. success
0.00
85.0 [83.1,86.8]
10.8 [9.0,12.6]
Anchor match
–
Not attained on validation
Pooled raw
Max. success
0.00
86.6 [84.8,88.3]
13.0 [10.2,15.9]
Anchor match
0.10
85.5 [83.6,87.2]
22.9 [19.8,26.3]
Table 4: Validation-selected routing operating points. Anchor matching allows at most 0.1 percentage point validation success loss, then minimizes cost; unattained rows mean no fixed-grid point met that constraint. Brackets are paired bootstrap intervals over base problem IDs; choices are frozen before test.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Language
DeepL
Qwen3.5-27B
Arabic
0.7758
0.7677
Chinese
0.8050
0.8058
French
0.8293
0.8286
Italian
0.8295
0.8278
Russian
0.8139
0.8121
Swahili
0.7913
0.7834
Appendix
Table 5: Mean COMET-Kiwi scores over 3,000 MATH problems per target language and translation system. Higher is better; bold indicates the stronger system within each row.
Figure 3: Ridge coefficient selected at every relative representation depth. The top row shows the same-model probe and the bottom row shows ETD target-specific heads; columns separate English and pooled supervision. In each cell, the selected α minimizes rollout-level validation Brier score over {10,102,103,104,105} . Each row denotes the model supplying the representation for the same-model probe and the predicted target model for ETD; each cell denotes one layer. ETD heads use the shared Qwen3-4B representation.
Probe
peasy
phard
AUROC
ECE
Brier
NLL
Constant
0.50
0.50
0.50
0.00
0.250
0.693
Overconfident
1.00
0.00
0.75
0.25
0.250
∞
Well specified
0.75
0.25
0.75
0.00
0.188
0.562
Appendix
Table 6: Expected metrics for three hypothetical probes. The dataset contains equally many easy and hard problems with true success probabilities 0.75 and 0.25 , respectively. The infinite NLL for the overconfident probe is the theoretical value before numerical clipping.
Model
Input price
Output price
Avg. input
Avg. output
Exp. cost
Relative
($/M tok.)
($/M tok.)
tokens
tokens
($/k gen.)
Exp. cost
Qwen3-0.6B
0.010
0.060
205.12
593.45
0.0377
0.0111
Qwen3-1.7B
0.028
0.170
205.12
2361.62
0.4073
0.1196
Qwen3-4B
0.067
0.400
205.12
2056.09
0.8361
0.2456
gpt-oss-low
0.204
1.224
226.57
946.03
1.2042
0.3537
Qwen3.5-4B
0.067
0.400
186.49
3042.49
1.2294
0.3611
Appendix
Table 7: Proxy prices used in the routing simulation. Average input lengths are computed over the balanced pooled training sample used for routing-cost estimation (210 prompts per language and 2,100 prompts per model), using each model’s native tokenizer and chat template. Average output lengths are computed over the same sample. Relative expected costs are normalized by the average standalone input-plus-output cost of gpt-oss-high .
Shared encoder
Layer
Max. validation success
Prefill / anchor
Qwen3-0.6B
21
88.1%
0.06%
Qwen3-1.7B
19
87.9%
0.17%
Qwen3-4B
23
88.3%
0.40%
Qwen3-8B
26
88.1%
0.80%
Qwen3.5-4B
10
88.2%
0.37%
Qwen3.5-9B
13
88.2%
0.81%
Appendix
Table 8: Validation-only shared-encoder comparison under pooled ETD supervision. For each encoder, the table reports its validation-selected layer and the maximum success attained over the fixed validation routing grid. No test outcomes enter the comparison. Prefill cost is the encoder’s standalone input-only cost as a percentage of the standalone gpt-oss-high anchor cost, using the prices and training-mean input lengths in Appendix D .
Language
Qwen3-0.6B
Qwen3-1.7B
Qwen3-4B
Qwen3-8B
Qwen3.5-4B
Qwen3.5-9B
gpt-oss-low
gpt-oss-high
Mean
English
0.279
0.508
0.716
0.728
0.916
0.935
0.895
0.925
0.738
Chinese
0.220
0.408
0.630
0.638
0.851
0.898
0.754
0.886
0.661
French
0.179
0.388
0.613
0.663
0.848
0.895
0.789
0.911
0.661
Russian
0.170
0.384
0.611
0.661
0.837
0.889
0.761
0.884
0.650
Italian
0.160
0.378
0.621
0.654
0.835
0.880
0.755
0.897
0.648
Arabic
0.134
0.312
0.537
0.600
0.785
0.841
0.749
0.885
0.605
Appendix
Table 9: Empirical success rates of the evaluated models across languages. Each value is the average empirical success score over the test split, where each problem-level score is computed from five sampled generations. Languages are ordered by decreasing average success rate across models. The final row reports the average success rate for each model across languages, and the final column reports the average success rate for each language across models.
Rollout AUROC ↑
Rollout Brier ↓
ECE ↓
Model
Layer
Eng. (Δpp)
Pool (Δpp)
Oracle
Eng. (Δrel)
Pool (Δrel)
Oracle
Eng. (Δrel)
Pool (Δrel)
Oracle
Qwen3-0.6B
14
0.710 ( −5.9 )
0.718 ( −5.1 )
0.769
0.120 ( +26% )
0.099 ( +4% )
0.095
0.116 ( +274% )
0.038 ( +23% )
0.031
Qwen3-1.7B
20
0.741 ( −3.7 )
0.766 ( −1.2 )
0.778
0.190 ( +21% )
0.161 ( +3% )
0.157
0.129 ( +249% )
0.042 ( +14% )
0.037
Qwen3-4B
23
0.729 ( −2.9 )
0.741 ( −1.7 )
0.758
0.213 ( +16% )
0.190 ( +4% )
0.183
0.125 ( +303% )
0.043 ( +39% )
0.031
Qwen3-8B
21
0.739 ( −1.6 )
0.743 ( −1.2 )
0.755
0.212 ( +10% )
0.196 ( +2% )
0.192
0.118 ( +237% )
0.037 ( +6% )
0.035
Qwen3.5-4B
19
0.684 ( −7.2 )
0.737 ( −1.9 )
0.756
0.178 ( +19% )
0.155 ( +3% )
0.150
0.083 ( +186% )
0.042 ( +45% )
0.029
Appendix
Table 10: Self-probe quality at the shared reporting layer. Values are macro-averages over the nine non-English test languages. Eng. uses 2,100 English prompts, Pool uses 210 prompts per language (2,100 total), and Oracle uses 2,100 prompts in each target language. Parentheses report deltas relative to Oracle. For AUROC, Δpp=100(X−Oracle) is the absolute difference in percentage points. For Brier and ECE, Δrel=100(X−Oracle)/Oracle is the relative percentage difference. Red indicates worse performance than Oracle and blue indicates better performance. AUROC treats the five rollouts as weighted binary outcomes.
Target
Predictor
AUROC ↑
Brier ↓
NLL ↓
ECE ↓
Qwen3-0.6B
English raw
0.769
0.109
0.465
0.112
English + cal.
0.769
0.106
0.385
0.093
Pooled raw
0.773
0.088
0.400
0.035
Qwen3-1.7B
English raw
0.759
0.176
0.547
0.103
English + cal.
0.759
0.172
0.539
0.085
Pooled raw
0.777
0.158
0.520
0.051
Appendix
Table 11: Prediction quality of the ETD routing heads, macro-averaged over the nine non-English test languages. English calibration is fit on multilingual validation data with a positive slope. Bold marks the best value per metric within each target model’s three predictor settings.
Figure 4: Per-language reliability diagrams for Qwen3-0.6B at its shared reporting layer. The English-trained probe is transferred unchanged to each target language, whereas the pooled probe uses the same total training budget distributed equally across the ten languages. The diagonal denotes perfect calibration.
Figure 5: Per-language reliability diagrams for Qwen3-1.7B at its shared reporting layer. The English-trained probe is transferred unchanged to each target language, whereas the pooled probe uses the same total training budget distributed equally across the ten languages. The diagonal denotes perfect calibration.
Figure 6: Per-language reliability diagrams for Qwen3-8B at its shared reporting layer. The English-trained probe is transferred unchanged to each target language, whereas the pooled probe uses the same total training budget distributed equally across the ten languages. The diagonal denotes perfect calibration.
Figure 7: Per-language reliability diagrams for Qwen3.5-4B at its shared reporting layer. The English-trained probe is transferred unchanged to each target language, whereas the pooled probe uses the same total training budget distributed equally across the ten languages. The diagonal denotes perfect calibration.
Figure 8: Per-language reliability diagrams for Qwen3.5-9B at its shared reporting layer. The English-trained probe is transferred unchanged to each target language, whereas the pooled probe uses the same total training budget distributed equally across the ten languages. The diagonal denotes perfect calibration.
Figure 9: Per-language reliability diagrams for gpt-oss-20b-low at its shared reporting layer. The English-trained probe is transferred unchanged to each target language, whereas the pooled probe uses the same total training budget distributed equally across the ten languages. The diagonal denotes perfect calibration.
Figure 10: Per-language reliability diagrams for gpt-oss-20b-high at its shared reporting layer. The English-trained probe is transferred unchanged to each target language, whereas the pooled probe uses the same total training budget distributed equally across the ten languages. The diagonal denotes perfect calibration.
Despite the widespread multilingual deployment of large language models, post-training pipelines remain predominantly English-centric, contributing to performance disparities across languages. We present a systematic, controlled study of the interplay between training language coverage, model scale, and task domain, based on 220 supervised fine-tuning runs on parallel translated multilingual data mixtures spanning mathematical reasoning and API calling tasks, with models up to 8B parameters. We find that English-only post-training is typically suboptimal: incorporating even a single non-English language improves both English performance and cross-lingual generalization. Increasing language diversity during post-training generally yields further gains, particularly for low-resource languages, while performance on high-resource languages tends to plateau rather than degrade. Moreover, greater language diversity enables strong zero-shot transfer to unseen languages, reducing the need for direct inclusion, though gains remain limited for typologically distant, low-resource languages.
Latent language identification is often used to argue that multilingual language models route computation through language-specific states, such as English pivots. However, existing probes infer latent language from different signals, such as the geometry of hidden states or what can be decoded from intermediate representations. Since such claims shape conclusions about how models share and route information across languages, we ask whether these probes measure the same phenomenon or expose distinct aspects of multilingual computation. We study this question across model families, training regimes, domains, tasks, checkpoints, and up to 27 languages. We find that identification probes systematically disagree: the GMM-based representation probe, which draws evidence from hidden state geometry, shows earlier cross-lingual mixing, whereas decoding-based probes, which rely on output-space decodability, retain sharper language-specific and more English-biased signals. These differences track model multilinguality and training progression, but are comparatively stable across domains. Our results suggest a more cautious interpretation of latent language identification, where current probes expose different aspects of multilingual processing, rather than directly revealing a single internal lingua franca.
Translation-based prompting is widely used in multilingual LLMs, yet its effectiveness varies across languages and tasks. We evaluate prompting strategies across ten languages of different resource levels and four benchmarks. Our analysis shows that no single strategy is universally optimal. Translation strongly benefits low-resource languages even when translation quality is imperfect, high-resource languages gain little, and prompt-based self-routing underperforms explicit translation. Motivated by these findings, we formulate prompting strategy selection as a learned decision problem and introduce lightweight classifiers that predict whether native or translation-based prompting is optimal for each instance. The classifiers achieve statistically significant improvements over fixed strategies across four benchmarks and generalize to unseen task formats not observed during training. Further analysis reveals that language resource level, rather than translation quality alone, determines when translation is beneficial.
Wei-Chi Wu, Sheng-Lun Wei, Hen-Hsen Huang +1
Department of Computer Science and Information Engineering, National Taiwan University, Taiwan · Institute of Information Science, Academia Sinica, Taiwan · AI Research Center (AINTU), National Taiwan University, Taiwan