Pre-generation success probes estimate response correctness from a language model's hidden activations before decoding, enabling cost-aware routing. While prior work has demonstrated their utility primarily on English inputs, we study their reliability across languages along three dimensions: (1) whether they preserve the ranking of likely successes and failures (DISCRIMINATION); (2) whether they retain probabilities that match observed success frequencies (CALIBRATION); and (3) whether they produce scores comparable enough across candidate models for cost-aware multilingual routing (UTILITY). Using 3,000 MATH problems in 10 languages and 8 open-weight model configurations, we compare cross-lingual transfer from English-trained probes and equal-budget pooled multilingual probes. English-trained probes retain useful cross-lingual discrimination but become less well calibrated after transfer. Pooled multilingual supervision improves both properties and yields more reliable estimates of success. In routing experiments, the pooled router achieves a 0.7% higher test success rate while reducing modeled cost by 13.0% relative to always selecting the model with the highest average success. These results show that multilingual routing requires success estimates that remain well calibrated and comparable across languages and models.
Figures & tables
Figure 1: Multilingual routing with pre-generation success probes. Test rollout success versus modeled cost saving relative to always selecting gpt-oss-high , on the balanced ten-language test set. Each point corresponds to a different value of the routing trade-off coefficient λ . Curves compare English-trained, recalibrated English-trained, and equal-budget pooled multilingual probes used for routing.
Language
Qwen3-4B
English
0.716
Chinese
0.630
Italian
0.621
French
0.613
Russian
0.611
Arabic
0.537
Table 1: Empirical success rates of Qwen3-4B across languages. Each value is the average empirical success score over the test split, where each problem-level score is computed from five sampled generations.
Figure 2: Per-language reliability diagrams for Qwen3-4B at its shared reporting layer. The English-trained probe is transferred unchanged to each target language, whereas the pooled probe uses the same total training budget distributed equally across the ten languages. The diagonal denotes perfect calibration.
Evaluation regime
Rollout AUROC ↑
Rollout Brier ↓
ECE ↓
English
0.729
0.213
0.125
Pooled
0.741
0.190
0.043
Oracle
0.758
0.183
0.031
Table 2: Self-probe quality for Qwen3-4B at the shared reporting layer (L23). Values are macro-averages over the nine non-English test languages. AUROC treats the five rollouts as weighted binary outcomes.
Predictor
AUROC ↑
Brier ↓
NLL ↓
ECE ↓
English raw
0.729
0.213
0.624
0.125
English + cal.
0.729
0.208
0.614
0.099
Pooled raw
0.741
0.190
0.573
0.043
Table 3: Prediction quality of the Qwen3-4B ETD routing head at the shared encoder’s reporting layer (L23), macro-averaged over the nine non-English test languages. Bold marks the best value for each metric.
Router
Val. choice
λ
Test success [95% CI]
Saving [95% CI]
English raw
Max. success
0.00
80.8 [78.8,82.8]
33.7 [31.8,35.5]
Anchor match
–
Not attained on validation
English + cal.
Max. success
0.00
85.0 [83.1,86.8]
10.8 [9.0,12.6]
Anchor match
–
Not attained on validation
Pooled raw
Max. success
0.00
86.6 [84.8,88.3]
13.0 [10.2,15.9]
Anchor match
0.10
85.5 [83.6,87.2]
22.9 [19.8,26.3]
Table 4: Validation-selected routing operating points. Anchor matching allows at most 0.1 percentage point validation success loss, then minimizes cost; unattained rows mean no fixed-grid point met that constraint. Brackets are paired bootstrap intervals over base problem IDs; choices are frozen before test.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Language
DeepL
Qwen3.5-27B
Arabic
0.7758
0.7677
Chinese
0.8050
0.8058
French
0.8293
0.8286
Italian
0.8295
0.8278
Russian
0.8139
0.8121
Swahili
0.7913
0.7834
Appendix
Table 5: Mean COMET-Kiwi scores over 3,000 MATH problems per target language and translation system. Higher is better; bold indicates the stronger system within each row.
Figure 3: Ridge coefficient selected at every relative representation depth. The top row shows the same-model probe and the bottom row shows ETD target-specific heads; columns separate English and pooled supervision. In each cell, the selected α minimizes rollout-level validation Brier score over {10,102,103,104,105} . Each row denotes the model supplying the representation for the same-model probe and the predicted target model for ETD; each cell denotes one layer. ETD heads use the shared Qwen3-4B representation.
Probe
peasy
phard
AUROC
ECE
Brier
NLL
Constant
0.50
0.50
0.50
0.00
0.250
0.693
Overconfident
1.00
0.00
0.75
0.25
0.250
∞
Well specified
0.75
0.25
0.75
0.00
0.188
0.562
Appendix
Table 6: Expected metrics for three hypothetical probes. The dataset contains equally many easy and hard problems with true success probabilities 0.75 and 0.25 , respectively. The infinite NLL for the overconfident probe is the theoretical value before numerical clipping.
Model
Input price
Output price
Avg. input
Avg. output
Exp. cost
Relative
($/M tok.)
($/M tok.)
tokens
tokens
($/k gen.)
Exp. cost
Qwen3-0.6B
0.010
0.060
205.12
593.45
0.0377
0.0111
Qwen3-1.7B
0.028
0.170
205.12
2361.62
0.4073
0.1196
Qwen3-4B
0.067
0.400
205.12
2056.09
0.8361
0.2456
gpt-oss-low
0.204
1.224
226.57
946.03
1.2042
0.3537
Qwen3.5-4B
0.067
0.400
186.49
3042.49
1.2294
0.3611
Appendix
Table 7: Proxy prices used in the routing simulation. Average input lengths are computed over the balanced pooled training sample used for routing-cost estimation (210 prompts per language and 2,100 prompts per model), using each model’s native tokenizer and chat template. Average output lengths are computed over the same sample. Relative expected costs are normalized by the average standalone input-plus-output cost of gpt-oss-high .
Shared encoder
Layer
Max. validation success
Prefill / anchor
Qwen3-0.6B
21
88.1%
0.06%
Qwen3-1.7B
19
87.9%
0.17%
Qwen3-4B
23
88.3%
0.40%
Qwen3-8B
26
88.1%
0.80%
Qwen3.5-4B
10
88.2%
0.37%
Qwen3.5-9B
13
88.2%
0.81%
Appendix
Table 8: Validation-only shared-encoder comparison under pooled ETD supervision. For each encoder, the table reports its validation-selected layer and the maximum success attained over the fixed validation routing grid. No test outcomes enter the comparison. Prefill cost is the encoder’s standalone input-only cost as a percentage of the standalone gpt-oss-high anchor cost, using the prices and training-mean input lengths in Appendix D .
Language
Qwen3-0.6B
Qwen3-1.7B
Qwen3-4B
Qwen3-8B
Qwen3.5-4B
Qwen3.5-9B
gpt-oss-low
gpt-oss-high
Mean
English
0.279
0.508
0.716
0.728
0.916
0.935
0.895
0.925
0.738
Chinese
0.220
0.408
0.630
0.638
0.851
0.898
0.754
0.886
0.661
French
0.179
0.388
0.613
0.663
0.848
0.895
0.789
0.911
0.661
Russian
0.170
0.384
0.611
0.661
0.837
0.889
0.761
0.884
0.650
Italian
0.160
0.378
0.621
0.654
0.835
0.880
0.755
0.897
0.648
Arabic
0.134
0.312
0.537
0.600
0.785
0.841
0.749
0.885
0.605
Appendix
Table 9: Empirical success rates of the evaluated models across languages. Each value is the average empirical success score over the test split, where each problem-level score is computed from five sampled generations. Languages are ordered by decreasing average success rate across models. The final row reports the average success rate for each model across languages, and the final column reports the average success rate for each language across models.
Rollout AUROC ↑
Rollout Brier ↓
ECE ↓
Model
Layer
Eng. (Δpp)
Pool (Δpp)
Oracle
Eng. (Δrel)
Pool (Δrel)
Oracle
Eng. (Δrel)
Pool (Δrel)
Oracle
Qwen3-0.6B
14
0.710 ( −5.9 )
0.718 ( −5.1 )
0.769
0.120 ( +26% )
0.099 ( +4% )
0.095
0.116 ( +274% )
0.038 ( +23% )
0.031
Qwen3-1.7B
20
0.741 ( −3.7 )
0.766 ( −1.2 )
0.778
0.190 ( +21% )
0.161 ( +3% )
0.157
0.129 ( +249% )
0.042 ( +14% )
0.037
Qwen3-4B
23
0.729 ( −2.9 )
0.741 ( −1.7 )
0.758
0.213 ( +16% )
0.190 ( +4% )
0.183
0.125 ( +303% )
0.043 ( +39% )
0.031
Qwen3-8B
21
0.739 ( −1.6 )
0.743 ( −1.2 )
0.755
0.212 ( +10% )
0.196 ( +2% )
0.192
0.118 ( +237% )
0.037 ( +6% )
0.035
Qwen3.5-4B
19
0.684 ( −7.2 )
0.737 ( −1.9 )
0.756
0.178 ( +19% )
0.155 ( +3% )
0.150
0.083 ( +186% )
0.042 ( +45% )
0.029
Appendix
Table 10: Self-probe quality at the shared reporting layer. Values are macro-averages over the nine non-English test languages. Eng. uses 2,100 English prompts, Pool uses 210 prompts per language (2,100 total), and Oracle uses 2,100 prompts in each target language. Parentheses report deltas relative to Oracle. For AUROC, Δpp=100(X−Oracle) is the absolute difference in percentage points. For Brier and ECE, Δrel=100(X−Oracle)/Oracle is the relative percentage difference. Red indicates worse performance than Oracle and blue indicates better performance. AUROC treats the five rollouts as weighted binary outcomes.
Target
Predictor
AUROC ↑
Brier ↓
NLL ↓
ECE ↓
Qwen3-0.6B
English raw
0.769
0.109
0.465
0.112
English + cal.
0.769
0.106
0.385
0.093
Pooled raw
0.773
0.088
0.400
0.035
Qwen3-1.7B
English raw
0.759
0.176
0.547
0.103
English + cal.
0.759
0.172
0.539
0.085
Pooled raw
0.777
0.158
0.520
0.051
Appendix
Table 11: Prediction quality of the ETD routing heads, macro-averaged over the nine non-English test languages. English calibration is fit on multilingual validation data with a positive slope. Bold marks the best value per metric within each target model’s three predictor settings.
Figure 4: Per-language reliability diagrams for Qwen3-0.6B at its shared reporting layer. The English-trained probe is transferred unchanged to each target language, whereas the pooled probe uses the same total training budget distributed equally across the ten languages. The diagonal denotes perfect calibration.
Figure 5: Per-language reliability diagrams for Qwen3-1.7B at its shared reporting layer. The English-trained probe is transferred unchanged to each target language, whereas the pooled probe uses the same total training budget distributed equally across the ten languages. The diagonal denotes perfect calibration.
Figure 6: Per-language reliability diagrams for Qwen3-8B at its shared reporting layer. The English-trained probe is transferred unchanged to each target language, whereas the pooled probe uses the same total training budget distributed equally across the ten languages. The diagonal denotes perfect calibration.
Figure 7: Per-language reliability diagrams for Qwen3.5-4B at its shared reporting layer. The English-trained probe is transferred unchanged to each target language, whereas the pooled probe uses the same total training budget distributed equally across the ten languages. The diagonal denotes perfect calibration.
Figure 8: Per-language reliability diagrams for Qwen3.5-9B at its shared reporting layer. The English-trained probe is transferred unchanged to each target language, whereas the pooled probe uses the same total training budget distributed equally across the ten languages. The diagonal denotes perfect calibration.
Figure 9: Per-language reliability diagrams for gpt-oss-20b-low at its shared reporting layer. The English-trained probe is transferred unchanged to each target language, whereas the pooled probe uses the same total training budget distributed equally across the ten languages. The diagonal denotes perfect calibration.
Figure 10: Per-language reliability diagrams for gpt-oss-20b-high at its shared reporting layer. The English-trained probe is transferred unchanged to each target language, whereas the pooled probe uses the same total training budget distributed equally across the ten languages. The diagonal denotes perfect calibration.
Department of Computer Science and Information Engineering, National Taiwan University, Taiwan · Institute of Information Science, Academia Sinica, Taiwan · AI Research Center (AINTU), National Taiwan University, Taiwan