What Does Post-Training Change in Multilingual Reasoning?
Organizations: University of Luxembourg · Seafill Open-Source Community
Abstract
Open-source reasoning models provide unequal access to reasoning capability across languages. When a model can solve a problem but cannot deliver a complete solution in the user's language, language becomes an access barrier rather than merely a source of performance variation. We audit Qwen3 checkpoints on competition-mathematics tasks in eleven languages. Across the ten non-English languages, only 15.4-17.9% of problems receive a correct, terminating solution with visible reasoning in the requested language in any of 16 samples, compared with 92.9% in English. To identify the source of this disparity, we evaluate thirteen endpoints from one model family, spanning released checkpoints, multilingual supervised fine-tuning (SFT) at two scales, controlled SFT ablations, and three reinforcement-learning (RL) reward formulations. We jointly track correctness, language adherence, termination, and delivery efficiency. The dominant bottleneck shifts across post-training stages. Released models often reason in English. Multilingual SFT restores target-language reasoning, but accuracy declines across multilingual, English-only, and single-language SFT runs, showing that this cost is not specific to multilingual mixing; non-English reasoning traces additionally become prone to non-terminating loops. RL restores termination in both arms at no cost in accuracy, but only the arm whose reward includes a language term delivers: rewarding correctness alone returns the model to English. Together, these stages establish a constructive post-training path from English-pivoted capability to multilingual reasoning that is reliably delivered.
Figures & tables
| stage | intervention | endpoint | eff. vs En | dominant residual | |||||
|---|---|---|---|---|---|---|---|---|---|
| I | released | Qwen3-8B | language selection | ||||||
| II | multilingual SFT | + SFT (8B) | termination | ||||||
| I | released | Qwen3-4B | language selection | ||||||
| II | multilingual SFT | + SFT | termination | ||||||
| III | correctness-only RL | RL: correctness only | returns to English | ||||||
| III | target-language RL | RL: additive gated | answer correctness |
| run | initialized from | examples | batch | epochs (cfg.) | evaluated at |
| supervised | |||||
| multilingual 4B | Qwen3-4B | 43,218 | 16 | 5 | epoch 2, step 5,402 |
| multilingual 8B | Qwen3-8B | 56,307 | 16 | 2 | epoch 1, step 3,519 |
| English-only | Qwen3-4B | 9,800 | 4 | 3 | epoch 1, step 2,450 |
| specialist zh | Qwen3-4B | 9,334 | 4 | 3 | epoch 1 |
| specialist es | Qwen3-4B | 9,491 | 4 | 3 | epoch 1 |
| all attempts | -only | ||||
|---|---|---|---|---|---|
| stage | adjusted | raw | adjusted | raw | token yield |
| Qwen3-4B | – | – | |||
| + SFT | |||||
| RL: correctness only | – | – | |||
| RL: gated | |||||
| Qwen3-8B | – | – | |||
| what the prompt supplies | zh | es | fr | ar | ru |
|---|---|---|---|---|---|
| Qwen3-4B (released) | |||||
| problem, asks | |||||
| English problem, asks | |||||
| multilingual SFT | |||||
| problem, asks | |||||
| English problem, asks | |||||
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| checkpoint | language | @16 [95% CI] | @16 [95% CI] | cap only | looping | tokens | ||
|---|---|---|---|---|---|---|---|---|
| Qwen3-4B | English | [85.7, 98.6] | [85.7, 98.6] | 9,568 | ||||
| Chinese | [84.3, 97.1] | [84.3, 97.1] | 8,127 | |||||
| Spanish | [84.3, 97.1] | [0.0, 4.3] | 7,928 | |||||
| French | [85.7, 98.6] | [0.0, 0.0] | 7,158 | |||||
| Arabic | [84.3, 97.1] | [0.0, 0.0] | 7,613 | |||||
| Russian | [77.1, 92.9] | [77.1, 92.9] | 8,380 |
| outcome | two instruments for | |||||||
|---|---|---|---|---|---|---|---|---|
| model | lang | gap | langid- | judge-en | judge-tgt | judge-frac | ||
| Qwen3-4B | is | |||||||
| sv | ||||||||
| sw | ||||||||
| ta | ||||||||
| ur | ||||||||
| instructed language | target frac. | |||
|---|---|---|---|---|
| en | 120 | |||
| zh | 120 | |||
| es | 120 | |||
| fr | 120 | |||
| ar | 120 | |||
| ru | 120 |
| mechanical fidelity | blind | |||||
| language | translator | boxed | digits | L a T e X | trunc. | faithfulness |
| Chinese | Qwen3-14B | |||||
| Hunyuan-MT-7B | ||||||
| Spanish | Qwen3-14B | |||||
| Hunyuan-MT-7B | ||||||
| French | Qwen3-14B | |||||
| language | coherence | fidelity | reaches \boxed | math kept | no major leak |
|---|---|---|---|---|---|
| Chinese | |||||
| Spanish | |||||
| French | |||||
| Arabic | |||||
| Russian | |||||
| English (untranslated) | — | — |
| Hyperparameter | Value |
|---|---|
| Framework | verl SFT trainer, FSDP |
| Base models | Qwen3-4B, Qwen3-8B (released, post-trained) |
| Chat template | Qwen3, thinking enabled |
| Sequence cutoff length | 32,768 (Arabic specialist: 24,576) |
| Padding | no_padding |
| Truncation | left |
| Hyperparameter | Value |
|---|---|
| Framework | verl, FSDP |
| Algorithm | PPO, GAE advantages |
| Initialized from | multilingual SFT on Qwen3-4B / additive @ step 200 (gated) |
| Training dataset | DAPO-Math-17K, localized (16,754 prompts; 16,751 after the over-length filter) |
| Max prompt length | 2048 |
| Max response length | 16384 / 8192 (gated) |
| comparison | |||
|---|---|---|---|
| Qwen3-4B SFT | |||
| Qwen3-8B SFT | |||
| Qwen3-4B English-only | |||
| mixed spec. (own) | – | – | |
| mixed spec. (en) | – | – |
| edge | lang | term. fail. | ||
|---|---|---|---|---|
| Qwen3-4B multilingual SFT | en | |||
| zh | ||||
| es | ||||
| fr | ||||
| ar | ||||
| ru |
| English | Chinese | Spanish | French | Arabic | Russian | |
| Qwen3-4B multilingual SFT-4B | ||||||
| looping | ∗ | ∗ | ∗ | ∗ | ∗ | |
| ran out of room | ||||||
| Qwen3-8B multilingual SFT-8B | ||||||
| looping | ∗ | ∗ | ∗ | ∗ | ∗ | |
| ran out of room | ||||||
| English | Chinese | Spanish | French | Arabic | Russian | non-English | |
|---|---|---|---|---|---|---|---|
| Qwen3-4B | |||||||
| + SFT | |||||||
| RL: gated | |||||||
| Qwen3-8B | |||||||
| + SFT (8B) |
| diagnostic | additive endpoint | gated continuation | scope |
|---|---|---|---|
| task–language component | training configuration | ||
| positive reward mass on wrong | recorded rollouts | ||
| correct and target-language | late training rollouts | ||
| mean response tokens | 1960 | 1983 | late training rollouts |
| AIME accuracy | monitored development |
| en | zh | es | fr | ar | ru | |
|---|---|---|---|---|---|---|
| en | — | |||||
| zh | — | |||||
| es | — | |||||
| fr | — | |||||
| ar | — | |||||
| ru | — |
| endpoint | RL corpus | (70) | (56) | change | excess over SFT |
|---|---|---|---|---|---|
| Qwen3-4B | no | — | |||
| + SFT | no | — | |||
| RL: correctness only | yes | ||||
| RL: additive | yes | ||||
| RL: gated | yes |
| endpoint | en | zh | es | fr | ar | ru |
|---|---|---|---|---|---|---|
| Qwen3-4B | ( ) | ( ) | ( ) | ( ) | ( ) | ( ) |
| Qwen3-8B | ( ) | ( ) | ( ) | ( ) | ( ) | ( ) |
| multilingual SFT | ( ) | ( ) | ( ) | ( ) | ( ) | ( ) |
| multilingual SFT (8B) | ( ) | ( ) | ( ) | ( ) | ( ) | ( ) |
| English-only SFT | ( ) | ( ) | ( ) | ( ) | ( ) | ( ) |
| specialist zh | ( ) | ( ) | – | – | – | – |
| stage | tokens charged | zh | es | fr | ar | ru |
|---|---|---|---|---|---|---|
| Qwen3-4B | all attempts | (108) | (0) | (0) | (0) | (98) |
| -only | (131) | – | – | (128) | ||
| + SFT | all attempts | (63) | (71) | (69) | (48) | (59) |
| -only | (120) | (88) | (86) | (92) | (82) | |
| RL: correctness only | all attempts | (44) | (0) | (0) | (0) | (1) |
| -only | (134) | – | – | – |
| stage | en | zh | es | fr | ar | ru |
|---|---|---|---|---|---|---|
| Qwen3-4B | ||||||
| + SFT | ||||||
| RL: correctness only | ||||||
| RL: additive | ||||||
| RL: gated | ||||||
| Qwen3-8B |