cs.CLSep 29, 2026

What Does Post-Training Change in Multilingual Reasoning?

Authors: Hongyang Li, Xiao Li, Caesar Wu, Grégoire Danoy, Pascal Bouvry

Organizations: University of Luxembourg · Seafill Open-Source Community

Abstract

Open-source reasoning models provide unequal access to reasoning capability across languages. When a model can solve a problem but cannot deliver a complete solution in the user's language, language becomes an access barrier rather than merely a source of performance variation. We audit Qwen3 checkpoints on competition-mathematics tasks in eleven languages. Across the ten non-English languages, only 15.4-17.9% of problems receive a correct, terminating solution with visible reasoning in the requested language in any of 16 samples, compared with 92.9% in English. To identify the source of this disparity, we evaluate thirteen endpoints from one model family, spanning released checkpoints, multilingual supervised fine-tuning (SFT) at two scales, controlled SFT ablations, and three reinforcement-learning (RL) reward formulations. We jointly track correctness, language adherence, termination, and delivery efficiency. The dominant bottleneck shifts across post-training stages. Released models often reason in English. Multilingual SFT restores target-language reasoning, but accuracy declines across multilingual, English-only, and single-language SFT runs, showing that this cost is not specific to multilingual mixing; non-English reasoning traces additionally become prone to non-terminating loops. RL restores termination in both arms at no cost in accuracy, but only the arm whose reward includes a language term delivers: rewarding correctness alone returns the model to English. Together, these stages establish a constructive post-training path from English-pivoted capability to multilingual reasoning that is reliably delivered.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Rethinking the Multilingual Reasoning Gap with Layer Swap

    May 26, 2026Maxence Lasbordes, Amélie Chatelain, Djamé SeddahLLM Reasoning StrategiesLanguage Pairs

  2. What Makes Good Multilingual Reasoning? Disentangling Traces with Measurable Features

    Apr 6, 2026Dayeon Ki, Kevin Duh, Marine CarpuatMultilingual BenchmarkReasoning Traces

  3. LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance

    May 21, 2026Yuchun Fan, Bei Li, Peiguang Li +9LLM Reasoning StrategiesMultilingual Language Models