Cross-Lingual Alignment for Decoder-Only Models using MoE Routers
Organizations: University of California, Los Angeles∗ · Haverford College · MILA - Quebec AI Institute & McGill University
Abstract
Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is not possible in decoder-only LLMs because of varying multilingual tokenization. However, a growing amount of research suggests that even in LLMs, higher cross-lingual representational alignment leads to improved cross-lingual transfer. In this paper, we propose a novel approach to reimagine cross-lingual contrastive learning given the architectural constraints of modern LLMs. Rather than applying an auxiliary alignment loss on hidden states, we propose using the outputs of the mixture-of-experts (MoE) routers as the target for alignment. Router outputs lend themselves better to pooling over many tokens, enabling more reliable cross-lingual comparisons at the sequence-level. Controlled continual pre-training experiments on four open-source MoEs show that incorporating this routing loss also aligns the underlying hidden representations across languages. Most importantly, this loss improves multilingual performance on our diverse evaluation suite, demonstrating the potential of cross-lingual MoE router alignment.
Figures & tables
| Model | Condition | Telugu (8) | Kyrgyz (4) | Kannada (4) | Thai (7) |
|---|---|---|---|---|---|
| Qwen3-30B-A3B | Original checkpoint | 44.6 | 40.8 | 49.3 | 43.0 |
| Baseline | 45.4 | 43.5 | 49.6 | 44.9 | |
| + aux routing loss (ours) | 46.3 | 45.6 | 51.4 | 45.4 | |
| Router-only training | 45.1 | 42.2 | 49.9 | 43.8 | |
| GPT-OSS-20B | Original checkpoint | 42.5 | 35.9 | 47.5 | 34.9 |
| Baseline | 46.5 | 42.0 | 55.5 | 39.9 |
| Condition | Sinhala (5) | Hungarian (6) |
|---|---|---|
| Original checkpoint | 24.0 | 36.3 |
| Baseline | 26.1 | 37.4 |
| + routing loss (ours) | 26.8 | 38.1 |
| + hidden-state loss | 26.3 | 36.9 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| (a) Training configuration and aggregate results | |||||
| Language | Condition | Learning rate | LM loss | AVG | |
| Telugu (tel) | Original | – | – | 0.77 | 44.6 |
| Baseline | – | 0.65 | 45.4 | ||
| + routing | 2 | 0.7 | 46.3 | ||
| Router-only | 10 | 0.78 | 45.1 | ||
| Kyrgyz (kir) | Original | – | – | 3.06 | 40.8 |
| (a) Training configuration and aggregate results | |||||
| Language | Condition | Learning rate | LM loss | AVG | |
| Telugu (tel) | Original | – | – | 3.05 | 42.5 |
| Baseline | – | 2.69 | 46.5 | ||
| + routing | 40 | 2.87 | 47.0 | ||
| Router-only | 40 | 3.04 | 43.2 | ||
| Kyrgyz (kir) | Original | – | – | 4.28 | 35.9 |
| (a) Training configuration and aggregate results | |||||
| Language | Condition | Learning rate | LM loss | AVG | |
| Vietnamese (vie) | Original | – | – | 2.93 | 34.8 |
| Baseline | – | 2.31 | 36.5 | ||
| + routing | 200 | 2.402 | 36.8 | ||
| Router-only | 200 | 2.85 | 36.5 | ||
| Sinhala (sin) | Original | – | – | 1.09 | 31.3 |
| (a) Training configuration and aggregate results | |||||
| Language | Condition | Learning rate | LM loss | AVG | |
| Vietnamese (vie) | Original | – | – | 3.46 | 44.2 |
| Baseline | – | 3.16 | 45.6 | ||
| + routing | 10 | 3.23 | 45.6 | ||
| Router-only | 40 | 3.45 | 44.8 | ||
| Sinhala (sin) | Original | – | – | 1.79 | 24.0 |
| Model | Params (Active) | Num. Layers | Active Experts / Total | Layers for |
|---|---|---|---|---|
| Qwen3-30B-A3B | 31B (A3B) | 48 | 8 / 128 | 7–34 |
| GPT-OSS-20B | 22B (A4B) | 24 | 4 / 32 | 4–17 |
| Marco-Nano | 8B (A0.6B) | 28 | 8 / 238 | 7–19 |
| Granite-4.0-H-Tiny | 7B (1B) | 40 | 6 / 64 | 10–35 |
| Language | Training-data sources |
|---|---|
| Vietnamese | MTet ( Ngo et al., 2022 ) ; Bactrian-X ( Li et al., 2023 ) ; NLLB English bitext ( NLLB Team et al., 2022 ) ; OPUS-100 ( Zhang et al., 2020 ) . |
| Sinhala | Bactrian-X ( Li et al., 2023 ) ; NLLB English bitext ( NLLB Team et al., 2022 ) ; OPUS-100 ( Zhang et al., 2020 ) . |
| Hungarian | NLLB English bitext ( NLLB Team et al., 2022 ) ; OPUS-100 ( Zhang et al., 2020 ) . |
| Telugu | Updesh ( Chitale et al., 2026 ) ; Bactrian-X ( Li et al., 2023 ) ; Samanantar ( Ramesh et al., 2022 ) ; NLLB English bitext ( NLLB Team et al., 2022 ) ; OPUS-100 ( Zhang et al., 2020 ) . |
| Kannada | Updesh ( Chitale et al., 2026 ) ; Samanantar ( Ramesh et al., 2022 ) ; NLLB English bitext ( NLLB Team et al., 2022 ) ; OPUS-100 ( Zhang et al., 2020 ) . |
| Thai | Bactrian-X ( Li et al., 2023 ) ; SCB-MT-EN-TH-2020 ( Lowphansirikul et al., 2022 ) ; OPUS-100 ( Zhang et al., 2020 ) . |