Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is not possible in decoder-only LLMs because of varying multilingual tokenization. However, a growing amount of research suggests that even in LLMs, higher cross-lingual representational alignment leads to improved cross-lingual transfer. In this paper, we propose a novel approach to reimagine cross-lingual contrastive learning given the architectural constraints of modern LLMs. Rather than applying an auxiliary alignment loss on hidden states, we propose using the outputs of the mixture-of-experts (MoE) routers as the target for alignment. Router outputs lend themselves better to pooling over many tokens, enabling more reliable cross-lingual comparisons at the sequence-level. Controlled continual pre-training experiments on four open-source MoEs show that incorporating this routing loss also aligns the underlying hidden representations across languages. Most importantly, this loss improves multilingual performance on our diverse evaluation suite, demonstrating the potential of cross-lingual MoE router alignment.
Figures & tables
Figure 1: Impact of our auxiliary alignment loss on routing alignment during CPT of the post-trained Qwen3-30B-A3B. In this case, vanilla CPT did not alter routing behavior (the blue and gray lines overlap), whereas our method (orange) substantially reduced routing divergence. The shaded orange region indicates the middle layers to which our loss was applied.
Figure 2: Comparison of the naive padded implementation (left) and our shared packed implementation (right). In the packed pass, the (typically shorter) source tokens exit early.
Figure 3: Layer-wise routing divergence from English for GPT-OSS-20B on Kannada (top), Granite-4.0-H-Tiny on Hungarian (middle), and Marco-Nano on Sinhala (bottom), comparing the base model, baseline training, and training with the auxiliary routing loss. Divergence is measured by mean entropy-normalized JS-divergence. The orange zone is the layers M where LXLR is applied.
Figure 4: Layer-wise hidden-state similarity to English for Qwen3-30B-A3B. Solid lines show the base model and dashed lines show the model trained with the auxiliary cross-lingual loss for Telugu, Kyrgyz, Kannada, and Thai.
Model
Condition
Telugu (8)
Kyrgyz (4)
Kannada (4)
Thai (7)
Qwen3-30B-A3B
Original checkpoint
44.6
40.8
49.3
43.0
Baseline
45.4
43.5
49.6
44.9
+ aux routing loss (ours)
46.3
45.6
51.4
45.4
Router-only training
45.1
42.2
49.9
43.8
GPT-OSS-20B
Original checkpoint
42.5
35.9
47.5
34.9
Baseline
46.5
42.0
55.5
39.9
Table 1: Average benchmark scores for each model, language, and training condition. Parentheses in the language headers indicate the number of tasks included in each average. Bold marks the best score within each model–language pair, including ties at the reported precision. Original checkpoint denotes the released post-trained model before our training. The baseline uses only target-language LM training; our method adds the routing-alignment loss. Router-only training applies the auxiliary loss while updating only the routers in the selected middle layers.
Condition
Sinhala (5)
Hungarian (6)
Original checkpoint
24.0
36.3
Baseline
26.1
37.4
+ routing loss (ours)
26.8
38.1
+ hidden-state loss
26.3
36.9
Table 2: Routing versus hidden-state alignment for Marco-Nano. Entries are the reported average benchmark scores over the benchmarks available in each language; parentheses in the headers give the number of tasks in each average. Bold marks the best score in each column. Hidden-state alignment uses α=5000 ; routing alignment uses α=10 for both languages. Full hyperparameters and task-level scores appear in Table 6 .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Overview of our cross-lingual alignment method to supplement the formal description in Section 3.2 .
(a) Training configuration and aggregate results
Language
Condition
α
Learning rate
LM loss
AVG
Telugu (tel)
Original
–
–
0.77
44.6
Baseline
–
3×10−6
0.65
45.4
+ routing
2
3×10−6
0.7
46.3
Router-only
10
3×10−6
0.78
45.1
Kyrgyz (kir)
Original
–
–
3.06
40.8
Appendix
Table 3: Complete results for Qwen3-30B-A3B. Auxiliary alignment is applied to layers 7–34. Bold AVG values mark the best result per language, including ties.
(a) Training configuration and aggregate results
Language
Condition
α
Learning rate
LM loss
AVG
Telugu (tel)
Original
–
–
3.05
42.5
Baseline
–
4×10−7
2.69
46.5
+ routing
40
4×10−7
2.87
47.0
Router-only
40
8×10−7
3.04
43.2
Kyrgyz (kir)
Original
–
–
4.28
35.9
Appendix
Table 4: Complete results for GPT-OSS-20B. Auxiliary alignment is applied to layers 4–17. Bold AVG values mark the best result per language, including ties.
(a) Training configuration and aggregate results
Language
Condition
α
Learning rate
LM loss
AVG
Vietnamese (vie)
Original
–
–
2.93
34.8
Baseline
–
4×10−7
2.31
36.5
+ routing
200
5×10−7
2.402
36.8
Router-only
200
4×10−7
2.85
36.5
Sinhala (sin)
Original
–
–
1.09
31.3
Appendix
Table 5: Complete results for Granite-4.0-H-Tiny. Auxiliary alignment is applied to layers 10–35. Bold AVG values mark the best result per language, including ties.
(a) Training configuration and aggregate results
Language
Condition
α
Learning rate
LM loss
AVG
Vietnamese (vie)
Original
–
–
3.46
44.2
Baseline
–
1×10−6
3.16
45.6
+ routing
10
1.2×10−6
3.23
45.6
Router-only
40
2×10−6
3.45
44.8
Sinhala (sin)
Original
–
–
1.79
24.0
Appendix
Table 6: Complete results for Marco-Nano. Auxiliary alignment is applied to layers 7–19. Bold AVG values mark the best result per language, including ties. Hidden-state alignment is evaluated only for Sinhala and Hungarian.
Figure 6: MoE router divergence for the four models used in experiments. Lower values represent higher alignment, and corresponds to higher-resource languages.
Model
Params (Active)
Num. Layers
Active Experts / Total
Layers for LXLR
Qwen3-30B-A3B
31B (A3B)
48
8 / 128
7–34
GPT-OSS-20B
22B (A4B)
24
4 / 32
4–17
Marco-Nano
8B (A0.6B)
28
8 / 238
7–19
Granite-4.0-H-Tiny
7B (1B)
40
6 / 64
10–35
Appendix
Table 7: MoE architecture and router-layer details for each model. Active parameter counts are shown in parentheses.
Language
Training-data sources
Vietnamese
MTet ( Ngo et al., 2022 ) ; Bactrian-X ( Li et al., 2023 ) ; NLLB English bitext ( NLLB Team et al., 2022 ) ; OPUS-100 ( Zhang et al., 2020 ) .
Sinhala
Bactrian-X ( Li et al., 2023 ) ; NLLB English bitext ( NLLB Team et al., 2022 ) ; OPUS-100 ( Zhang et al., 2020 ) .
Hungarian
NLLB English bitext ( NLLB Team et al., 2022 ) ; OPUS-100 ( Zhang et al., 2020 ) .
Telugu
Updesh ( Chitale et al., 2026 ) ; Bactrian-X ( Li et al., 2023 ) ; Samanantar ( Ramesh et al., 2022 ) ; NLLB English bitext ( NLLB Team et al., 2022 ) ; OPUS-100 ( Zhang et al., 2020 ) .
Kannada
Updesh ( Chitale et al., 2026 ) ; Samanantar ( Ramesh et al., 2022 ) ; NLLB English bitext ( NLLB Team et al., 2022 ) ; OPUS-100 ( Zhang et al., 2020 ) .
Thai
Bactrian-X ( Li et al., 2023 ) ; SCB-MT-EN-TH-2020 ( Lowphansirikul et al., 2022 ) ; OPUS-100 ( Zhang et al., 2020 ) .
Appendix
Table 8: Parallel training-data sources for each target language. The table records source inclusion, rather than nominal or realized mixture ratios.
Mixture-of-Experts (MoE) models enable efficient LLM scaling, yet adapting them to non-English downstream tasks remains challenging. Standard multilingual fine-tuning largely ignores their heterogeneous routing structure. Across multiple MoE models and tasks, we find strong cross-lingual routing alignment in middle layers, with routing divergence associated with target-language performance gaps. Motivated by this observation, we propose RA-MoE (Routing-Aligned MoE Fine-Tuning), a three-stage framework for multilingual MoE adaptation. RA-MoE categorizes parallel examples into four correctness groups (cc/ci/ic/ii) and identifies task-relevant experts in middle layers. It then selectively aligns target-language routing on ci examples toward successful English routing patterns, jointly matching the total routing mass assigned to task experts and its relative allocation among them. Experiments across three MoE models, three downstream tasks, and six target languages show that RA-MoE consistently outperforms standard SFT and strong routing-aware baselines. Further analyses confirm the intended routing changes and reveal that middle-layer task routing is largely shared and transferable across languages, providing mechanistic evidence for the cross-language transferability of task-specific routing.
Guanzhi Deng, Kuan Wu, Haibo Wang +6
City University of Hong Kong, Hong Kong, China · Carnegie Mellon University, Pittsburgh, USA · The University of Hong Kong, Hong Kong, China
Mixture-of-Experts (MoE) architectures enable efficient model scaling, yet expert routing behavior across underrepresented languages remains poorly understood. We analyze routing dynamics in two architecturally distinct MoE models -- a pure Transformer (Qwen3-30B-A3B) and a hybrid Mamba-Transformer (Nemotron-3-Nano-30B-A3B) -- using Hebrew as a morphologically rich, low-resource testbed. Both pre-trained models exhibit \emph{deep-layer routing collapse}: usage entropy drops sharply in final layers and tokens concentrate on a narrow expert subset, a pattern largely absent for English. Continual pre-training (CPT) on balanced bilingual data substantially corrects this imbalance, increasing entropy and shifting routing toward shared, language-agnostic experts; supervised fine-tuning (SFT) alone achieves less complete correction. Extending the analysis to Japanese reveals quantitatively consistent collapse signatures, providing cross-linguistic evidence that the phenomenon is a systematic consequence of pre-training underrepresentation rather than any language-intrinsic property. Routing improvements correlate with consistent downstream benchmark gains, positioning routing entropy and expert specialization as principled diagnostics for multilingual capacity in MoE systems.
Mixture-of-Experts (MoE) models exhibit striking performance disparities across languages, yet the internal mechanisms driving these gaps remain poorly understood. In this work, we conduct a systematic analysis of expert routing patterns in MoE models, revealing a phenomenon we term Language Routing Isolation, in which high- and low-resource languages tend to activate largely disjoint expert sets. Through layer-stratified analysis, we further show that routing patterns exhibit a layer-wise convergence-divergence pattern across model depth. Building on these findings, we propose RISE (Routing Isolation-guided Subnetwork Enhancement), a framework that exploits routing isolation to identify and adapt language-specific expert subnetworks. RISE applies a tripartite selection strategy, using specificity scores to identify language-specific experts in shallow and deep layers and overlap scores to select universal experts in middle layers. By training only the selected subnetwork while freezing all other parameters, RISE substantially improves low-resource language performance while preserving capabilities in other languages. Experiments on 10 languages demonstrate that RISE achieves target-language F1 gains of up to 10.85% with minimal cross-lingual degradation.
Kening Zheng, Wei-Chieh Huang, Jiahao Huo +9
University of Illinois Chicago · University of Maryland · HKUST (Guangzhou)