cs.CLOct 1, 2026

Cross-Lingual Alignment for Decoder-Only Models using MoE Routers

Authors: Lucas Bandarkar, Clark Peng, Ahmed Haj Ahmed, Aditi Khandelwal, Nanyun Peng

Organizations: University of California, Los Angeles∗ · Haverford College · MILA - Quebec AI Institute & McGill University

Abstract

Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is not possible in decoder-only LLMs because of varying multilingual tokenization. However, a growing amount of research suggests that even in LLMs, higher cross-lingual representational alignment leads to improved cross-lingual transfer. In this paper, we propose a novel approach to reimagine cross-lingual contrastive learning given the architectural constraints of modern LLMs. Rather than applying an auxiliary alignment loss on hidden states, we propose using the outputs of the mixture-of-experts (MoE) routers as the target for alignment. Router outputs lend themselves better to pooling over many tokens, enabling more reliable cross-lingual comparisons at the sequence-level. Controlled continual pre-training experiments on four open-source MoEs show that incorporating this routing loss also aligns the underlying hidden representations across languages. Most importantly, this loss improves multilingual performance on our diverse evaluation suite, demonstrating the potential of cross-lingual MoE router alignment.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RA-MoE: Routing-Aligned Fine-Tuning for Multilingual Adaptation of Mixture-of-Experts Models

    May 27, 2026Guanzhi Deng, Kuan Wu, Haibo Wang +6Multilingual Language ModelsDownstream Reasoning

  2. Mixture of Experts for Low-Resource LLMs

    May 17, 2026Ori Bar Joseph, Smadar Arvatz, Noam Kayzer +2Mixture-Of-ExpertsMultilingual Language Models

  3. Unveiling Language Routing Isolation in Multilingual MoE Models for Interpretable Subnetwork Adaptation

    Apr 4, 2026Kening Zheng, Wei-Chieh Huang, Jiahao Huo +9Mixture-Of-Experts Large Language ModelsMultilingual Language Models