Mixture-of-Experts Language Models

Latest papers 147

All topics
CardsList
  1. RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing

    Oct 8, 2026Ilya Lasy, Nora Yinuo Cai, Kola AyonrindeMixture of ExpertsExpert Routing

  2. Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts

    Oct 8, 2026Yunkai Chai, Tong Zhu, Xiaoye Qu +4Mixture-of-Experts Language ModelsSparse Mixture-of-Experts

  3. Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression

    Oct 7, 2026Tianxiao Cao, Jiahe Shao, Yuning Qiu +3LLM CompressionMixture-of-Experts Language Models

  4. Distributionally Robust Mixture-of-Experts Training

    Oct 5, 2026Xin Teng, Muxiao Li, Hongyi WenMixture-of-Experts Language ModelsSparse Mixture-of-Experts

  5. How Sparse Probability Maps Shape Mixture-of-Experts Routing

    Oct 5, 2026Tomás Brogueira, Marcos Treviso, Miguel CouceiroMixture-of-Experts Language ModelsMixture-of-Experts Inference

  6. Cross-Lingual Alignment for Decoder-Only Models using MoE Routers

    Oct 1, 2026Lucas Bandarkar, Clark Peng, Ahmed Haj Ahmed +2Mixture-of-Experts Language ModelsCross-Lingual Representation Alignment

  7. ITC-MoE: Importance-guided Token-aware Compression for MoE Diffusion Language Models

    Oct 1, 2026Lianjun Liu, Shipeng Li, You Huang +5Mixture-of-Experts Language ModelsLLM Compression

  8. Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

    Oct 1, 2026Di He, Pengxiang Li, Da Chang +3Mixture-of-Experts Language ModelsRecurrent Transformers

  9. Score the Update, Not the Token: Descent-Aligned Routing for Combinatorial LoRA Experts

    Sep 30, 2026Priya Nair, Lukas Brenner, Maya Lindqvist +5Mixture-of-Experts Language ModelsParameter-Free Mixture-of-Experts Routing

  10. Scaling Laws for Looped Mixture of Experts

    Sep 30, 2026Yanbei Chen, Anirudh Goyal, Raghuraman KrishnamoorthiMixture-of-Experts Language ModelsLanguage Model Scaling Laws

  11. ID Balancing: Stable Training of Extremely Sparse MoE via PID-Based Load Control

    Sep 30, 2026Peng Jin, Zihan Qiu, Zekun Wang +8Expert Load BalancingMixture-of-Experts Language Models

  12. Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models

    Sep 29, 2026Yury Nahshan, Nati Daniel, Jacob Goldberger +1Mixture-of-Experts Language ModelsParameter-Free Mixture-of-Experts Routing

  13. How to Loop MoE: Flatten the Experts, Untie the Attention

    Sep 28, 2026Shouren Wang, Chuang Ma, Mohsen Hariri +6Mixture-of-Experts Language ModelsSparse Mixture-of-Experts

  14. MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference

    Sep 28, 2026Junfeng Wu, Zehao Fan, Hadjer Benmeziane +3Expert OffloadingMixture-of-Experts Language Models

  15. EAT: Expert Account Tracker for Efficient MoE Inference

    Sep 27, 2026Yuexian Li, Yifei Yang, Zouying Cao +1Mixture-of-Experts PruningMixture-of-Experts Language Models

  16. RAZOR: Pruning Replaceable Experts in LLMs

    Sep 24, 2026Mingyang Song, Mao ZhengMixture-of-Experts PruningLLM Pruning

  17. You Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMs

    Sep 22, 2026Yuanteng Chen, Qiwei Lai, Chen Tianqi +7Mixture-of-Experts PruningMixture-of-Experts Language Models

  18. Higher-order pruning of experts in mixture-of-experts language models

    Sep 16, 2026Alex M. Tseng, Prannay Kaul, Luca Zancato +2Mixture-of-Experts PruningMixture-of-Experts Language Models

  19. MoRE: Mixture of Reused Experts

    Sep 16, 2026Eric S. Qiu, Utku Umur Acikalin, Justin Lovelace +4Mixture-of-Experts Language ModelsSparse Mixture-of-Experts

  20. Beyond Magnitude: Contrastive Routing for Modular Mixture-of-Experts

    Sep 1, 2026Nikolaos Xiros, Dimitrios Damianos, Maria-Eleni Zoumpoulidi +3Mixture-of-Experts Language ModelsExpert Routing

  21. Instella-MoE Technical Report

    Sep 1, 2026Jiang Liu, Sudhanshu Ranjan, Prakamya Mishra +10Mixture-of-Experts Language ModelsLanguage Model Post-Training

  22. Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs

    Sep 1, 2026Seungwoo Jung, Dohyeok Kwon, Seungmin Cha +4Mixture-of-Experts Language ModelsLLM Compression

  23. Q-Strata: Hierarchical Bit Allocation for Mixed-Precision Quantization of Mixture-of-Experts LLMs

    Aug 31, 2026Deokjae Lee, Sihun Chu, Hyun Oh SongMixed-Precision QuantizationLLM Quantization

  24. A.X K2 Technical Report

    Aug 31, 2026Cheolseung Baek, Dhammiko Arya, Eunki Kim +40Mixture-of-Experts Language ModelsLong-Context Language Modeling