Mixture-of-Experts Language Models

Latest papers 147

All topics
CardsList
  1. Evolutionary Soups: Evolving Mixture-of-Experts for Multi-Objective LLM Alignment

    Aug 30, 2026Lingxiao Kong, Steffen Staab, Cong Yang +2Mixture-of-Experts Language ModelsMulti-Objective Language Model Alignment

  2. TradingMoE: Routing the Right Experts in Evolving Markets

    Aug 12, 2026Chang Zhou, Xingtong Yu, Minbin Huang +4Mixture-of-Experts Language ModelsSparse Mixture-of-Experts

  3. DistMoE: Private-data Rehearsal-free Routing in Mixture-of-Experts for Distributed Instruction Tuning

    Aug 10, 2026Mainak Singha, Niccolò Biondi, Elisa Ricci +1Mixture-of-Experts Language ModelsVLM Adaptation

  4. MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts

    Aug 10, 2026Peiwen Li, Shiyang Zhang, Yangtian Zhang +3Multi-Agent LLM SystemsMixture-of-Experts Language Models

  5. Motif 3: Technical Report

    Aug 10, 2026Junghwan Lim, Joon Son Chung, Sungmin Lee +24Mixture-of-Experts Language ModelsLong-Context Language Modeling

  6. RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space Attention

    Aug 8, 2026Anthony. Lui, Mohamed. Elsaied, N. P. SavaniExpert OffloadingMixture-of-Experts Language Models

  7. AFD-Ledger: Deployment Provisioning for Attention--FFN Disaggregation

    Aug 5, 2026Chengyu Qiu, Xiao Fu, Fengcun Li +6Mixture-of-Experts Language ModelsDisaggregated LLM Serving

  8. LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

    Aug 4, 2026Fengqi Zhu, Shaoxuan Xu, Jingyang Ou +11Mixture-of-Experts Language ModelsLanguage Model Scaling Laws

  9. TEXAS: Task-Expert-Aware Supervision for Downstream Mixture-of-Experts LLM Adaptation

    Jul 31, 2026Guanzhi Deng, Haibo Wang, Kuan Wu +5Mixture-of-Experts Language ModelsMixture of Experts

  10. MoE2^2-LoRA: When MoE Models Meet MoE-style Low-Rank Adaptation

    Jul 24, 2026Qingyu Yang, Haonan He, Minglei Li +4Mixture-of-Experts Language ModelsLow-Rank Adaptation

  11. Multi-level context Modeling for consistent expert selection in Mixture-of-Experts

    Jul 17, 2026Shuhan Huang, Naifan Zhang, Yuanbo Tang +2Mixture-of-Experts Language ModelsExpert Routing

  12. PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization

    Jul 17, 2026Yuchen Yang, Yifan Zhao, Anisha Dasgupta +1LLM QuantizationMixture-of-Experts Language Models

  13. Loop the Loopies!

    Jul 17, 2026Zitian Gao, Yilong Chen, Yihao Xiao +4Mixture-of-Experts Language ModelsRecurrent Transformers

  14. A Sovereign, Open-Source Foundation Model for German and English

    Jul 10, 2026The Soofi-Team, :, Benedikt Droste +29Language Model PretrainingMultilingual Language Models

  15. It Takes a MAESTRO To Prune Bad Experts

    Jul 9, 2026Palaash Goel, Ayush Maheshwari, Tanmoy ChakrabortyMixture-of-Experts PruningLLM Pruning

  16. Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs

    Jul 5, 2026Akhiad Bercovich, Talor Abramovich, Daniel Afrimi +67Mixture-of-Experts PruningLLM Pruning

  17. EPnG: Adaptive Expert Prune-and-Grow for Parameter-Efficient MoE Fine-tuning

    Jul 2, 2026Ahin Lee, Sehyun Yun, Taesik GongMixture-of-Experts PruningMixture-of-Experts Language Models