Mixture-of-Experts Language Models

Latest papers 147

All topics
CardsList
  1. dMoE: dLLMs with Learnable Block Experts

    May 29, 2026Sicheng Feng, Zigeng Chen, Gongfan Fang +2Mixture-of-Experts Language ModelsMemory-Efficient Inference

  2. Leveraging Routing Dynamics in Mixture-of-Experts Models for Efficient Language Adaptation

    May 28, 2026Aditi Khandelwal, Marius Mosbach, Verna Dankers +2Multilingual Language ModelsContinual Learning for LLMs

  3. Understanding Safety-Sensitive Expert Behavior in Mixture-of-Experts LLMs

    May 28, 2026Zhibo Zhang, Yuxi Li, Zhen Ouyang +2Mixture-of-Experts Language ModelsLLM Auditing

  4. RA-MoE: Routing-Aligned Fine-Tuning for Multilingual Adaptation of Mixture-of-Experts Models

    May 27, 2026Guanzhi Deng, Kuan Wu, Haibo Wang +6Mixture-of-Experts Language ModelsLLM Fine-Tuning

  5. PrunePath: Towards Highly Structured Sparse Language Models

    May 27, 2026Zhexuan Gu, Zixun Fu, Yancheng YuanMixture-of-Experts PruningLLM Pruning

  6. Pruning and Distilling Mixture-of-Experts into Dense Language Models

    May 27, 2026Junhyuck Kim, Jihun Yun, Haechan Kim +3Mixture-of-Experts PruningLLM Pruning

  7. Extracting Small Translation Specialists from LLMs by Aggressively Pruning Experts

    May 27, 2026Liu O. Martin, Lucas Bandarkar, Nanyun PengMixture-of-Experts PruningLLM Pruning

  8. FPMoE: A Sparse Mixture-of-Experts Approach to Functional Code Generation

    May 27, 2026Loc Pham, Lang Hong Nguyet Anh, Thanh Le-CongMixture-of-Experts Language ModelsCode Generation

  9. L2Rec: Towards Dual-View Understanding of LLMs for Personalized Recommendation

    May 26, 2026Pingjun Pan, Tingting Zhou, Peiyao Lu +3Mixture-of-Experts Language ModelsLLM Personalization

  10. RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry

    May 24, 2026Bo Lv, Zhiheng Xu, KeDong Xiu +4Mixture-of-Experts Language ModelsLLM Auditing

  11. Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs

    May 23, 2026Bo Li, Tianyu Dong, Shaolin Zhu +1Multilingual Language ModelsMixture-of-Experts Language Models

  12. Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts

    May 22, 2026Md Nurul Absar SiddikyMixture-of-Experts Language ModelsExpert Routing

  13. BitsMoE: Cost-Aware Bit Allocation in Spectral Space for MoE LLM Quantization

    May 22, 2026Jiayu Zhao, Zihan Teng, Minhao Fan +4Mixed-Precision QuantizationLLM Quantization

  14. GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs

    May 21, 2026Jianing Deng, Song Wang, Dongwei Wang +4Mixed-Precision QuantizationLLM Quantization

  15. Dynamic Mixture of Latent Memories for Self-Evolving Agents

    May 21, 2026Dianzhi Yu, Vireo Zhang, Hongru Wang +7Continual Learning for LLMsMixture-of-Experts Language Models

  16. A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAMΔΔ Integration into Upcycled MoE

    May 18, 2026Hao Zhou, Tianhao Li, Zhijun Wang +6Multilingual Language ModelsMixture-of-Experts Language Models

  17. CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning

    May 18, 2026Yang Liu, Toan Nguyen, Flora D. SalimContinual Learning for LLMsVision-Language Models

  18. Mixture of Experts for Low-Resource LLMs

    May 17, 2026Ori Bar Joseph, Smadar Arvatz, Noam Kayzer +2Mixture-of-Experts Language ModelsExpert Routing

  19. Scalable Knowledge Editing for Mixture-of-Experts LLMs via Tensor-Structured Updates

    May 15, 2026Roman Maksimov, Vladimir Aletov, Dmitry Bylinkin +3Knowledge EditingMixture-of-Experts Language Models

  20. ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems

    May 12, 2026Wenyong Zhou, Yuannuo Feng, Yizhe Chen +6Compute-in-MemoryMixture-of-Experts Language Models

  21. Slicing and Dicing: Configuring Optimal Mixtures of Experts

    May 12, 2026Margaret Li, Sneha Kudugunta, Danielle Rothermel +1Mixture-of-Experts Language ModelsMixture of Experts

  22. HELLoRA: Hot Experts Layer-Level Low-Rank Adaptation for Mixture-of-Experts Models

    May 11, 2026Jia Wei, Zhonghao Zhang, Ping Chen +5Mixture-of-Experts Language ModelsFine-Tuning

  23. Sparse Layers are Critical to Scaling Looped Language Models

    May 9, 2026Ryan Lee, Jacob Biloki, Edward J. Hu +1Mixture-of-Experts Language ModelsRecurrent Transformers