Mixture-of-Experts Models

Latest papers 297

All topics
CardsList
  1. DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition

    Oct 8, 2026Yuxuan Lou, Kai Yang, Geng Zhang +2Mixture-of-Experts ModelsSparse Mixture-of-Experts

  2. MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling

    Oct 6, 2026Mingyuan Zhang, Yue Bai, Zhongruo Wang +5Expert RoutingSparse Neural Networks

  3. CoRE: Learning Collaboration-Role Experts for Decentralized Collaborative Manipulation with One Policy

    Oct 6, 2026Yanan Zhou, Zhaoyan Qian, Zihao Li +3Multi-Robot SystemsMulti-Task Robotic Manipulation

  4. Structuring MoE Expert Selection for Agentic Reinforcement Learning

    Oct 5, 2026Bolian Li, Ting-Yao Hu, Cheng-Yu Hsieh +3Agentic RLExpert Routing

  5. Distributionally Robust Mixture-of-Experts Training

    Oct 5, 2026Xin Teng, Muxiao Li, Hongyi WenMixture-of-Experts Language ModelsSparse Mixture-of-Experts

  6. BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models

    Oct 5, 2026Gang Fu, Adel Javanmard, MohammadHossein Bateni +1Expert Load BalancingMixture-of-Experts Models

  7. CIPHER-MoE: Balancing Efficiency and Routing Fidelity in Trillion-Scale MoE Training

    Oct 5, 2026Jing Li, Jian Meng, Yingmeng Gao +8Expert Load BalancingExpert Routing

  8. RoMod: Temporal Routing Modulation via Mixture-of-Experts for Video Anomaly Detection

    Oct 4, 2026Chao Huang, Pengfei Wei, Benfeng Wang +5Video Anomaly DetectionMultimodal Foundation Models

  9. InstMoE: Adaptive Multimodal Routing with Specialized Experts

    Oct 4, 2026Guimin Hu, Xiang He, Yingjian Li +3Cross-Modal AlignmentMultimodal Sentiment Analysis

  10. Conditional Capacity and Routing in Mixture-of-Experts Particle Transformers

    Oct 2, 2026Kaushik Pendiyala, Haris Zia, Trevin Lee +8Mixture-of-Experts ModelsSparse Mixture-of-Experts

  11. Harnessing Domain Specialists in Multimodal Mixture-of-Experts for Efficient Adaptation

    Oct 1, 2026Damiano Marsili, Raphi Kang, Aditya Mehta +2Mixture of ExpertsMixture-of-Experts Models

  12. MoLE: Mixture of Latent Experts for Complementary Visual Reasoning

    Oct 1, 2026Yingcheng Liu, Tianyi Jiang, Yujuan Ding +5Large Vision-Language ModelsLatent Visual Reasoning

  13. SpikeMoE: Brain-Inspired Competitive Routing for Flexible Spiking Mixture-of-Experts

    Oct 1, 2026Xiaoli Liu, Yujie Liang, Jialin Li +1Spiking Neural NetworksExpert Routing

  14. From Task Mixtures to Specialized Experts

    Sep 30, 2026Hojat Allah Salehi, Mehrdad Mahdavi, Andrew Arash Mahyari +1Multi-Task LearningMixture-of-Experts Models

  15. Redundancy Meets Synergy: Dependency-aware Expert Selection for MoE via Submodular Optimization

    Sep 30, 2026Zheng Lin, Shaoke Fang, Yuxin Zhang +6Submodular OptimizationMixture-of-Experts Pruning

  16. Score the Update, Not the Token: Descent-Aligned Routing for Combinatorial LoRA Experts

    Sep 30, 2026Priya Nair, Lukas Brenner, Maya Lindqvist +5Mixture-of-Experts Language ModelsParameter-Free Mixture-of-Experts Routing

  17. OpenTSLM TeeMoE: A Unified Time-Series Language Model for Forecasting, Contextual Prediction, and Reasoning

    Sep 30, 2026Tony Chen, Timo Stoffregen, Maxwell Xu +8Temporal Reasoning in Language ModelsTime Series Forecasting

  18. HAPMoE: Heterogeneity-Aware Automatic Parallelism Planning for Mixture-of-Experts Models Training

    Sep 30, 2026Mengyuan Fan, Peizhuang Cong, Zixiao Huang +7Expert ParallelismHeterogeneous Computing

  19. ElectrolyteFM: Unifying Electrolyte Property Prediction through Cross-Property Knowledge Learning

    Sep 30, 2026Jiaxin Yu, Shuo Wang, Peng Wang +2Multi-Task LearningMaterials Property Prediction

  20. MoRA: MoE Pruning via Router Bias Learning and Expert Approximation

    Sep 30, 2026Yushuai Sun, Zikun Zhou, Lin Gao +2Mixture-of-Experts PruningLLM Compression

  21. Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

    Sep 29, 2026Yu Xu, Yuxin Zhang, Xiao Yang +7Video Diffusion ModelsVideo Generation

  22. E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models

    Sep 29, 2026Arseny Ivanov, Alexander Kolesov, Alexander Korotin +2Masked Diffusion ModelsFew-Step Diffusion Sampling

  23. MoRE: Scaling mixture of experts with hardware-aware low-rank routing

    Sep 28, 2026Honam Wong, Surbhi Goel, Enric Boix-AdseràParameter-Free Mixture-of-Experts RoutingLow-Rank Matrix Decomposition

  24. TopoEP: Topology-Aware Load Balancing for Expert-Parallel MoE Training

    Sep 28, 2026Jiacheng Zhu, Xie Zhao, Gongming Zhao +3Expert ParallelismExpert Load Balancing

  25. Role-Guided MOE for Encoder-Level Pathology Representation Learning in WSI Classification

    Sep 28, 2026Xinyu Ma, Xing Yang, Hongtao Jin +4Pathology Foundation ModelsMixture-of-Experts Models

  26. A Persistent State for Auditable Mixture-of-Experts Routing

    Sep 28, 2026Abdurrahman Javat, Allan KazakovLLM AuditingState Tracking

  27. X-MoD: Practical Scaling Laws for Sparse-Depth Routing Beyond Mixture-of-Depths

    Sep 28, 2026Bowen Dong, Yilong Fan, Tengyu Pan +6Language Model Scaling LawsMixture-of-Experts Models