cs.CVOct 1, 2026

Harnessing Domain Specialists in Multimodal Mixture-of-Experts for Efficient Adaptation

Authors: Damiano Marsili, Raphi Kang, Aditya Mehta, Pietro Perona, Georgia Gkioxari

Organizations: California Institute of Technology

Abstract

Mixture-of-Experts (MoE) architectures scale model capacity through sparse computation, routing each token through only a small subset of experts. In this work, we explore whether this sparsity gives rise to emergent intrinsic organization in multimodal MoEs. We find that experts develop strong semantic specialization across modalities and domains despite not being explicitly trained for modularity. Building on this structure, we introduce ExpertLens, a data-free method that identifies domain-specialized experts directly from pretrained model weights by decoding router weights into semantically meaningful vocabulary tokens. We leverage this specialization for efficient multimodal adaptation by selectively fine-tuning experts relevant to a target domain. Across math, medical, and remote sensing tasks, ExpertLens matches or surpasses full fine-tuning while updating only 21.7 - 47.0% of model parameters and achieving a 4.0x average training speedup, and outperforms LoRA in both adaptation performance and training efficiency. These results show that sparsity introduced for efficiency can give rise to semantic modularity that is directly useful for efficient adaptation.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. EMO: Pretraining Mixture of Experts for Emergent Modularity

    May 7, 2026Ryan Wang, Akshita Bhagia, Sewon MinMixture-Of-ExpertsExperts

  2. FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning

    May 10, 2026Xing Han, Shravan Chaudhari, Tanvi Ranade +2Mixture-Of-Experts Architectures