cs.LGOct 7, 2026

Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression

Authors: Tianxiao Cao, Jiahe Shao, Yuning Qiu, Kyohei Atarashi, Hisashi Kashima, Qibin Zhao

Organizations: Kyoto University · The University of Tokyo · RIKEN AIP

Abstract

Mixture-of-Experts (MoE) large language models decouple capacity from compute through sparse routing, but their large parameter count creates storage and serving challenges. We analyze three MoE compression families: expert pruning, expert merging, and weight reconstruction, and derive structural error bounds showing that pruning and merging can incur non-vanishing errors tied to routing and expert heterogeneity. In contrast, weight reconstruction avoids these structural costs by preserving expert structure and routing. Motivated by the analysis, we propose Shared Low-rank Basis Factorization (SLBF), a data-free weight reconstruction method that uses rank-kk bases shared among experts, enabling richer cross-expert sharing, faster convergence, and lower reconstruction error. A post-hoc gauge fixing removes redundant parameters at no representational cost. Across five MoE architectures spanning 16B to 122B parameters, SLBF consistently outperforms methods from all three compression families.

Explore similar work

CardsList
  1. Shape Mutating Expert Compression:LorExperts and BTExperts

    Aug 7, 2026Inesh Chakrabarti, Sourjya Roy, Bowen Bao +3Mixture-Of-Experts Large Language ModelsExperts

  2. PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference

    Nov 6, 2025Yushu Zhao, Zheng Wang, Minjia ZhangMixture-Of-Expert InferenceMixture-Of-Experts