cs.LGSep 28, 2026

MoRE: Scaling mixture of experts with hardware-aware low-rank routing

Authors: Honam Wong, Surbhi Goel, Enric Boix-Adserà

Organizations: University of Pennsylvania · The Wharton School, University of Pennsylvania

Abstract

Mixture-of-Experts (MoE) layers are central to frontier language models, and recent architectures push toward more and smaller experts. In this regime, the standard linear router becomes a bottleneck: with MM experts and hidden dimension hh, its per-token cost Θ(Mh)Θ(Mh) dominates the MoE layer once MM is large. We introduce MoRE (Mixture of Rank-reduced-routed Experts), which factorizes the router weight matrix at rank rr and reduces the routing cost to O((h+M)r)O((h + M)r). We prove that rank logarithmic in MM suffices for routing expressivity when the number of active experts is fixed, and is necessary up to precision factors. We also prove that logarithmic rank preserves load balance in a Gaussian memorization model, and training on a synthetic phonebook task shows that low rank does not hurt memorization. At matched active FLOPs, the factorization allows a factor of Θ(h/r)Θ(h/r) more experts. To realize this gain in wall-clock time, we design a fused Triton kernel at inference that avoids expensive memory operations on HBM. Empirically, MoRE improves memorization on the phonebook task and performance on knowledge-intensive Q&A benchmarks after pretraining, while matching reasoning ability. Code available at https://github.com/Matheart/MoRE_code.

Figures & tables

Appendix figures & tables29 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MoRE: Mixture of Reused Experts

    Sep 16, 2026Eric S. Qiu, Utku Umur Acikalin, Justin Lovelace +4Mixture-Of-ExpertsExperts

  2. Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models

    Aug 8, 2026Ali Janati, Kaoutar El Maghraoui, Chengke Zou +2ExpertsLow-Rank Adaptation Framework