cs.LGOct 5, 2026

BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models

Authors: Gang Fu, Adel Javanmard, MohammadHossein Bateni, Vahab Mirrokni

Organizations: Google Research · University of Southern California

Abstract

Mixture-of-experts (MoE) layers increase model capacity without a proportional increase in per-example computation. However, conventional flat routers can yield imbalanced expert utilization and treat experts as an unstructured collection, whose indices carry no topological meaning. We introduce {\bf BRANCH-MoE}, a routing architecture that places EE experts at the leaves of a binary decision tree of depth log⁡2E\log_2 E. At each internal node the branching probability is centered on the arrival-weighted mean score of the traffic reaching that node. This mean is estimated using an exponential moving average, which promotes utilization of both child subtrees without an auxiliary load-balancing loss. We show that this moving-average estimate admits an explicit noise-lag trade-off. We prove that for linear node maps and log-concave arrival distributions, this mechanism prevents routing-mass collapse. We further establish that, under a frozen router, an expert's execution frequency controls its stochastic-gradient convergence rate, and that confident decisions near the root bound cross-device communication when experts are assigned to devices by tree prefix. We evaluate BRANCH-MoE against Switch softmax, DeepSeek-V3 dynamic-bias, Skywork logit-normalized, and deterministic hash routing on Criteo click-through-rate prediction, Forest Covertype, HIGGS, and YearPredictionMSD, using E=16E=16, top-44 routing, and five random seeds. Our results show that hierarchical routing can preserve task quality and balanced utilization while inducing a topology that supports localized expert co-activation and reduced communication.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Self-Routing: Parameter-Free Expert Routing from Hidden States

    Apr 1, 2026Jama Hussein Mohamud, Drew Wagner, Mirco RavanelliMixture-Of-ExpertsLarge Language Model Routing

  2. ProbMoE: Differentiable Probabilistic Routing for Mixture-of-Experts

    Jun 1, 2026Heng Zhao, Zilei Shao, Guy Van den Broeck +1Large Language Model RoutingProbabilistic Inference

  3. Elbow-Based MoE Routing: A Training-Free Inference Time Plugin for Expert Selection

    Aug 5, 2026Robin Pan, Raymond Liu, Daniel Fang +2Mixture-Of-Expert InferenceMixture-Of-Experts