BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models
Organizations: Google Research · University of Southern California
Abstract
Mixture-of-experts (MoE) layers increase model capacity without a proportional increase in per-example computation. However, conventional flat routers can yield imbalanced expert utilization and treat experts as an unstructured collection, whose indices carry no topological meaning. We introduce {\bf BRANCH-MoE}, a routing architecture that places experts at the leaves of a binary decision tree of depth . At each internal node the branching probability is centered on the arrival-weighted mean score of the traffic reaching that node. This mean is estimated using an exponential moving average, which promotes utilization of both child subtrees without an auxiliary load-balancing loss. We show that this moving-average estimate admits an explicit noise-lag trade-off. We prove that for linear node maps and log-concave arrival distributions, this mechanism prevents routing-mass collapse. We further establish that, under a frozen router, an expert's execution frequency controls its stochastic-gradient convergence rate, and that confident decisions near the root bound cross-device communication when experts are assigned to devices by tree prefix. We evaluate BRANCH-MoE against Switch softmax, DeepSeek-V3 dynamic-bias, Skywork logit-normalized, and deterministic hash routing on Criteo click-through-rate prediction, Forest Covertype, HIGGS, and YearPredictionMSD, using , top- routing, and five random seeds. Our results show that hierarchical routing can preserve task quality and balanced utilization while inducing a topology that supports localized expert co-activation and reduced communication.
Figures & tables
| Symbol | Definition |
|---|---|
| Number of resident tokens and hidden dimension | |
| Routing-tree depth and number of leaf experts | |
| Binary tree, internal nodes, and leaves | |
| Matrix of learned node directions | |
| Node score anchor and induced routing bias | |
| Arrival-weighted centroid of the tokens reaching node |
| Method | Criteo (logloss) | Covertype (CE) | HIGGS (logloss) | MSD (MSE) |
|---|---|---|---|---|
| (a) Best held-out loss | ||||
| BRANCH-MoE (best map) | 0.00015 h4 | 0.0013 h8 | 0.0014 h8 | 0.0027 h8 |
| BRANCH-MoE (linear) | 0.00025 | 0.0041 | 0.0008 | 0.0035 |
| Skywork | 0.00030 lin | 0.0036 h8 | 0.0005 lin | 0.0015 h2 |
| DeepSeek-V3 | 0.00005 h8 | 0.0033 h8 | 0.0013 lin | 0.0043 h8 |
| Switch | 0.00019 lin | 0.0028 h8 | 0.0003 h4 | 0.0038 h8 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Task | Examples | Features | Seeds | Metric |
|---|---|---|---|---|---|
| Criteo (small) | CTR (binary) | M 2 2 2 The held-out split is exactly M examples (one full evaluation pass of steps at batch ). The training split is not counted directly: we infer M from the ratio of on-disk sizes ( GB train against GB held-out), which assumes a common mean record size across the two splits. All reported results depend on the training step count, not on this figure, so the corpus size should be read as approximate. | fields | Logloss / AUC | |
| Forest Covertype | -class | K | Cross-entropy / Acc. | ||
| HIGGS | Binary | K | Logloss / AUC | ||
| YearPredictionMSD | Regression | K | MSE / RMSE (yr) |
| Method | Best val. logloss | Best val. AUC | Load CV | Tree | Index |
|---|---|---|---|---|---|
| Linear node map | |||||
| Skywork | |||||
| DeepSeek-V3 | |||||
| BRANCH-MoE (ours) | |||||
| Best val. cross-entropy | ||||
|---|---|---|---|---|
| Method | linear | |||
| BRANCH-MoE (ours) | ||||
| Skywork | ||||
| Switch softmax | ||||
| HIGGS | YearPredictionMSD | |||||
|---|---|---|---|---|---|---|
| Method | Logloss | Tree | Index | MSE | Tree | Index |
| Skywork | (lin) | ( ) | ||||
| BRANCH-MoE ( ) | ||||||
| BRANCH-MoE (linear) | ||||||
| BRANCH-MoE (linear) | BRANCH-MoE ( ) | Flat routers | ||||
|---|---|---|---|---|---|---|
| Dataset | Tree | Index | Tree | Index | Tree | Index |
| Forest Covertype | – | – | ||||
| HIGGS | – | – | ||||
| YearPredictionMSD | – | – | ||||
| Criteo | – | – | ||||
| Random baseline | ||||||