Mixture-of-experts (MoE) layers increase model capacity without a proportional increase in per-example computation. However, conventional flat routers can yield imbalanced expert utilization and treat experts as an unstructured collection, whose indices carry no topological meaning. We introduce {\bf BRANCH-MoE}, a routing architecture that places E experts at the leaves of a binary decision tree of depth log2E. At each internal node the branching probability is centered on the arrival-weighted mean score of the traffic reaching that node. This mean is estimated using an exponential moving average, which promotes utilization of both child subtrees without an auxiliary load-balancing loss. We show that this moving-average estimate admits an explicit noise-lag trade-off. We prove that for linear node maps and log-concave arrival distributions, this mechanism prevents routing-mass collapse. We further establish that, under a frozen router, an expert's execution frequency controls its stochastic-gradient convergence rate, and that confident decisions near the root bound cross-device communication when experts are assigned to devices by tree prefix. We evaluate BRANCH-MoE against Switch softmax, DeepSeek-V3 dynamic-bias, Skywork logit-normalized, and deterministic hash routing on Criteo click-through-rate prediction, Forest Covertype, HIGGS, and YearPredictionMSD, using E=16, top-4 routing, and five random seeds. Our results show that hierarchical routing can preserve task quality and balanced utilization while inducing a topology that supports localized expert co-activation and reduced communication.
Figures & tables
Figure 1: Architectural progression from conventional flat routing to BRANCH-MoE. (a) A flat Top- K router scores all E experts independently in O(E) work and treats expert indices as unstructured labels. (b) A depth- D binary routing tree with the experts as leaves gives O(log2E) sequential routing depth and a structured root-to-leaf binary address per leaf ( 00 , 01 , …), but unanchored internal splits can starve entire subtrees; the full soft distribution over the leaves still requires evaluating all E−1 internal nodes. (c) BRANCH-MoE centers each internal split j on its arrival-weighted EMA score anchor ( pj=σ(fj(x)−νj) ), supporting non-starvation under the conditions of Section 4.2 without an auxiliary load-balancing loss and inducing prefix-aligned co-routing locality within hardware devices.
Symbol
Definition
B,din
Number of resident tokens and hidden dimension
D,E=2D
Routing-tree depth and number of leaf experts
T,J,E
Binary tree, internal nodes, and leaves
W∈R(2D−1)×din
Matrix of learned node directions
νj,bj=−νj
Node score anchor and induced routing bias
xˉj
Arrival-weighted centroid of the tokens reaching node j
Table 1: Principal notation.
Method
Criteo (logloss)
Covertype (CE)
HIGGS (logloss)
MSD (MSE)
(a) Best held-out loss ↓
BRANCH-MoE (best map)
0.44596± 0.00015 h4
0.2372± 0.0013 h8
0.5256± 0.0014 h8
0.6767± 0.0027 h8
BRANCH-MoE (linear)
0.44615± 0.00025
0.2545± 0.0041
0.5259± 0.0008
0.6843± 0.0035
Skywork
0.44603± 0.00030 lin
0.2406± 0.0036 h8
0.5247± 0.0005 lin
0.6757± 0.0015 h2
DeepSeek-V3
0.44605± 0.00005 h8
0.2421± 0.0033 h8
0.5264± 0.0013 lin
0.6746± 0.0043 h8
Switch
0.44623± 0.00019 lin
0.2406± 0.0028 h8
0.5276± 0.0003 h4
0.6766± 0.0038 h8
Table 2: Main results ( E=16 , top- 4 , 5 seeds; mean ± sd). (a) Best held-out loss, each router at its best of four node maps (tag), selected on the reporting split; BRANCH-MoE also with the linear map covered by the theory. Bold: best arm. Hash ignores the node map (UCI: best of its four identical runs). (b) Normalized tree distance (random pairs 0.817 ; bold: most local on UCI) / load CV of the same arms. † Not read as locality ( Section C.1 ). Per-map results: Appendix C .
Figure 2: Node-map effect. (a) Covertype loss of the learned routers vs. node map, ± sd (Hash, at 0.447 , is off-scale). (b) BRANCH-MoE’s tree distance vs. node map on UCI; dashed: random pairs ( 0.817 ); shaded: flat routers at their best maps.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Task
Examples
Features
Seeds
Metric
Criteo (small)
CTR (binary)
≈41.5 M 2 2 2 The held-out split is exactly 4.42 M examples (one full evaluation pass of 2,158 steps at batch 2048 ). The training split is not counted directly: we infer ≈37 M from the ratio of on-disk sizes ( 5.93 GB train against 0.71 GB held-out), which assumes a common mean record size across the two splits. All reported results depend on the training step count, not on this figure, so the corpus size should be read as approximate.
39 fields
5
Logloss / AUC
Forest Covertype
7 -class
581 K
54
5
Cross-entropy / Acc.
HIGGS
Binary
600 K
28
5
Logloss / AUC
YearPredictionMSD
Regression
515 K
90
5
MSE / RMSE (yr)
Appendix
Table 3: Benchmark suite. All MoE conditions use E=16 experts with top- 4 routing. “Examples” is the full corpus size (train + held-out); “Seeds” is the number of independent replications per condition.
Method
Best val. logloss ↓
Best val. AUC ↑
Load CV ↓
Tree ↓
Index ↓
Linear node map
Skywork
0.446034
0.805827
0.0040
0.8184
0.3778
±0.000298
±0.000184
DeepSeek-V3
0.446082
0.805799
0.0057
0.8188
0.3849
±0.000270
±0.000174
BRANCH-MoE (ours)
0.446149
0.805707
0.1152
0.8126
0.3843
Appendix
Table 4: Criteo, E=16 , top- 4 , capacity factor 1.25 , 5 seeds (mean ± std). Logloss is binary cross-entropy; logloss and AUC are the best held-out values over the 15 evaluations. Load CV and the normalized tree and index distances are training-time routing statistics; random baselines are 0.8167 and 0.3778 . Upper block: every router with the linear node map. Lower block: the two routers whose lowest-loss node map is an MLP, at that map (for Skywork and Switch it is the linear map). Bold marks the best logloss and AUC. At 5 vs. 5 seeds, differences below ≈3.7×10−4 in logloss or ≈2.7×10−4 in AUC are not resolvable (Welch t -test, p<0.05 ).
Best val. cross-entropy ↓
Method
linear
h=2
h=4
h=8
BRANCH-MoE (ours)
0.25451
0.25206
0.24394
0.23720
±0.00408
±0.00487
±0.00356
±0.00125
Skywork
0.25862
0.25169
0.24277
0.24062
±0.00365
±0.00324
±0.00494
±0.00358
Switch softmax
0.26150
0.25646
0.24286
0.24062
Appendix
Table 5: Forest Covertype, E=16 , top- 4 , 20 epochs, batch 512 , 5 paired seeds (mean ± std). h is the MLP node-map hidden width ( Section B.3 ). Hash ignores the node map by construction; its four arms are the same configuration run four times and serve as an exact negative control.
HIGGS
YearPredictionMSD
Method
Logloss ↓
Tree ↓
Index ↓
MSE ↓
Tree ↓
Index ↓
Skywork
0.5247 (lin)
0.812
0.370
0.6757 ( h=2 )
0.813
0.365
±0.0005
±0.0015
BRANCH-MoE ( h=8 )
0.5256
0.736
0.322
0.6767
0.749
0.324
±0.0014
±0.0027
BRANCH-MoE (linear)
0.5259
0.710
0.292
0.6843
0.718
0.301
Appendix
Table 6: HIGGS (logloss) and YearPredictionMSD (MSE), E=16 , top- 4 , 5 paired seeds (mean ± std). Each router is shown at its best node map, given in parentheses; for BRANCH-MoE we show both h=8 (best loss) and linear (best locality), since the two differ by 0.0003 on HIGGS. Tree and index distances are normalized against the 0.8167 / 0.3778 random baselines. Hash ignores the node map, so its four arms are one configuration run four times; the best is shown.
BRANCH-MoE (linear)
BRANCH-MoE ( h=8 )
Flat routers
Dataset
Tree
Index
Tree
Index
Tree
Index
Forest Covertype
0.718
0.312
0.742
0.334
0.813 – 0.819
0.371 – 0.385
HIGGS
0.710
0.292
0.736
0.322
0.812 – 0.817
0.370 – 0.380
YearPredictionMSD
0.718
0.301
0.749
0.324
0.800 – 0.817
0.365 – 0.378
Criteo
0.813
0.384
0.804
0.370
0.817 – 0.821
0.378 – 0.382
Random baseline
0.817
0.378
0.817
0.378
0.817
0.378
Appendix
Table 7: Normalized co-routing distances ( E=16 , top- 4 ; mean over 5 seeds). Flat routers: range over Switch, DeepSeek-V3, Skywork, and Hash, each at the arm shown in Table 2 . Bold: most local.
Department of Computer Science, University of Virginia, Charlottesville, USA · Department of Computer Science, University of California, Los Angeles, Los Angeles, USA