Mixture-of-Experts (MoE) layers are central to frontier language models, and recent architectures push toward more and smaller experts. In this regime, the standard linear router becomes a bottleneck: with M experts and hidden dimension h, its per-token cost Θ(Mh) dominates the MoE layer once M is large. We introduce MoRE (Mixture of Rank-reduced-routed Experts), which factorizes the router weight matrix at rank r and reduces the routing cost to O((h+M)r). We prove that rank logarithmic in M suffices for routing expressivity when the number of active experts is fixed, and is necessary up to precision factors. We also prove that logarithmic rank preserves load balance in a Gaussian memorization model, and training on a synthetic phonebook task shows that low rank does not hurt memorization. At matched active FLOPs, the factorization allows a factor of Θ(h/r) more experts. To realize this gain in wall-clock time, we design a fused Triton kernel at inference that avoids expensive memory operations on HBM. Empirically, MoRE improves memorization on the phonebook task and performance on knowledge-intensive Q&A benchmarks after pretraining, while matching reasoning ability. Code available at https://github.com/Matheart/MoRE_code.
Figures & tables
Figure 1: MoRE unlocks more scaling with the same active FLOPs. (Top) Replacing the full-rank h×M router with a low-rank factorization R1,R2 inside a fused kernel cuts routing cost from O(hM) to O(r(h+M′)) , enabling M′≫M experts at matched active FLOPs. (Bottom Left) Inference (prefilling) wall-clock time of one layer at h=512 , split into the router and the rest (setup as in Figure 4 ). The standard router’s cost grows linearly with M and dominates at large M , while MoRE’s fused low-rank router grows far more slowly. (Bottom Right) At matched FLOPs, MoRE memorizes more on the synthetic phonebook task, and the gap widens with h , as predicted by the Θ(h/r) scaling of Section 4.1 .
Figure 2: (Left) Sufficient rank does not affect memorization. (Middle, Right) MoRE with r=32 and MoE have similar load balancing at the end of training. On the Phonebook task, at M=192 , h=128 , s=8 , k=4 , L=6 . Left: maximum phonebook size MoRE memorizes against the router rank. The dashed line is MoE (standard router) at the same configuration, which MoRE already matches at r=16 . Middle, Right: routing entropy Hnorm per layer against training step for MoRE ( r=32 ) and MoE. 1 means equal load balancing and 0 means collapse onto one expert. We define Hnorm and show similar results for other setups in Appendix G.1 .
Figure 3: (Left) Fusing with top- k eliminates the need to materialize the scores for MoRE. (Middle) The low-rank router is faster than the standard router, while its cost grows with r . (Right) The gap between MoRE and MoE widens as h grows. M=4,096 , N=262,144 ( B=256 , T=1024 ), single NVIDIA B200, mean of 200 trials after 50 warm-ups. Left and middle time the router alone at h=2048 (left: r=16 , middle: k=4 ). Right times the whole layer at s=64 , k=4 , r=64 .
Figure 4: Inference (prefilling): At a fixed configuration the MoRE layer is 1.83 to 4.54× faster at M=8,192 , and at matched active FLOPs it has comparable speed to the standard MoE layer while holding h/r times more experts. Forward pass of a single layer. Left: MoE layer. Middle: MoRE layer ( r=64 ). Right: speedup TMoE/TMoRE . s=64 , k=4 , N=262,144 ( B=256 , T=1024 ), mean of 200 trials after 50 warm-ups on a single NVIDIA B200. The comparison where both layers are fused is in Figure 22 . We provide additional measurements on an A100, with similar speedups, in Appendix H.2 .
Figure 5: At fixed active FLOPs τ , MoRE memorizes more phonebook entries than MoE, and the gap widens with h as Proposition 2 predicts. Left: sweep over hidden dimension h∈{64,128,256,512} at expert width s=8 . Right: sweep over expert width s∈{8,16,32,64} at h=64 . Dotted lines are log-log linear fits. Configurations and sizes are in Tables 6 and 7 (Appendix C.1 ).
Figure 6: With active FLOPs fixed, a smaller rank allows more experts, and validation perplexity falls as the rank shrinks. The number of experts M grows from 64 to 8704 as r falls from 256 to 16, at h=512 , k=4 , L=12 and 10B training tokens.
s,h
Mo(R)E- M
Active MoE FLOPs / layer
Total Param
Val PPL ↓
TriviaQA F1 ↑
NQ-Open F1 ↑
Hella. acc_norm ↑
ARC-e acc ↑
SciQ acc ↑
Wino. acc ↑
32, 1024
MoE-128
1.57M
287.0M
14.01
16.57 ± 0.24
6.88 ± 0.32
36.91 ± 0.48
56.31 ± 1.02
79.60 ± 1.27
50.20 ± 1.41
MoRE-1536
1.57M
1.50B
11.96
20.21 ± 0.25
7.63 ± 0.34
42.28 ± 0.49
63.30 ± 0.99
80.20 ± 1.26
53.51 ± 1.40
64, 512
MoE-256
1.57M
362.9M
14.01
16.29 ± 0.24
5.72 ± 0.31
36.95 ± 0.48
57.32 ± 1.01
75.80 ± 1.36
51.46 ± 1.40
MoRE-1792
1.57M
1.91B
12.66
17.12 ± 0.23
6.88 ± 0.32
40.15 ± 0.49
60.44 ± 1.00
77.70 ± 1.32
50.28 ± 1.41
128, 512
MoE-256
2.36M
664.9M
12.81
18.09 ± 0.24
5.99 ± 0.29
40.00 ± 0.49
59.68 ± 1.01
76.30 ± 1.35
51.22 ± 1.40
MoRE-1792
2.26M
3.76B
11.92
20.79 ± 0.26
7.62 ± 0.34
42.07 ± 0.49
60.27 ± 1.00
78.70 ± 1.30
54.70 ± 1.40
Table 1: Pretraining results for MoE Transformers on 40B tokens. Each MoRE configuration ( r=64 ) matches the active FLOPs of its MoE counterpart, while total parameters are scaled up. All tasks are evaluated 0-shot, and ± denotes the standard error over test samples. Active MoE FLOPs shown are at inference and we have matched FLOPs at training too. See Appendices C.2 and C.3 for details.
Scale
Model
#Experts
Total params
Prefill (ms)
s=32 , h=1024
MoE
128
0.29 B
115.4
MoRE ( r=64 )
1,536
1.50 B
115.1
s=64 , h=512
MoE
256
0.36 B
101.4
MoRE ( r=64 )
1,792
1.91 B
95.3
s=128 , h=512
MoE
256
0.66 B
102.0
MoRE ( r=64 )
1,792
3.76 B
98.0
Table 2: MoRE prefills as fast as the MoE it matches in active FLOPs, while holding 5 to 6 times as many parameters. Prefill wall-clock (ms) for the pairs of Table 1 : the forward pass of the whole model on one NVIDIA B200 at the per-device batch the training runs use, T=1024 , mean of 200 iterations after 100 warm-ups. Training time for the same pairs is in Table 12 .
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Total
Active
L
M
k
h
s
Mixtral 8x22B Mistral AI (2024)
141B
39B
56
8
2
6144
16384
Grok-1 xAI (2024)
314B
79B
64
8
2
6144
32768
DeepSeek-V2 DeepSeek-AI (2024a)
236B
21B
60
160
6
5120
1536
DeepSeek-V3 DeepSeek-AI (2024b)
671B
37B
61
256
8
7168
2048
Llama 4 Maverick Meta AI (2025)
400B
17B
48
128
1
5120
8192
Qwen3 235B-A22B Yang et al. (2025)
235B
22B
94
128
8
4096
1536
Appendix
Table 3: Configurations of frontier MoE models. M : total experts; k : active experts per token; h : hidden dimension; s : expert FFN intermediate dimension. The trend shows greater sparsity and granularity is preferred.
Forward
Forward + backward
Tokens per step
MoRE
PEER
MoRE
PEER
4,096
1.06
5.56 ( 5.2× )
2.94
16.30 ( 5.5× )
16,384
1.23
22.06 ( 18.0× )
3.01
60.96 ( 20.3× )
65,536
1.93
87.79 ( 45.6× )
4.75
244.01 ( 51.4× )
262,144
4.92
350.86 ( 71.4× )
14.19
974.41 ( 68.7× )
Appendix
Table 4: Single-layer wall-clock time at matched active FLOPs. The ratio grows with batch size. Single-layer wall-clock in milliseconds at the matched configuration, on one NVIDIA B200 in bf16, median of 20 iterations after 5 warm-ups.
Architecture
Validation PPL
Seconds per step
Training time
MoE-256
24.26
0.321
20.4 min
MoRE-1792
24.13
0.405
25.8 min
PEER
24.17
6.135
390.1 min
Appendix
Table 5: Validation perplexity at 1 B tokens, h=512 , L=12 , four B200s, 262,144 tokens per optimizer step.
Configuration
M/M′
k
h
s
r
Max memorizable
MoE
32
4
64
8
—
10K
MoE
32
4
128
8
—
30K
MoE
32
4
256
8
—
100K
MoE
32
4
512
8
—
200K
MoRE
96
4
64
8
16
30K
MoRE
192
4
128
8
16
100K
Appendix
Table 6: Phonebook per-dot configurations for the vary- h sweep at s=8 . M ( M′ for MoRE) is the number of experts; k active per token; h hidden dim; s expert intermediate dim; r router rank (MoRE only). “Max memorizable” is the largest phonebook size at which the calibrated training run reached the early-stopping threshold.
Configuration
M/M′
k
h
s
r
Max memorizable
MoE
32
4
64
8
—
10K
MoE
64
4
64
16
—
30K
MoE
128
4
64
32
—
150K
MoE
256
4
64
64
—
500K
MoRE
96
4
64
8
16
30K
MoRE
224
4
64
16
16
100K
Appendix
Table 7: Phonebook per-dot configurations for the vary- s sweep at h=64 . Columns as in Table 6 . The (h=64,s=8) row is shared with Table 6 .
Figure 7: Expert coverage is insensitive to the router learning rate, and the optimal router learning rate is the global learning rate at every rank. We therefore do not tune the router learning rate separately, and set it equal to the global learning rate. Switch balancing loss, M=1792 , s=56 , h=512 , k=4 , L=12 , trained on 500M tokens of the mixture corpus at global learning rate η=3×10−3 . The dashed vertical line in every panel marks that global rate, and stars mark each rank’s best cell. Top: the optimum coincides with the global learning rate at all three ranks. The 5×10−4 cell is omitted at every rank because its apparent advantage at this budget does not survive to 2B tokens. Bottom: coverage holds at 1.000 at every rank up to 5×10−3 , and at r=32 across the entire grid. It falls only at 10−2 , and only for r=64 and r=128 .
Scale
Model
M
k
h
s
r
s =32, h =1024
Standard
128
4
1024
32
–
Low-rank ( r =64)
1536
4
1024
24
64
s =64, h =512
Standard
256
4
512
64
–
Low-rank ( r =64)
1792
4
512
56
64
s =128, h =512
Standard
256
4
512
128
–
Low-rank ( r =64)
1792
4
512
112
64
Appendix
Table 8: Architectural configurations for each variant in Table 1 . M : number of experts. k : experts active per token. h : model hidden dimension. s : expert FFN hidden dimension. r : router rank (Low-rank only). All variants use the GPT-2 tokenizer. Within each Scale group, the Low-rank variant scales M up while reducing s slightly so that active FLOPs approximately match the Standard reference; the (s=128,h=512) group is the exception, obtained by scaling s and s′ of (s=64,h=512) by 2 with M and M′ fixed.
(s,h)
Model
Training FLOPs
Inference FLOPs
(32, 1024)
MoE-128
4.19M
1.57M
MoRE-1536
4.06M
1.57M
(64, 512)
MoE-256
4.19M
1.57M
MoRE-1792
4.13M
1.57M
(128, 512)
MoE-256
6.55M
2.36M
MoRE-1792
6.19M
2.26M
Appendix
Table 9: Active MoE FLOPs per token per layer for the configurations in Table 1 . Training FLOPs are 18khs+14Mh for MoE and 18khs′+14(M′+h)r for MoRE. Inference FLOPs are 6khs+6Mh and 6khs′+6(M′+h)r .
Figure 8: Router learning-rate sweep on FineWeb-edu at the optimal global learning rate η=10−3 , comparing the default per-entry initialization ( left ) against the variance-matching initialization of Lemma 1 ( right ) for low-rank ranks r∈{16,32,64} . Each curve trains for 400 M FineWeb-edu tokens. The variance-matching initialization achieves lower validation perplexity across ranks, and also has stabler optimum across this range, although it might also shift leftwards when r is larger.
Figure 9: Operations in the forward pass of one MoE layer. The first expert linear projects to 2s because SwiGLU uses separate gate and up projections: SiLU(Xwgate)⊙Xwup . Expert linear layers use fused scatter-gather kernels ( Tan et al., 2024 ) that eliminate O(Nkh) memory traffic from token dispatch, making the standard router the dominant remaining overhead and motivating the low-rank factorization.
Figure 10: The two implementations measured in Table 10 . Both shard the experts M/E per GPU and hold a full copy of the router weights. In (a) the batch is replicated, so every GPU scores all N tokens and one all-reduce sums the partial outputs; the routing work is therefore duplicated E times. In (b) each GPU keeps only its own N/E tokens and scores only those, so the routing work is divided by E , at the cost of shipping each token to the GPU owning its selected expert and back with two all-to-alls.
Router
Shared phases
Layer total
E
M
MoE
MoRE
experts
comm.
MoE
MoRE
speedup
(a) expert parallelism only
2
2,048
8.20
1.68
3.86
0.61
12.68
6.15
2.1×
2
8,192
27.90
6.09
4.15
0.61
32.66
10.85
3.0×
2
16,384
45.30
11.95
4.41
0.62
50.33
16.98
3.0×
4
2,048
9.01
1.68
3.81
0.96
13.77
6.45
2.1×
Appendix
Table 10: The standard MoE router is the bottleneck of one layer under expert parallelism and dominates communication, and MoRE achieves a notable speedup by replacing it with a low-rank router. Forward pass of one MoE layer. The experts column includes the local dispatch bookkeeping, which moves nothing between devices: it turns expert ids into a sorted pair list, and in (b) it also gathers the rows that each of the other GPUs will receive. comm. is the collectives themselves, a single all-reduce of the (N,h) output in (a) against two all-to-alls in (b), the second of which also carries the (M) balancing statistic. All runs use the Switch load-balancing loss. h=512 , s=64 , k=4 , r=64 , bf16 , N=262,144 , mean of 30 trials on B200 GPUs in one node.
Figure 11: MoRE has comparable decoding time with MoE for the same configuration. Decode wall-clock of one MoE layer on the (h,M) grid of Figure 4 , one row per decode batch size B . Left: MoE layer, standard router. Middle: MoRE layer, fused low-rank router. Right: the ratio of the two. Both arms use the dispatch of this section. One NVIDIA B200 in bf16 , s=64 , k=4 , r=64 , one token per sequence, median of ten timed calls after five warm-ups. The h=4096 , M=8192 cell is blank because the expert stride overflows the int32 offsets ScatterMoE uses.
Scale
Model
#Experts
B=4
B=16
B=64
B=256
s=32 , h=1024
MoE
128
17.9
17.9
17.9
18.2
MoRE ( r=64 )
1,536
18.5
18.3
18.4
18.6
s=64 , h=512
MoE
256
18.0
17.8
17.9
18.1
MoRE ( r=64 )
1,792
18.5
18.5
18.4
18.6
s=128 , h=512
MoE
256
17.9
17.8
17.9
18.0
MoRE ( r=64 )
1,792
18.5
18.3
18.5
18.5
Appendix
Table 11: Decoding: MoRE matches MoE in wall-clock time even with a much larger number of experts. Per-step decode wall-clock (ms) for MoE/MoRE pairs used in our pretraining experiments with matched active FLOPs (Table 1 ), across decode batch size B . Each MoRE row holds 6 – 12× the experts of the MoE it is matched against, and stays within 3% of it at every scale and batch size. One NVIDIA B200 in bf16 , 12 layers, k=4 , prefill 128 , median of the last six of twelve decode steps. Both use the expert dispatch of this section.
Scale
Model
#Experts
Total params
Training (ms)
Training time ratio
s=32 , h=1024
MoE
128
0.29 B
241.0
MoRE ( r=64 )
1,536
1.50 B
312.5
1.30 ×
s=64 , h=512
MoE
256
0.36 B
165.0
MoRE ( r=64 )
1,792
1.91 B
221.1
1.34 ×
s=128 , h=512
MoE
256
0.66 B
182.7
MoRE ( r=64 )
1,792
3.76 B
259.3
1.42 ×
Appendix
Table 12: MoRE costs 15 to 42% more per training step than the MoE it matches in active FLOPs. Training wall-clock (ms) for the pairs of Table 1 , measured over the 40 B-token runs themselves, on 4× NVIDIA B200. Prefill time for the same pairs is in Table 2 .
Configuration
Model
Total params
Training (ms)
Ratio
M=1,536 , s=24 , h=1024
MoRE ( r=64 )
1.50 B
312.5
MoE
1.51 B
344.2
1.10 ×
M=1,792 , s=56 , h=512
MoRE ( r=64 )
1.91 B
221.1
MoE
1.92 B
249.6
1.13 ×
M=1,792 , s=112 , h=512
MoRE ( r=64 )
3.76 B
259.3
MoE
3.77 B
289.6
1.12 ×
Appendix
Table 13: MoRE and MoE at the same configuration: MoE is 1.10 to 1.23× slower than MoRE. Training step time (ms) of each MoRE model of Table 12 and of a standard MoE with the same number M and width s of experts, on 4× NVIDIA B200. The last pair is trained with expert parallelism (EP).
Figure 12: Flowchart for inference (prefilling) for low-rank router. The blue block is one Triton kernel, holding one BN×BM tile of ℓ at a time, so neither the (N,r) intermediate z nor the (N,M) logits reaches HBM. TopIdxs names the k experts to run for each token; a softmax over the k entries of TopVals gives the weights that combine their outputs. Nothing here depends on all M experts, which is why one sweep is enough.
Figure 13: Fused low-rank router for training. Blue blocks run inside the kernel, which holds one BN×BM tile of ℓ at a time and never forms the full N×M matrix. Purple marks what the Switch loss adds. Sweep 2 computes the quantity required for the load-balancing loss: psumi=∑nPn,i , the softmax probability of expert i summed over tokens. TopVals is the gradient of the prediction loss with respect to the k selected scores, and is sparse in M . pˉ is the gradient of the balancing loss with respect to psum , one number per expert and dense in M .
Figure 14: Similar results at a larger expert pool. Phonebook, h=64 , s=8 , M=1024 , k=4 , L=6 , panels as in Figure 2 . Left: largest phonebook memorized, searched on a 25 K grid. Middle, Right: Hnorm per layer against training step, for r=32 and for MoE, both trained on the 200 K phonebook. This configuration is off the matched-FLOP curve of Proposition 2 : it places several ranks below log2M , and both arms hold the same experts, so only the router differs.
Figure 15: Adding more experts on MoRE keeps improving quality. Validation perplexity of MoRE ( r=64 ) after 5 B training tokens, against the number of experts M . The curve is monotone across a 16× increase in the pool size.
Figure 16: The perplexity gap between the two configurations grows with the token budget. Validation perplexity of MoE-1024 minus that of MoRE-8192 on the 5,000 held-out mixture sequences, against training tokens. Positive values mean MoRE has the lower perplexity. Both models have matched active FLOPs and 1.27 B against 7.92 B total parameters. We plot the first 25 B tokens before the cosine schedule enters its annealing tail.
Figure 17: The fused MoRE layer stays faster at larger expert widths. Inference (prefilling) of a single layer at s=128 , 256 and 512 (rows). (a) MoE Layer. (b) MoRE Layer ( r=64 ). (c) speedup TMoE/TMoRE . The blank cell at s=512 ran out of memory on one B200. k=4 , r=64 , N=262,144 ( B=256 , T=1024 ), bf16 , mean of 30 trials after 10 warm-ups on a single NVIDIA B200.
Figure 18: The A100 reproduces the B200 result, with the fused MoRE layer 2.60 to 3.07× faster at M=8,192 . Inference (prefilling) of a single layer, the two arms at equal M . (a) MoE Layer. (b) MoRE Layer ( r=64 ). (c) speedup TMoE/TMoRE . Single NVIDIA A100-80GB, s=64 , k=4 , N=262,144 ( B=256 , T=1024 ), bf16 , mean of 30 trials after 10 warm-ups.
Figure 19: Fusing the low-rank router reduces peak memory by up to 14× for the whole layer. Ratio of peak memory over a forward and a backward, unfused MoRE against fused MoRE, on the same grid.
Figure 20: On both A100 and B200, fusing Standard MoE does not get much acceleration at large h . Inference (prefilling) of a single layer, both arms a standard MoE layer at the same M . (a) unfused router. (b) fused router. (c) speedup. (Top) A100, 0.86 to 2.83× . (Bottom) B200, 1.01 to 4.21× . At h=4,096 , M=8,192 fusing takes the layer from 174.8 to 202.5 ms on the A100 and from 61.1 to 60.4 ms on the B200. s=64 , k=4 , N=262,144 ( B=256 , T=1024 ), bf16 , mean of 30 trials after 10 warm-ups.
Figure 21: Fusing keeps saving time for MoRE as h grows, but stops saving time for MoE. Forward pass of a single layer, the unfused-router time minus the fused-router time, against the hidden dimension h at a fixed pool of M=4,096 experts. What fusing removes is the round trip of the (N,M) scores, which depends on N and M but not on h . The standard router gives it back because its fused kernel reloads a (BM,h) tile of the router weight for every token block. s=64 , k=4 , r=64 , N=262,144 ( B=256 , T=1024 ), bf16 , mean of 30 trials after 10 warm-ups on a single NVIDIA B200.
Figure 22: Comparing fused MoE and fused MoRE, the low-rank factorization gives speedup up to 2.82× on an A100 and 1.72× on a B200. Inference (prefilling) of a single layer, both arms fused, at the same M . (a) MoE Layer with the fused router. (b) MoRE Layer ( r=64 ) with the fused router. (c) speedup. (Top) A100, 1.00 to 2.82× . (Bottom) B200, 0.99 to 1.72× . s=64 , k=4 , N=262,144 ( B=256 , T=1024 ), bf16 , mean of 30 trials after 10 warm-ups.
Figure 23: End-to-end forward-pass breakdown at h=512 , sweeping M . Each bar decomposes the wall-time of a 12-layer transformer forward pass into Attention, Router, Experts (dispatch + grouped MLP), RMSNorm, and the Liger-fused ( Hsu et al., 2024 ) LM head + cross-entropy. (Left) Standard MoE router (unfused h×M matmul + top- k + softmax + load-balancing-loss as separate kernels): the router slice grows from 16% of total at M=1,024 to 38% at M=8,192 , becoming the largest single component and the dominant driver of the rising end-to-end cost. (Right) Fused low-rank MoRE router ( r=64 ): the router slice stays at 5–12% across the same sweep, and total wall-time grows much more slowly with M . At M=8,192 , MoRE is 30% faster end-to-end ( 108.2 ms vs. 152.8 ms), with router cost reduced from 57.9 ms to 13.4 ms. Mean of 200 trials with 50 warm-ups, 95% C.I. on a single NVIDIA B200.
Figure 24: The benefit of fusing the MoRE router grows with the batch size. Forward-pass time of the low-rank router, unfused and fully fused, with the speedup of the fused kernel above each pair. h=2048 , M=4,096 , r=16 , k=4 , T=1024 , bf16 , mean of 200 trials after 50 warm-ups on a single NVIDIA B200.
Mixture-of-Experts (MoE) architectures decouple model capacity from computational cost, yet incur high memory footprints as parameters grow linearly with the number of experts. Recurrent Transformers achieve parameter efficiency by reusing layer weights, but typically lack the capacity for competitive language modeling. We propose Mixture of Reused Experts (MoRE), a hybrid that shares expert pools across groups of adjacent layers. Each layer retains its own router but selects from a larger shared pool, expanding the diversity of routing combinations without additional parameters. To enable shared experts to distinguish between layers, we introduce lightweight learnable depth embeddings that condition each layer's input before routing. Experiments across three model scales (114M-1.15B parameters) show that MoRE consistently achieves lower perplexity and stronger downstream performance than standard MoEs and state-of-the-art weight-sharing architectures at matched compute and parameter budgets, with only minimal modifications to existing MoE implementations.
Eric S. Qiu, Utku Umur Acikalin, Justin Lovelace +4
Department of Computer Science, Cornell University
Mixture-of-Experts (MoE) models scale neural networks by conditionally activating a small subset of experts, where the router plays a central role in determining expert specialization and overall model performance. However, many modern MoE systems still adopt linear routers in raw high-dimensional representation spaces, where representation mismatch, angular concentration, and scale-sensitive scoring can jointly undermine routing discriminability and stable expert specialization. In this work, we propose Low-rank & Lipschitz-controlled Routing (L2R), a unified routing framework that reshapes both the routing space and scoring geometry. L2R performs expert assignment in a shared low-rank latent routing space and introduces Saturated Inner-Product Scoring (SIPS) to explicitly control the Lipschitz behavior of routing functions, yielding smoother and more stable routing geometry. In addition, L2R incorporates a parameter-efficient multi-anchor routing mechanism to enhance expert expressiveness. Experiments on an OLMoE-based language MoE model and a ViT-based ImageNet setting show improved overall performance in both domains; OLMoE diagnostics further show improved routing geometry and expert discrimination.
Mixture-of-Experts (MoE) models decouple parameter count from per-token compute, but deployment still requires hosting every expert in memory. Recent theory shows that experts whose router weights change least during fine-tuning can be pruned with provable accuracy preservation, yet the guarantee assumes full fine-tuning. We show that the signal can be elicited through a brief parameter-efficient adaptation. We fine-tune with a lightweight adapter, rank experts by the induced router change, and prune the least-changed experts in one shot. On Mixtral-8×7B-Instruct, router-only LoRA trains 0.002% of parameters and retains 27.54% MMLU-Pro accuracy with half the experts removed, against roughly 16% for magnitude and random pruning. Signal quality improves monotonically with adapter size, reaching 28.76%, and declines as adaptation spreads beyond the router. Under their shared budget, IA3 reaches 28.04% while Houlsby reaches 25.39%. The criterion transfers to Qwen1.5-MoE fine-tuned for mathematical reasoning, retaining 49.7% mean accuracy over eleven benchmarks with half the experts removed. Structural pruning reduces memory by 49% and per-token latency by 37%. Lightweight router sensitivity therefore makes provably motivated, task-conditioned expert pruning practical at scale.
Ali Janati, Kaoutar El Maghraoui, Chengke Zou +2
Data Science Institute, Columbia University, New York, NY, USA · Department of Computer Science, Columbia University, New York, NY, USA