Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token. However, traditional training paradigms enforce a static choice of top-k experts, which converts a continuous routing distribution into a rigid step function. This constraint introduces a brittle boundary where highly competitive experts are arbitrarily separated into full-supervision and zero-feedback zones based on minor score fluctuations. To address this issue, we propose Elastic Expert Routing, which stochastically samples the active expert budget from a localized discrete distribution centered at k. Over multiple training iterations, this mechanism softens the sharp threshold into a gradual probability distribution. Because the sampling neighborhood remains symmetric, this approach matches the expected computational cost of deterministic training, while preserving the inference budget. Extensive experiments demonstrate the efficacy of our method on both supervised fine-tuning and from-scratch pretraining settings. During supervised fine-tuning, elastic routing improves downstream macro-averages on OLMoE-1B-7B and Qwen3-30B-A3B by +0.84 and +2.02 points, respectively. In addition, in from-scratch pretraining, it outperforms the static top-k baseline by 1.6 points on average across downstream tasks.
Figures & tables
Figure 1: Routing boundaries of static top- k and Elastic Expert Routing.
Task
Metric
OLMoE-1B-7B
Qwen3-30B-A3B
6--10
5--11
7--9
top-8
6--10
5--11
7--9
top-8
MMLU
Accuracy
51.25
50.91
50.14
50.47
78.49
78.40
78.17
78.24
AGIEval (En)
Accuracy
23.96
23.44
22.48
23.29
52.28
51.90
52.44
51.74
GPQA (Main)
Accuracy
26.34
26.79
26.12
25.00
39.06
38.39
37.28
37.72
BBH
Exact Match
35.28
36.05
34.54
34.86
63.29
57.69
55.68
49.29
GSM8K
Exact Match
21.46
21.38
19.86
20.09
82.03
82.94
82.71
81.43
Table 1: Main benchmark results across two sparse MoE backbones. Avg(9) is the unweighted average over the nine reported tasks. Bold marks the best method within each backbone. All elastic checkpoints are evaluated with the same static top-8 inference rule as their corresponding top-8 baseline.
Dataset
top-8
6--10
7--9
5--11
LAMBADA PPL ↓
54.2
42.7
44.9
48.7
LAMBADA
31.3
33.1
32.2
31.7
HellaSwag
41.9
43.7
43.5
42.2
PIQA
68.4
68.6
68.3
66.9
WinoGrande
49.3
52.6
53.0
53.4
CommonsenseQA
19.4
20.1
19.8
20.2
Table 2: From-scratch pretraining validation under matched architecture, data, token budget, and expected active expert count. PPL denotes perplexity, where lower is better. All metrics except LAMBADA perplexity are percentages; Avg excludes perplexity.
Figure 2: Layerwise diagnostic comparing 6--10 against static top-8 . We plot mean pairwise JS divergence across tasks. The shaded region marks the internal block L3–L7 used by the layerwise variants.
Variant
MMLU
AGI
GPQA
BBH
GSM
Min
MBPP
IFE
TQA
Avg
Δ
top-8
50.47
23.29
25.00
34.86
20.09
6.40
24.20
50.46
40.32
30.57
–
6--10
51.25
23.96
26.34
35.28
21.46
6.44
25.40
51.02
41.58
31.41
+0.84
L3--L7
50.56
24.17
27.68
35.11
21.91
6.16
26.00
50.09
41.81
31.50
+0.93
L8--L12
50.90
23.39
28.12
34.77
18.65
6.26
25.60
49.72
41.00
30.94
+0.37
Table 3: Task-level depth-region ablation on OLMoE. All methods are evaluated with static top-8 inference. L3--L7 corresponds to an exploratory diagnostic-motivated variant. Only the middle diagnostic region uses the 6--10 training neighborhood, while all other layers remain static at top-8. L8--L12 applies the same elastic rule to a later block as a depth-control experiment. Bold marks the best value in each task or summary column.
Figure 3: Per-layer expert-selection disagreement between 6--10 and static top-8 . Top-1 disagreement exceeds normalized top-8 set disagreement, indicating changes in expert priority and primary assignment within a largely preserved candidate set.
Condition
Source
Target
Avg.
Target model: 6--10
Original
6--10
6--10
31.41
Swapped
top-8
6--10
31.29
Target model: top-8
Original
top-8
top-8
30.57
Swapped
6--10
top-8
30.86
Table 4: Router-swap intervention with the remaining target model parameters fixed and only the router replaced.
Figure 4: OLMoE SFT results at different inference top- k values.
Training schedule
σ
Avg(9)
6--10
1.0
31.30
6--10
1.5
31.41
6--10
3.0
30.66
Table 5: Distribution shape ablation on OLMoE SFT.
Training rule
Evaluation rule
Avg(9)
top- p
static top-8
29.72
top- p
top- p
29.70
top-9
static top-8
30.71
top-10
static top-8
30.81
6--10 , σ=1.5
static top-8
31.41
Table 6: Alternative budget controls on OLMoE SFT. The final row is the default elastic expert routing strategy from Table 1 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Item
Protocol
Schedules
Static top-8 , 7--9 , 6--10 , and 5--11 . Layer-restricted OLMoE variants apply the dynamic rule only in the specified layer block.
Compute matching
Elastic schedules are centered at eight experts, preserving the expected active-expert count during training. All main evaluations use static top-8 inference.
Generation settings follow the task-level lm-evaluation-harness configs.
Appendix
Table A.1: Shared evaluation and routing conventions.
Setting
Method
sec/iter
Rel.
Max alloc. MB
OLMoE SFT
OLMoE
top-8
3.297
1.00x
91067.81
OLMoE
6--10
3.394
1.03x
91045.26
OLMoE
7--9
3.328
1.01x
91044.89
OLMoE
5--11
3.356
1.02x
91025.12
Qwen3-30B-A3B SFT
Appendix
Table A.2: Measured training cost from the completed training runs. Qwen3-SFT runs use global batch size 16, micro batch size 1, sequence length 8192, all-to-all token dispatch, no explicit expert capacity factor, no padding to expert capacity, and static top-8 inference for all reported evaluations. Pretraining runs use global batch size 256, micro batch size 4, sequence length 4096, and 25,034 training iterations. Memory is reported as maximum allocated CUDA memory when available.
Item
Value
Total parameters
1.42B
Active parameters / token
0.36B
Transformer layers
20
Hidden size
768
Attention heads
24
GQA groups
6
Appendix
Table A.3: Architecture and data configuration for the from-scratch MoE pretraining experiment.
Task
σ=1.0
σ=1.5
σ=3.0
MMLU
50.65
51.25
50.65
AGIEval
23.62
23.96
23.31
GPQA
27.90
26.34
27.68
BBH
35.68
35.28
35.23
GSM8K
19.79
21.46
19.26
MATH
6.30
6.44
5.70
Appendix
Table A.4: Full task-level results for the distribution-shape ablation in Table 5 . All rows use the same 6--10 budget support and static top-8 evaluation; only the Gaussian scale σ changes.
Task
top- p /8
top- p /top- p
top-9
top-10
6--10
MMLU
50.06
50.06
50.66
50.25
51.25
AGIEval
23.44
23.44
23.36
23.29
23.96
GPQA
27.90
27.90
26.34
26.12
26.34
BBH
34.00
34.00
33.71
35.16
35.28
GSM8K
17.29
17.29
19.33
20.32
21.46
MATH
4.72
4.72
6.22
7.04
6.44
Appendix
Table A.5: Full task-level results for the alternative budget controls in Table 6 . All models use static top-8 evaluation unless otherwise stated; top- p /8 denotes top- p training with static top-8 evaluation.
Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling model capacity while preserving efficient inference in large foundation models. However, most MoE models use a fixed top-k expert selection policy, assigning the same expert budget to every token even when fewer experts may be sufficient. Inference-time dynamic top-k routing can reduce computation without retraining, but existing methods often overlook the distributional shift caused by deviating from the training-time routing configuration. We show that reducing the number of activated experts consistently increases the RMS scale and variance of SMoE outputs, inducing a representation mismatch that contributes to downstream performance degradation in addition to the loss of expert capacity. To address this correctable component, we propose Layer-wise Distribution Alignment (LDA), a lightweight inference-time correction that uses layer-wise calibration statistics to align reduced-routing representations with the default configuration. Across multiple SMoE LLMs, benchmarks, and routing strategies, LDA recovers much of the performance lost induced by the distributional shift under reduced routing while preserving sparse-inference efficiency with negligible overhead.
Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized with load balancing terms and learn affinity scores through the language-model objective. However, these objectives do not provide direct alignment between routing affinities and token-level error. We introduce token-error supervision for sparse routing in two forms. The first form predicts an error score per expert. The affinity-weighted aggregate of these scores is aligned to the next-token cross-entropy loss, while the individual scores attenuate affinity before top-K selection. The second directly aligns the router's affinities to the model's objective without requiring an additional head or inference-time modification. Both formulations use the Itakura--Saito divergence or an exponential negative log-likelihood for aligning affinities and token errors. Across two sparse MoE backbones and four multiple-choice question-answering benchmarks, we evaluate both supervision mechanisms. On Granite, our method improves accuracy by approximately 2.3 percentage points on average over a parameter-matched routing baseline. With stronger supervision, the gain on ARC-Challenge reaches 2.94 points. Both mechanisms preserve the native sparse execution budget and aggregation policy. Our code is available in the supplementary materials.
Yury Nahshan, Nati Daniel, Jacob Goldberger +1
Bar-Ilan University, Ramat-Gan, Israel · NVIDIA, Israel
Mixture-of-experts (MoE) routers typically apply softmax to the router scores and keep the top-K experts, making every token use exactly K experts. Sparsity-inducing probability maps such as sparsemax, alpha-entmax and normmax can adaptively assign exact zeros to selected experts, and therefore appear to offer token-dependent expert participation, even when using the same top-K machinery. In this work, we study whether and how this sparsity survives training. We train matched 300M and 1B top-2 MoE language models with softmax, 1.5-entmax, sparsemax and 2-normmax, and find that the maps behave very differently once trained: at 1B, entmax discards 30% less probability mass than softmax while almost never dropping a selected expert, sparsemax retains the most mass, and normmax routes 21% of tokens to a single expert. These outcomes are not properties of the maps alone. Each map drops a selected expert only when the gap between the two largest scores reaches a fixed threshold, and the trained routers differ in the score distribution they learn: the entmax router learns scores with roughly half the spread of softmax's, which keeps its top-2 gaps below its threshold, while sparsemax and normmax, which share the same threshold, learn different gap distributions and hence different participation. Routers thus co-adapt their scores to the map, and a map's capacity to produce zeros does not by itself determine expert participation. While none of the sparse maps improves validation loss over softmax, they make the trained models far less sensitive to selecting more experts at inference: sparsemax trained with K=2 loses 0.02 nats when run with K=8, where softmax loses 0.58. Our results indicate that adaptive MoE routing has to be designed around the joint behavior of the probability map and the learned scores, rather than around the map alone.
Tomás Brogueira, Marcos Treviso, Miguel Couceiro
Técnico, Universidade de Lisboa · INESC-ID · Instituto de Telecomunicações +2