Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token. However, traditional training paradigms enforce a static choice of top-k experts, which converts a continuous routing distribution into a rigid step function. This constraint introduces a brittle boundary where highly competitive experts are arbitrarily separated into full-supervision and zero-feedback zones based on minor score fluctuations. To address this issue, we propose Elastic Expert Routing, which stochastically samples the active expert budget from a localized discrete distribution centered at k. Over multiple training iterations, this mechanism softens the sharp threshold into a gradual probability distribution. Because the sampling neighborhood remains symmetric, this approach matches the expected computational cost of deterministic training, while preserving the inference budget. Extensive experiments demonstrate the efficacy of our method on both supervised fine-tuning and from-scratch pretraining settings. During supervised fine-tuning, elastic routing improves downstream macro-averages on OLMoE-1B-7B and Qwen3-30B-A3B by +0.84 and +2.02 points, respectively. In addition, in from-scratch pretraining, it outperforms the static top-k baseline by 1.6 points on average across downstream tasks.
Figures & tables
Figure 1: Routing boundaries of static top- k and Elastic Expert Routing.
Task
Metric
OLMoE-1B-7B
Qwen3-30B-A3B
6--10
5--11
7--9
top-8
6--10
5--11
7--9
top-8
MMLU
Accuracy
51.25
50.91
50.14
50.47
78.49
78.40
78.17
78.24
AGIEval (En)
Accuracy
23.96
23.44
22.48
23.29
52.28
51.90
52.44
51.74
GPQA (Main)
Accuracy
26.34
26.79
26.12
25.00
39.06
38.39
37.28
37.72
BBH
Exact Match
35.28
36.05
34.54
34.86
63.29
57.69
55.68
49.29
GSM8K
Exact Match
21.46
21.38
19.86
20.09
82.03
82.94
82.71
81.43
Table 1: Main benchmark results across two sparse MoE backbones. Avg(9) is the unweighted average over the nine reported tasks. Bold marks the best method within each backbone. All elastic checkpoints are evaluated with the same static top-8 inference rule as their corresponding top-8 baseline.
Dataset
top-8
6--10
7--9
5--11
LAMBADA PPL ↓
54.2
42.7
44.9
48.7
LAMBADA
31.3
33.1
32.2
31.7
HellaSwag
41.9
43.7
43.5
42.2
PIQA
68.4
68.6
68.3
66.9
WinoGrande
49.3
52.6
53.0
53.4
CommonsenseQA
19.4
20.1
19.8
20.2
Table 2: From-scratch pretraining validation under matched architecture, data, token budget, and expected active expert count. PPL denotes perplexity, where lower is better. All metrics except LAMBADA perplexity are percentages; Avg excludes perplexity.
Figure 2: Layerwise diagnostic comparing 6--10 against static top-8 . We plot mean pairwise JS divergence across tasks. The shaded region marks the internal block L3–L7 used by the layerwise variants.
Variant
MMLU
AGI
GPQA
BBH
GSM
Min
MBPP
IFE
TQA
Avg
Δ
top-8
50.47
23.29
25.00
34.86
20.09
6.40
24.20
50.46
40.32
30.57
–
6--10
51.25
23.96
26.34
35.28
21.46
6.44
25.40
51.02
41.58
31.41
+0.84
L3--L7
50.56
24.17
27.68
35.11
21.91
6.16
26.00
50.09
41.81
31.50
+0.93
L8--L12
50.90
23.39
28.12
34.77
18.65
6.26
25.60
49.72
41.00
30.94
+0.37
Table 3: Task-level depth-region ablation on OLMoE. All methods are evaluated with static top-8 inference. L3--L7 corresponds to an exploratory diagnostic-motivated variant. Only the middle diagnostic region uses the 6--10 training neighborhood, while all other layers remain static at top-8. L8--L12 applies the same elastic rule to a later block as a depth-control experiment. Bold marks the best value in each task or summary column.
Figure 3: Per-layer expert-selection disagreement between 6--10 and static top-8 . Top-1 disagreement exceeds normalized top-8 set disagreement, indicating changes in expert priority and primary assignment within a largely preserved candidate set.
Condition
Source
Target
Avg.
Target model: 6--10
Original
6--10
6--10
31.41
Swapped
top-8
6--10
31.29
Target model: top-8
Original
top-8
top-8
30.57
Swapped
6--10
top-8
30.86
Table 4: Router-swap intervention with the remaining target model parameters fixed and only the router replaced.
Figure 4: OLMoE SFT results at different inference top- k values.
Training schedule
σ
Avg(9)
6--10
1.0
31.30
6--10
1.5
31.41
6--10
3.0
30.66
Table 5: Distribution shape ablation on OLMoE SFT.
Training rule
Evaluation rule
Avg(9)
top- p
static top-8
29.72
top- p
top- p
29.70
top-9
static top-8
30.71
top-10
static top-8
30.81
6--10 , σ=1.5
static top-8
31.41
Table 6: Alternative budget controls on OLMoE SFT. The final row is the default elastic expert routing strategy from Table 1 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Item
Protocol
Schedules
Static top-8 , 7--9 , 6--10 , and 5--11 . Layer-restricted OLMoE variants apply the dynamic rule only in the specified layer block.
Compute matching
Elastic schedules are centered at eight experts, preserving the expected active-expert count during training. All main evaluations use static top-8 inference.
Generation settings follow the task-level lm-evaluation-harness configs.
Appendix
Table A.1: Shared evaluation and routing conventions.
Setting
Method
sec/iter
Rel.
Max alloc. MB
OLMoE SFT
OLMoE
top-8
3.297
1.00x
91067.81
OLMoE
6--10
3.394
1.03x
91045.26
OLMoE
7--9
3.328
1.01x
91044.89
OLMoE
5--11
3.356
1.02x
91025.12
Qwen3-30B-A3B SFT
Appendix
Table A.2: Measured training cost from the completed training runs. Qwen3-SFT runs use global batch size 16, micro batch size 1, sequence length 8192, all-to-all token dispatch, no explicit expert capacity factor, no padding to expert capacity, and static top-8 inference for all reported evaluations. Pretraining runs use global batch size 256, micro batch size 4, sequence length 4096, and 25,034 training iterations. Memory is reported as maximum allocated CUDA memory when available.
Item
Value
Total parameters
1.42B
Active parameters / token
0.36B
Transformer layers
20
Hidden size
768
Attention heads
24
GQA groups
6
Appendix
Table A.3: Architecture and data configuration for the from-scratch MoE pretraining experiment.
Task
σ=1.0
σ=1.5
σ=3.0
MMLU
50.65
51.25
50.65
AGIEval
23.62
23.96
23.31
GPQA
27.90
26.34
27.68
BBH
35.68
35.28
35.23
GSM8K
19.79
21.46
19.26
MATH
6.30
6.44
5.70
Appendix
Table A.4: Full task-level results for the distribution-shape ablation in Table 5 . All rows use the same 6--10 budget support and static top-8 evaluation; only the Gaussian scale σ changes.
Task
top- p /8
top- p /top- p
top-9
top-10
6--10
MMLU
50.06
50.06
50.66
50.25
51.25
AGIEval
23.44
23.44
23.36
23.29
23.96
GPQA
27.90
27.90
26.34
26.12
26.34
BBH
34.00
34.00
33.71
35.16
35.28
GSM8K
17.29
17.29
19.33
20.32
21.46
MATH
4.72
4.72
6.22
7.04
6.44
Appendix
Table A.5: Full task-level results for the alternative budget controls in Table 6 . All models use static top-8 evaluation unless otherwise stated; top- p /8 denotes top- p training with static top-8 evaluation.