DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition
Authors: Yuxuan Lou, Kai Yang, Geng Zhang, Yong Liu, Yang You
Organizations: School of Computing, National University of Singapore, Singapore · Department of Computer Science and Engineering, Shanghai Jiao Tong University, Shanghai, China
Mixture-of-Experts (MoE) architectures have become essential for scaling large language models, with recent work demonstrating the benefits of fine-grained expert designs. Training such models from scratch is expensive, and sparse upcycling from pre-trained dense models is an attractive alternative. However, we identify a structural pathology of fine-grained upcycling: when fine-grained experts are derived from a single source model, naive routing collapses and downstream accuracy drops to near-random (e.g., on Qwen3-1.7B, Drop-Upcycling-fine-grained reaches only 23.2% average accuracy across 15 benchmarks, essentially matching from-scratch training at 22.2%, while the same method's coarse-grained variant reaches 50.2%). We propose DivMoE, the first framework achieving fine-grained MoE Upcycling with structurally-balanced routing. DivMoE introduces domain-specialized fine-grained expert initialization, deriving experts from dense models that have undergone domain-adaptive continual pre-training, and diversity-constrained routing, a hard structural constraint guaranteeing that each token activates experts from distinct domain groups. Across two base models and 15 benchmarks, DivMoE consistently outperforms six upcycling baselines (55.6% vs. 51.6% for the strongest baseline on Qwen3-1.7B) and strictly improves over the dense base model on every benchmark after Stage 2 continual pre-training -- closing the regression gap that has plagued prior fine-grained upcycling. After supervised fine-tuning on a public reasoning mixture, our 12B-parameter DivMoE model matches Moonlight-MoE (16B) at 64.5% average accuracy while outperforming a controlled NVIDIA-Upcycling baseline by 6.1 percentage points.
Figures & tables
Method
Domain-specialized
Fine-grained
Structural routing
experts
partitioning
constraint
Sparse Upcycling [ 20 ]
✗
✗
✗
Branch-Train-Mix [ 33 ]
✓
✗
✗
NVIDIA Upcycling [ 11 ]
✗
✓
✗
Drop-Upcycling [ 26 ]
✗
✗
✗
Drop-Upcycling-fine-grained
✗
✓
✗
Table 1 : Design-space of dense-to-MoE upcycling. ✓indicates the method occupies the design choice. Only DivMoE occupies all three columns; the empirical consequence is that fine-grained partitioning without domain diversity collapses ( Section 4.2 ).
Figure 1 : DivMoE overview. Left: a single base model is domain-adaptively pre-trained into n=4 specialists. Middle: each specialist’s FFN is vertically sliced into m=2 shards, yielding N=8 experts in n=4 domain groups. Right: the router selects the top-1 expert within each group, then top- k groups, guaranteeing k experts from k distinct domains.
Figure 2 : Routing comparison. (a) Standard top- k may select both experts from one domain. (b) Diversity-constrained routing guarantees cross-domain selection.
Method
F.G.
ARC-c
BoolQ
MMLU
TriviaQ
HellaS.
GSM8K
MATH
HumanE.
MBPP
GPQA
BBH
Avg.
From Scratch
✓
12.2
30.5
40.5
33.0
42.5
5.5
2.3
6.8
11.2
22.6
20.5
20.7
Sparse Upcycling
✗
38.5
60.2
56.8
43.5
54.5
56.5
28.5
38.0
47.5
24.5
47.5
45.1
BTX
✗
43.0
71.0
61.5
47.5
61.0
64.0
36.5
44.0
53.0
26.0
51.5
50.8
NVIDIA Upcycling
✓
42.5
70.5
61.5
47.5
60.5
62.5
35.5
44.5
53.5
25.8
51.0
50.5
Drop-Upcycling
✗
41.5
71.5
61.0
47.0
59.5
60.0
33.5
41.0
50.5
25.0
50.0
49.1
Drop-Upcycling-f
✓
14.5
32.5
42.0
33.5
41.5
7.0
3.5
7.8
12.5
22.5
22.0
21.8
Table 2 : Upcycling comparison on Qwen3-1.7B-Base after 500B Stage 2 CPT tokens (540B for BTX and DivMoE , including Stage 1). Best across upcycling methods in bold . F.G. denotes fine-grained partitioning. “Avg.” averages over the 11 benchmarks shown; full 15-benchmark results in Appendix C .
Method
F.G.
ARC-c
BoolQ
MMLU
TriviaQ
HellaS.
GSM8K
MATH
HumanE.
MBPP
GPQA
BBH
Avg.
From Scratch
✓
13.0
27.5
38.0
31.0
36.0
4.0
1.5
4.5
8.0
21.8
18.0
18.5
Sparse Upcycling
✗
33.5
60.5
48.0
39.5
52.5
26.5
9.0
15.0
21.5
23.5
31.5
32.8
BTX
✗
36.0
66.0
50.5
39.5
56.0
31.5
11.0
17.5
24.5
25.0
33.0
35.5
NVIDIA Upcycling
✓
36.0
65.5
50.0
39.0
55.5
30.5
10.8
17.0
24.0
24.8
32.8
35.1
Drop-Upcycling
✗
35.0
66.0
49.5
39.0
55.0
29.5
10.5
16.5
23.5
24.5
32.5
34.7
Drop-Upcycling-f
✓
25.5
33.0
43.5
38.0
39.5
6.0
2.5
6.0
9.5
21.5
19.0
22.2
Table 3 : Upcycling comparison on Llama3.2-1B after 500B Stage 2 CPT tokens (540B for BTX and DivMoE ). Same conventions as Table 2 .
Model
Size
ARC-c
MMLU
TriviaQA
HumanE.
MBPP
GSM8K
MATH
Avg.
OLMoE
1B/7B
49.2
51.9
60.4
51.8
61.2
45.5
23.9
49.1
OLMoE + Tülu-3 SFT
1B/7B
52.4
55.8
63.5
53.5
62.8
56.8
31.2
53.7
DeepSeek-V2-Lite
2.4B/16B
52.1
58.3
65.1
29.9
43.2
41.1
17.1
43.8
NVIDIA Up. + Tülu-3 SFT
2.4B/5.4B
56.3
64.2
67.5
45.8
54.2
66.5
38.8
56.2
Moonlight-MoE
2.4B/16B
65.5
70.0
66.3
48.1
63.8
77.4
45.3
62.3
DivMoE
3.6B/12B
63.1
71.5
73.3
51.4
60.0
74.4
47.5
63.0
Table 4 : Comparison with open-source MoE models on a representative 7-benchmark subset. Size denotes active/total parameters. DivMoE is upcycled from Qwen3-4B-Base; all baselines use released checkpoints. Rows marked “+Tülu-3 SFT” apply the same fine-tuning to a baseline. Best in bold , second-best underlined . Full 8-benchmark results (including ARC-e) in Appendix D .
Table 7
Figure 3 : Expert routing patterns across layers and datasets. Sparse Upcycling exhibits routing collapse with few experts dominating. DivMoE without the constraint already shows improved utilization. With the constraint, routing is balanced across domain groups preserving dataset-specific preferences: GSM8K activates more math experts, while HumanEval activates more code experts.
Figure 4 : Effect of expert granularity ( m=k ) on Qwen3-1.7B-Base. Medium granularity ( m=k=2 ) provides the best balance between training-loss reduction and downstream accuracy.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Qwen3-4B
Qwen3-1.7B
Llama3.2-1B
Hidden size
2560
2048
2048
FFN intermediate size
6912
5632
8192
Num layers
36
28
16
Num attention heads
32
16
32
Num KV heads
4
4
8
Vocab size
151,936
151,936
128,256
Appendix
Table 7 : Model architecture configurations. DivMoE parameters are derived from each base model with expansion factor n=4 .
Table 10 : Evaluation format and shot count per benchmark.
Method
F.G.
ARC-c
ARC-e
BoolQ
COPA
MMLU
OBQA
TriviaQ
HellaS.
SQuAD2
GSM8K
MATH
HumanE.
MBPP
GPQA
BBH
Avg.
Reference rows (not upcycling baselines)
Qwen3-1.7B-Base
–
43.2
72.4
72.8
76.0
62.6
38.6
48.5
63.2
32.5
65.8
38.4
46.5
54.8
26.5
53.5
53.0
Math specialist
–
39.8
60.5
64.8
68.2
59.5
34.2
37.8
56.5
25.8
76.2
48.5
40.3
47.5
27.8
49.6
49.1
Code specialist
–
36.5
55.2
60.4
64.8
54.1
32.5
34.8
53.8
23.5
52.4
30.5
57.8
61.5
24.2
44.3
45.8
Science specialist
–
46.8
73.2
70.5
73.8
61.8
42.5
44.2
61.5
30.2
60.5
37.2
39.5
48.3
31.8
52.6
51.6
Commonsense specialist
–
42.2
68.5
76.2
80.5
59.2
40.8
46.5
66.2
33.5
55.3
33.8
37.2
46.1
25.0
54.8
51.1
Appendix
Table 11 : Full upcycling comparison on Qwen3-1.7B-Base after 500B Stage 2 CPT tokens (540B for BTX and DivMoE , including Stage 1). Best across upcycling methods in bold ; reference rows (above the rule) provide context. F.G. denotes fine-grained partitioning.
Method
F.G.
ARC-c
ARC-e
BoolQ
COPA
MMLU
OBQA
TriviaQ
HellaS.
SQuAD2
GSM8K
MATH
HumanE.
MBPP
GPQA
BBH
Avg.
Llama3.2-1B (ref.)
–
34.8
63.5
65.2
68.0
49.3
32.4
38.5
55.2
28.5
28.5
10.2
16.5
23.8
24.8
32.4
38.1
Dense baseline (same data)
–
35.2
63.8
65.7
68.5
50.0
32.8
39.0
55.4
28.8
29.0
10.5
17.0
24.2
24.9
32.7
38.5
From Scratch
✓
13.0
22.5
27.5
32.0
38.0
23.0
31.0
36.0
17.5
4.0
1.5
4.5
8.0
21.8
18.0
19.9
Sparse Upcycling
✗
33.5
56.0
60.5
65.0
48.0
30.5
39.5
52.5
26.5
26.5
9.0
15.0
21.5
23.5
31.5
35.9
BTX
✗
36.0
64.5
66.0
69.0
50.5
33.0
39.5
56.0
29.0
31.5
11.0
17.5
24.5
25.0
33.0
39.1
NVIDIA Upcycling
✓
36.0
64.0
65.5
68.5
50.0
33.0
39.0
55.5
28.5
30.5
10.8
17.0
24.0
24.8
32.8
38.7
Appendix
Table 12 : Full upcycling comparison on Llama3.2-1B after 500B Stage 2 CPT tokens (540B for BTX and DivMoE ). Same conventions as Table 11 .
Model
Size
ARC-c
ARC-e
MMLU
TriviaQA
HumanE.
MBPP
GSM8K
MATH
Avg.
OLMoE
1B/7B
49.2
76.9
51.9
60.4
51.8
61.2
45.5
23.9
52.6
OLMoE + Tülu-3 SFT
1B/7B
52.4
77.2
55.8
63.5
53.5
62.8
56.8
31.2
56.7
DeepSeek-V2-Lite
2.4B/16B
52.1
70.3
58.3
65.1
29.9
43.2
41.1
17.1
47.1
NVIDIA Up. + Tülu-3 SFT
2.4B/5.4B
56.3
73.8
64.2
67.5
45.8
54.2
66.5
38.8
58.4
Moonlight-MoE
2.4B/16B
65.5
79.3
70.0
66.3
48.1
63.8
77.4
45.3
64.5
DivMoE
3.6B/12B
63.1
75.0
71.5
73.3
51.4
60.0
74.4
47.5
64.5
Appendix
Table 13 : Comparison with open-source MoE models on the full 8-benchmark set. Same conventions as Table 4 .
Initialization
ARC-c
MMLU
HellaSwag
Avg. (9 bm.)
Mean of base + all specialists (default)
45.0
64.5
63.5
58.3
Base model only
44.5
64.0
63.2
57.9
Random specialist selection
43.5
63.0
62.5
57.0
Appendix
Table 14 : Attention initialization ablation on Qwen3-1.7B-Base after 500B Stage-2 CPT. Avg. is over the 9 general-knowledge benchmarks for direct comparability with prior reports.
Figure 5 : Stage 2 learning dynamics. DivMoE achieves lower training loss and validation perplexity throughout training and converges to higher downstream performance, on both Qwen3-1.7B (top) and Llama3.2-1B (bottom).
Method
@150B
@300B
@500B
NVIDIA Upcycling
47.5
49.5
51.1
Drop-Upcycling
46.5
48.5
50.2
DivMoE
52.0
54.0
55.6
Appendix
Table 15 : Average accuracy (15 benchmarks) on Qwen3-1.7B-Base at extended checkpoints. DivMoE ’s advantage is sustained: gaps narrow only mildly between 150B and 500B, indicating an architectural rather than initialization-only benefit.
Tokens (B)
0
10
20
30
40
50
60
70
80
90
100
110
120
130
140
150
Training Loss
From Scratch
3.50
3.23
3.00
2.79
2.64
2.50
2.41
2.32
2.23
2.17
2.12
2.08
2.04
2.01
1.98
1.95
Sparse Upcycling
2.38
2.25
2.15
2.07
2.00
1.94
1.89
1.85
1.81
1.78
1.75
1.72
1.70
1.68
1.69
1.65
BTX
2.40
2.34
2.25
2.20
2.10
2.02
1.95
1.89
1.84
1.80
1.76
1.73
1.70
1.67
1.65
1.63
NVIDIA Upcycling
2.70
2.58
2.46
2.24
2.16
2.01
1.91
1.83
1.78
1.74
1.70
1.67
1.64
1.61
1.59
1.57
Drop-Upcycling
2.64
2.40
2.25
2.13
2.02
1.92
1.85
1.79
1.74
1.70
1.66
1.63
1.60
1.57
1.55
1.53
Appendix
Table 16 : Learning dynamics for Qwen3-1.7B-Base. Values reported at 10B-token intervals.
Tokens (B)
0
10
20
30
40
50
60
70
80
90
100
110
120
130
140
150
Training Loss
From Scratch
3.65
3.35
3.12
2.91
2.74
2.60
2.50
2.41
2.33
2.26
2.20
2.15
2.11
2.07
2.04
2.01
Sparse Upcycling
2.45
2.32
2.21
2.12
2.05
1.99
1.94
1.90
1.86
1.83
1.80
1.77
1.75
1.73
1.74
1.70
BTX
2.48
2.41
2.31
2.25
2.15
2.07
2.00
1.94
1.89
1.85
1.81
1.78
1.75
1.72
1.70
1.68
NVIDIA Upcycling
2.78
2.65
2.52
2.30
2.21
2.06
1.96
1.88
1.83
1.79
1.75
1.72
1.69
1.66
1.64
1.62
Drop-Upcycling
2.72
2.47
2.31
2.18
2.07
1.97
1.90
1.84
1.79
1.75
1.71
1.68
1.65
1.62
1.60
1.58
Appendix
Table 17 : Learning dynamics for Llama3.2-1B. Values reported at 10B-token intervals.
Tokens (B)
0
10
20
30
40
50
60
70
80
90
100
110
120
130
140
150
Training Loss
m=k=1
2.80
2.52
2.35
2.22
2.12
2.04
1.97
1.91
1.86
1.82
1.78
1.75
1.72
1.69
1.67
1.65
m=k=4
2.97
2.63
2.48
2.21
2.07
1.97
1.87
1.81
1.76
1.71
1.67
1.64
1.61
1.58
1.56
1.54
m=k=2
2.90
2.58
2.36
2.15
2.01
1.94
1.85
1.76
1.71
1.66
1.62
1.59
1.56
1.53
1.51
1.49
Validation Perplexity
m=k=1
16.5
12.8
10.9
9.6
8.7
8.0
7.5
7.1
6.7
6.4
6.2
6.0
5.8
5.6
5.5
5.4
Appendix
Table 18 : Effect of expert granularity on training dynamics. m=k=1 corresponds to coarse-grained MoE (4 experts), m=k=2 is the default DivMoE (8 experts), and m=k=4 represents finer-grained MoE (16 experts).
Mixture-of-Experts (MoE) models scale capacity without proportional compute cost and have become a key architecture for frontier large language models (LLMs). Yet domain-specific post-training inherits an expert pool shaped by mixed-domain pre-training: a substantial subset of experts contributes little on the target domain, and standard supervised fine-tuning (SFT) leaves the composition of this pool unchanged. We propose a simple, budget-preserving pipeline that realigns the expert pool to the target domain before fine-tuning. Given a target domain, we (1) prune the experts with lowest domain-aligned saliency, (2) regrow the expert pool to its original size through perturbation-based expert expansion, and (3) apply standard SFT. The resulting model preserves the original expert count, parameter count, and inference cost. With a single frozen recipe and no per-domain hyperparameter tuning, UMoE consistently improves over direct sft across two MoE architectures (Qwen3-30B-A3B and Qwen3.5-35B-A3B), five domains (math, code, science, tool-use, and agentic coding), and 12 benchmarks. Representative improvements are 3.4 points in math average accuracy, 6.0 points on SWE-bench Verified. On a strong in-house math corpus, direct sft already surpasses Qwen3-30B-A3B-Thinking (82.81 vs.\ 81.06), yet UMoE further raises the average to 84.17, an additional 1.36 points, demonstrating robustness to a substantially stronger SFT regime. Data-scaling experiments further show that the gain persists as training data grows. Analysis reveals that the direct-SFT model allocates substantial routed-expert compute to a low-saliency subset that can be removed post hoc with little average degradation; UMoE turns this redundant capacity into useful domain capacity and achieves lower training loss, with gains spanning all difficulty levels in downstream evaluation.
Mixture-of-Experts (MoE) has become the dominant architecture for scaling large language models: frontier models routinely decouple total parameters from per-token computation through sparse expert routing. Scaling laws show that under fixed active computation, model quality scales predictably with total parameters, and MoEs realize this by increasing expert count. However, training large MoEs is expensive, as memory requirements and inter-device communication both scale with total parameter count. We propose expert upcycling, a method for progressively expanding MoE capacity by increasing the number of experts during continued pre-training (CPT). Given a trained E-expert model, the upcycling operator constructs an mE-expert model through expert duplication and router extension while holding top-K routing fixed, preserving per-token inference cost. Duplication provides a warm initialization: the expanded model inherits the source checkpoint's learned representations, starting from a substantially lower loss than random initialization. Subsequent CPT then breaks the symmetry among duplicated experts to drive specialization. We formalize the upcycling operator and develop a theoretical framework decomposing the quality gap into a capacity term and an initialization term. We further introduce utility-based expert selection, which uses gradient-based importance scores to guide non-uniform duplication, more than tripling gap closure when CPT is limited. In our 7B-13B total parameter experiments, the upcycled model matches the fixed-size baseline on validation loss while saving 32% of GPU hours. Comprehensive ablations across model scales, activation ratios, MoE architectures, and training budgets yield a practical recipe for deploying expert upcycling, establishing it as a principled, compute-efficient alternative to training large MoE models from scratch.
We present Marco-MoE, a suite of fully open multilingual sparse Mixture-of-Experts (MoE) models. Marco-MoE features a highly sparse design in which only around 5% of the total parameters are activated per input token. This extreme sparsity, combined with upcycling from dense models, enables efficient pre-training on 5T tokens. Our models surpass similarly-sized competitors on English and multilingual benchmarks, achieving a best-in-class performance-to-compute ratio. We further post-train these models to create Marco-MoE-\textsc{Instruct} variants, which surpass the performance of competing models possessing 3--14× more activated parameters. Our analysis reveals that Marco-MoE learns structured expert activation patterns shared across related languages, while maintaining highly specialized utilization for linguistically isolated ones. We further show that Marco-MoE allows for scalable language expansion without the interference typical of dense models. To support the community, we disclose our full training datasets, recipes, and model weights.