DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition
Authors: Yuxuan Lou, Kai Yang, Geng Zhang, Yong Liu, Yang You
Organizations: School of Computing, National University of Singapore, Singapore · Department of Computer Science and Engineering, Shanghai Jiao Tong University, Shanghai, China
Mixture-of-Experts (MoE) architectures have become essential for scaling large language models, with recent work demonstrating the benefits of fine-grained expert designs. Training such models from scratch is expensive, and sparse upcycling from pre-trained dense models is an attractive alternative. However, we identify a structural pathology of fine-grained upcycling: when fine-grained experts are derived from a single source model, naive routing collapses and downstream accuracy drops to near-random (e.g., on Qwen3-1.7B, Drop-Upcycling-fine-grained reaches only 23.2% average accuracy across 15 benchmarks, essentially matching from-scratch training at 22.2%, while the same method's coarse-grained variant reaches 50.2%). We propose DivMoE, the first framework achieving fine-grained MoE Upcycling with structurally-balanced routing. DivMoE introduces domain-specialized fine-grained expert initialization, deriving experts from dense models that have undergone domain-adaptive continual pre-training, and diversity-constrained routing, a hard structural constraint guaranteeing that each token activates experts from distinct domain groups. Across two base models and 15 benchmarks, DivMoE consistently outperforms six upcycling baselines (55.6% vs. 51.6% for the strongest baseline on Qwen3-1.7B) and strictly improves over the dense base model on every benchmark after Stage 2 continual pre-training -- closing the regression gap that has plagued prior fine-grained upcycling. After supervised fine-tuning on a public reasoning mixture, our 12B-parameter DivMoE model matches Moonlight-MoE (16B) at 64.5% average accuracy while outperforming a controlled NVIDIA-Upcycling baseline by 6.1 percentage points.
Figures & tables
Method
Domain-specialized
Fine-grained
Structural routing
experts
partitioning
constraint
Sparse Upcycling [ 20 ]
✗
✗
✗
Branch-Train-Mix [ 33 ]
✓
✗
✗
NVIDIA Upcycling [ 11 ]
✗
✓
✗
Drop-Upcycling [ 26 ]
✗
✗
✗
Drop-Upcycling-fine-grained
✗
✓
✗
Table 1 : Design-space of dense-to-MoE upcycling. ✓indicates the method occupies the design choice. Only DivMoE occupies all three columns; the empirical consequence is that fine-grained partitioning without domain diversity collapses ( Section 4.2 ).
Figure 1 : DivMoE overview. Left: a single base model is domain-adaptively pre-trained into n=4 specialists. Middle: each specialist’s FFN is vertically sliced into m=2 shards, yielding N=8 experts in n=4 domain groups. Right: the router selects the top-1 expert within each group, then top- k groups, guaranteeing k experts from k distinct domains.
Figure 2 : Routing comparison. (a) Standard top- k may select both experts from one domain. (b) Diversity-constrained routing guarantees cross-domain selection.
Method
F.G.
ARC-c
BoolQ
MMLU
TriviaQ
HellaS.
GSM8K
MATH
HumanE.
MBPP
GPQA
BBH
Avg.
From Scratch
✓
12.2
30.5
40.5
33.0
42.5
5.5
2.3
6.8
11.2
22.6
20.5
20.7
Sparse Upcycling
✗
38.5
60.2
56.8
43.5
54.5
56.5
28.5
38.0
47.5
24.5
47.5
45.1
BTX
✗
43.0
71.0
61.5
47.5
61.0
64.0
36.5
44.0
53.0
26.0
51.5
50.8
NVIDIA Upcycling
✓
42.5
70.5
61.5
47.5
60.5
62.5
35.5
44.5
53.5
25.8
51.0
50.5
Drop-Upcycling
✗
41.5
71.5
61.0
47.0
59.5
60.0
33.5
41.0
50.5
25.0
50.0
49.1
Drop-Upcycling-f
✓
14.5
32.5
42.0
33.5
41.5
7.0
3.5
7.8
12.5
22.5
22.0
21.8
Table 2 : Upcycling comparison on Qwen3-1.7B-Base after 500B Stage 2 CPT tokens (540B for BTX and DivMoE , including Stage 1). Best across upcycling methods in bold . F.G. denotes fine-grained partitioning. “Avg.” averages over the 11 benchmarks shown; full 15-benchmark results in Appendix C .
Method
F.G.
ARC-c
BoolQ
MMLU
TriviaQ
HellaS.
GSM8K
MATH
HumanE.
MBPP
GPQA
BBH
Avg.
From Scratch
✓
13.0
27.5
38.0
31.0
36.0
4.0
1.5
4.5
8.0
21.8
18.0
18.5
Sparse Upcycling
✗
33.5
60.5
48.0
39.5
52.5
26.5
9.0
15.0
21.5
23.5
31.5
32.8
BTX
✗
36.0
66.0
50.5
39.5
56.0
31.5
11.0
17.5
24.5
25.0
33.0
35.5
NVIDIA Upcycling
✓
36.0
65.5
50.0
39.0
55.5
30.5
10.8
17.0
24.0
24.8
32.8
35.1
Drop-Upcycling
✗
35.0
66.0
49.5
39.0
55.0
29.5
10.5
16.5
23.5
24.5
32.5
34.7
Drop-Upcycling-f
✓
25.5
33.0
43.5
38.0
39.5
6.0
2.5
6.0
9.5
21.5
19.0
22.2
Table 3 : Upcycling comparison on Llama3.2-1B after 500B Stage 2 CPT tokens (540B for BTX and DivMoE ). Same conventions as Table 2 .
Model
Size
ARC-c
MMLU
TriviaQA
HumanE.
MBPP
GSM8K
MATH
Avg.
OLMoE
1B/7B
49.2
51.9
60.4
51.8
61.2
45.5
23.9
49.1
OLMoE + Tülu-3 SFT
1B/7B
52.4
55.8
63.5
53.5
62.8
56.8
31.2
53.7
DeepSeek-V2-Lite
2.4B/16B
52.1
58.3
65.1
29.9
43.2
41.1
17.1
43.8
NVIDIA Up. + Tülu-3 SFT
2.4B/5.4B
56.3
64.2
67.5
45.8
54.2
66.5
38.8
56.2
Moonlight-MoE
2.4B/16B
65.5
70.0
66.3
48.1
63.8
77.4
45.3
62.3
DivMoE
3.6B/12B
63.1
71.5
73.3
51.4
60.0
74.4
47.5
63.0
Table 4 : Comparison with open-source MoE models on a representative 7-benchmark subset. Size denotes active/total parameters. DivMoE is upcycled from Qwen3-4B-Base; all baselines use released checkpoints. Rows marked “+Tülu-3 SFT” apply the same fine-tuning to a baseline. Best in bold , second-best underlined . Full 8-benchmark results (including ARC-e) in Appendix D .
Table 7
Figure 3 : Expert routing patterns across layers and datasets. Sparse Upcycling exhibits routing collapse with few experts dominating. DivMoE without the constraint already shows improved utilization. With the constraint, routing is balanced across domain groups preserving dataset-specific preferences: GSM8K activates more math experts, while HumanEval activates more code experts.
Figure 4 : Effect of expert granularity ( m=k ) on Qwen3-1.7B-Base. Medium granularity ( m=k=2 ) provides the best balance between training-loss reduction and downstream accuracy.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Qwen3-4B
Qwen3-1.7B
Llama3.2-1B
Hidden size
2560
2048
2048
FFN intermediate size
6912
5632
8192
Num layers
36
28
16
Num attention heads
32
16
32
Num KV heads
4
4
8
Vocab size
151,936
151,936
128,256
Appendix
Table 7 : Model architecture configurations. DivMoE parameters are derived from each base model with expansion factor n=4 .
Table 10 : Evaluation format and shot count per benchmark.
Method
F.G.
ARC-c
ARC-e
BoolQ
COPA
MMLU
OBQA
TriviaQ
HellaS.
SQuAD2
GSM8K
MATH
HumanE.
MBPP
GPQA
BBH
Avg.
Reference rows (not upcycling baselines)
Qwen3-1.7B-Base
–
43.2
72.4
72.8
76.0
62.6
38.6
48.5
63.2
32.5
65.8
38.4
46.5
54.8
26.5
53.5
53.0
Math specialist
–
39.8
60.5
64.8
68.2
59.5
34.2
37.8
56.5
25.8
76.2
48.5
40.3
47.5
27.8
49.6
49.1
Code specialist
–
36.5
55.2
60.4
64.8
54.1
32.5
34.8
53.8
23.5
52.4
30.5
57.8
61.5
24.2
44.3
45.8
Science specialist
–
46.8
73.2
70.5
73.8
61.8
42.5
44.2
61.5
30.2
60.5
37.2
39.5
48.3
31.8
52.6
51.6
Commonsense specialist
–
42.2
68.5
76.2
80.5
59.2
40.8
46.5
66.2
33.5
55.3
33.8
37.2
46.1
25.0
54.8
51.1
Appendix
Table 11 : Full upcycling comparison on Qwen3-1.7B-Base after 500B Stage 2 CPT tokens (540B for BTX and DivMoE , including Stage 1). Best across upcycling methods in bold ; reference rows (above the rule) provide context. F.G. denotes fine-grained partitioning.
Method
F.G.
ARC-c
ARC-e
BoolQ
COPA
MMLU
OBQA
TriviaQ
HellaS.
SQuAD2
GSM8K
MATH
HumanE.
MBPP
GPQA
BBH
Avg.
Llama3.2-1B (ref.)
–
34.8
63.5
65.2
68.0
49.3
32.4
38.5
55.2
28.5
28.5
10.2
16.5
23.8
24.8
32.4
38.1
Dense baseline (same data)
–
35.2
63.8
65.7
68.5
50.0
32.8
39.0
55.4
28.8
29.0
10.5
17.0
24.2
24.9
32.7
38.5
From Scratch
✓
13.0
22.5
27.5
32.0
38.0
23.0
31.0
36.0
17.5
4.0
1.5
4.5
8.0
21.8
18.0
19.9
Sparse Upcycling
✗
33.5
56.0
60.5
65.0
48.0
30.5
39.5
52.5
26.5
26.5
9.0
15.0
21.5
23.5
31.5
35.9
BTX
✗
36.0
64.5
66.0
69.0
50.5
33.0
39.5
56.0
29.0
31.5
11.0
17.5
24.5
25.0
33.0
39.1
NVIDIA Upcycling
✓
36.0
64.0
65.5
68.5
50.0
33.0
39.0
55.5
28.5
30.5
10.8
17.0
24.0
24.8
32.8
38.7
Appendix
Table 12 : Full upcycling comparison on Llama3.2-1B after 500B Stage 2 CPT tokens (540B for BTX and DivMoE ). Same conventions as Table 11 .
Model
Size
ARC-c
ARC-e
MMLU
TriviaQA
HumanE.
MBPP
GSM8K
MATH
Avg.
OLMoE
1B/7B
49.2
76.9
51.9
60.4
51.8
61.2
45.5
23.9
52.6
OLMoE + Tülu-3 SFT
1B/7B
52.4
77.2
55.8
63.5
53.5
62.8
56.8
31.2
56.7
DeepSeek-V2-Lite
2.4B/16B
52.1
70.3
58.3
65.1
29.9
43.2
41.1
17.1
47.1
NVIDIA Up. + Tülu-3 SFT
2.4B/5.4B
56.3
73.8
64.2
67.5
45.8
54.2
66.5
38.8
58.4
Moonlight-MoE
2.4B/16B
65.5
79.3
70.0
66.3
48.1
63.8
77.4
45.3
64.5
DivMoE
3.6B/12B
63.1
75.0
71.5
73.3
51.4
60.0
74.4
47.5
64.5
Appendix
Table 13 : Comparison with open-source MoE models on the full 8-benchmark set. Same conventions as Table 4 .
Initialization
ARC-c
MMLU
HellaSwag
Avg. (9 bm.)
Mean of base + all specialists (default)
45.0
64.5
63.5
58.3
Base model only
44.5
64.0
63.2
57.9
Random specialist selection
43.5
63.0
62.5
57.0
Appendix
Table 14 : Attention initialization ablation on Qwen3-1.7B-Base after 500B Stage-2 CPT. Avg. is over the 9 general-knowledge benchmarks for direct comparability with prior reports.
Figure 5 : Stage 2 learning dynamics. DivMoE achieves lower training loss and validation perplexity throughout training and converges to higher downstream performance, on both Qwen3-1.7B (top) and Llama3.2-1B (bottom).
Method
@150B
@300B
@500B
NVIDIA Upcycling
47.5
49.5
51.1
Drop-Upcycling
46.5
48.5
50.2
DivMoE
52.0
54.0
55.6
Appendix
Table 15 : Average accuracy (15 benchmarks) on Qwen3-1.7B-Base at extended checkpoints. DivMoE ’s advantage is sustained: gaps narrow only mildly between 150B and 500B, indicating an architectural rather than initialization-only benefit.
Tokens (B)
0
10
20
30
40
50
60
70
80
90
100
110
120
130
140
150
Training Loss
From Scratch
3.50
3.23
3.00
2.79
2.64
2.50
2.41
2.32
2.23
2.17
2.12
2.08
2.04
2.01
1.98
1.95
Sparse Upcycling
2.38
2.25
2.15
2.07
2.00
1.94
1.89
1.85
1.81
1.78
1.75
1.72
1.70
1.68
1.69
1.65
BTX
2.40
2.34
2.25
2.20
2.10
2.02
1.95
1.89
1.84
1.80
1.76
1.73
1.70
1.67
1.65
1.63
NVIDIA Upcycling
2.70
2.58
2.46
2.24
2.16
2.01
1.91
1.83
1.78
1.74
1.70
1.67
1.64
1.61
1.59
1.57
Drop-Upcycling
2.64
2.40
2.25
2.13
2.02
1.92
1.85
1.79
1.74
1.70
1.66
1.63
1.60
1.57
1.55
1.53
Appendix
Table 16 : Learning dynamics for Qwen3-1.7B-Base. Values reported at 10B-token intervals.
Tokens (B)
0
10
20
30
40
50
60
70
80
90
100
110
120
130
140
150
Training Loss
From Scratch
3.65
3.35
3.12
2.91
2.74
2.60
2.50
2.41
2.33
2.26
2.20
2.15
2.11
2.07
2.04
2.01
Sparse Upcycling
2.45
2.32
2.21
2.12
2.05
1.99
1.94
1.90
1.86
1.83
1.80
1.77
1.75
1.73
1.74
1.70
BTX
2.48
2.41
2.31
2.25
2.15
2.07
2.00
1.94
1.89
1.85
1.81
1.78
1.75
1.72
1.70
1.68
NVIDIA Upcycling
2.78
2.65
2.52
2.30
2.21
2.06
1.96
1.88
1.83
1.79
1.75
1.72
1.69
1.66
1.64
1.62
Drop-Upcycling
2.72
2.47
2.31
2.18
2.07
1.97
1.90
1.84
1.79
1.75
1.71
1.68
1.65
1.62
1.60
1.58
Appendix
Table 17 : Learning dynamics for Llama3.2-1B. Values reported at 10B-token intervals.
Tokens (B)
0
10
20
30
40
50
60
70
80
90
100
110
120
130
140
150
Training Loss
m=k=1
2.80
2.52
2.35
2.22
2.12
2.04
1.97
1.91
1.86
1.82
1.78
1.75
1.72
1.69
1.67
1.65
m=k=4
2.97
2.63
2.48
2.21
2.07
1.97
1.87
1.81
1.76
1.71
1.67
1.64
1.61
1.58
1.56
1.54
m=k=2
2.90
2.58
2.36
2.15
2.01
1.94
1.85
1.76
1.71
1.66
1.62
1.59
1.56
1.53
1.51
1.49
Validation Perplexity
m=k=1
16.5
12.8
10.9
9.6
8.7
8.0
7.5
7.1
6.7
6.4
6.2
6.0
5.8
5.6
5.5
5.4
Appendix
Table 18 : Effect of expert granularity on training dynamics. m=k=1 corresponds to coarse-grained MoE (4 experts), m=k=2 is the default DivMoE (8 experts), and m=k=4 represents finer-grained MoE (16 experts).