Revisiting Diffusion Fine-Tuning for Unsupervised Domain Adaptation
Authors: Xuan Qi, Yi Wei, Daniele Berardini, Vito Paolo Pastore, Vittorio Murino
Organizations: AI for Good (AIGO), Istituto Italiano di Tecnologia, Genoa, Italy · DITEN, University of Genoa, Genoa, Italy · National Key Laboratory of Novel Software Technology, Nanjing University, China · School of Intelligence Science and Technology, Nanjing University, China · MaLGa, DIBRIS, University of Genoa, Genoa, Italy · Department of Computer Science, University of Verona, Verona, Italy
Diffusion-based unsupervised domain adaptation (UDA) improves cross-domain transfer by generating target-specific synthetic data for downstream adaptation. Existing methods are largely designed for single-target adaptation: when a model trained on one labeled source domain must be adapted to multiple unlabeled target domains, they typically require separate diffusion fine-tuning for each source--target pair, causing training, storage, and deployment costs to grow with the number of targets. In this paper, we study multi-target data generation for diffusion-based UDA, where a single source-guided diffusion fine-tuning process is reused to generate target-specific synthetic data for multiple target domains. We propose MUSE (Multi-target UDA-oriented Synthesis with Efficient diffusion fine-tuning), a decoupled adaptation framework that separates source-supervised semantic adaptation from target-specific style adaptation. MUSE uses a shared semantic branch updated by labeled source data and target-private style branches specialized to individual target domains, enabling target-specific generation while avoiding repeated source-guided fine-tuning for each target. Experiments on standard UDA benchmarks show that MUSE achieves a stronger accuracy--efficiency trade-off than repeated per-target diffusion adaptation, reducing diffusion fine-tuning cost while improving average target-domain accuracy. The project page is available at https://xuanqi99.github.io/MUSE/.
Figures & tables
Figure 1: Overview of the proposed MUSE pipeline. MUSE performs branch-decoupled diffusion adaptation with a source-supervised shared semantic branch and target-private style branches, then activates the corresponding target branch to generate target-specific bridge samples for downstream UDA.
Figure 2: Bridge generation on miniDomainNet. (a) DDIM inversion translates a Real-domain lion into Clipart, Painting, and Sketch targets, i.e., R→{C,P,S} . (b) Gaussian-noise class-conditional generation synthesizes broccoli samples for Painting, Real, and Sketch targets under C→{P,R,S} .
Method
Diffusion Gen. Protocol
A → W
D → W
W → D
A → D
D → A
W → A
Avg.
ERM
—
77.07
96.60
99.20
81.08
64.11
64.01
80.35
DANN [ 8 ]
—
89.85
97.95
99.90
83.26
73.28
73.75
86.33
CDAN [ 25 ]
—
92.42
98.62
100.00
91.44
74.61
72.80
88.32
AFN [ 54 ]
—
91.82
98.77
100.00
95.12
72.43
70.71
88.14
MDD [ 58 ]
—
93.55
98.66
100.00
93.92
75.29
73.95
89.23
SDAT [ 39 ]
—
91.32
98.83
100.00
95.25
76.97
73.19
89.26
Table 1: Transfer accuracy (%) on Office-31. Avg. is over six tasks. Pairwise denotes per-source–target diffusion generation; Multi-target denotes one source-guided generator reused across targets.
Method
Diffusion Gen. Protocol
Ar → Cl
Ar → Pr
Ar → Rw
Cl → Ar
Cl → Pr
Cl → Rw
Pr → Ar
Pr → Cl
Pr → Rw
Rw → Ar
Rw → Cl
Rw → Pr
Avg.
ERM
—
44.06
67.12
74.26
53.26
61.96
64.54
51.91
38.90
72.94
64.51
43.84
75.39
59.39
DANN [ 8 ]
—
52.53
62.57
73.20
56.89
67.02
68.34
58.37
54.14
78.31
70.78
60.76
80.57
65.29
CDAN [ 25 ]
—
54.21
72.18
78.29
61.97
71.43
72.39
62.96
55.68
80.68
74.71
61.22
83.68
69.12
AFN [ 54 ]
—
52.58
72.42
76.96
64.90
71.14
72.91
64.08
51.29
77.83
72.21
57.46
82.09
67.99
MDD [ 58 ]
—
56.37
75.53
79.17
62.95
73.21
73.55
62.56
54.86
79.49
73.84
61.45
84.06
69.75
SDAT [ 39 ]
—
58.20
77.46
81.35
66.06
76.45
76.41
63.70
56.69
82.49
76.02
62.09
85.24
71.85
Table 2: Transfer accuracy (%) on Office-Home. Avg. is over twelve tasks. Pairwise denotes per-source–target diffusion generation; Multi-target denotes one source-guided generator reused across targets.
Method
Diffusion Gen. Protocol
C → P
C → R
C → S
P → C
P → R
P → S
R → C
R → P
R → S
S → C
S → P
S → R
Avg.
ERM
—
39.48
53.27
42.93
49.55
68.18
41.99
49.30
55.52
37.35
54.60
45.33
53.08
49.22
DANN [ 8 ]
—
45.94
56.12
49.40
50.72
65.61
50.07
55.15
60.55
49.95
58.54
54.64
58.99
54.64
AFN [ 54 ]
—
49.23
60.11
51.11
55.60
70.59
51.78
55.84
60.41
47.46
60.69
56.36
62.28
56.79
CDAN [ 25 ]
—
47.99
58.50
51.17
56.36
68.71
53.01
61.15
62.85
53.44
60.89
55.90
60.88
57.57
MDD [ 58 ]
—
48.53
61.75
52.32
59.74
70.62
55.43
62.18
62.22
54.04
63.07
58.55
64.50
59.41
SDAT [ 39 ]
—
50.97
62.42
53.91
60.57
69.97
55.85
64.39
64.83
55.86
64.07
59.43
64.28
60.55
Table 3: Transfer accuracy (%) on miniDomainNet. Avg. is over twelve tasks. Pairwise denotes per-source–target diffusion generation; Multi-target denotes one source-guided generator reused across targets.
Dataset
# Domains
Terra Time (h) ↓
MUSE Time (h) ↓
Speedup ↑
Office-31
3
96.78
51.74
1.87 ×
Office-Home
4
196.36
80.79
2.43 ×
miniDomainNet
4
197.68
81.04
2.44 ×
Table 4: SDXL fine-tuning time on one A100 80GB GPU. The maximum number of fine-tuning steps is 10000 for both Terra and MUSE. # Domains denotes the total number of domains in the benchmark.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Method
M
Peak allocated VRAM per process (GiB)
Aggregate VRAM across M concurrent processes (GiB)
Cumulative fine-tuning time (h)
Office-31
Independent fine-tuning (Terra)
2
32.433
64.865
96.78
Office-31
MUSE
2
27.849
27.849
51.74
Office-Home
Independent fine-tuning (Terra)
3
32.433
97.298
196.36
Office-Home
MUSE
3
27.835
27.835
80.79
miniDomainNet
Independent fine-tuning (Terra)
3
32.433
97.298
197.68
miniDomainNet
MUSE
3
27.835
27.835
81.04
Appendix
Table 5: Memory and training-time comparison for multi-target diffusion adaptation. Peak allocated VRAM is measured per training process. Aggregate VRAM for independent fine-tuning corresponds to concurrently executing the M target-specific processes and is obtained by summing their unrounded per-process peaks. Cumulative fine-tuning time is measured over the complete benchmark protocol.
Jtar=λmbLtgt(mb)+λorthLorth+λcoreLcore
Appendix
Algorithm 1 MUSE Diffusion Adaptation
B(m)=Binv(m)∪Bcls(m)
Appendix
Algorithm 2 Target-Specific Bridge Generation and Downstream Adaptation
Method
Diffusion Gen. Protocol
A → W
D → W
W → D
A → D
D → A
W → A
ERM
—
0.11
0.00
0.00
1.22
0.15
0.11
DANN [ 8 ]
—
1.34
0.06
0.08
0.68
0.65
0.39
CDAN [ 25 ]
—
1.75
0.18
0.00
1.19
0.79
0.45
AFN [ 54 ]
—
0.63
0.07
0.00
0.53
0.50
0.32
MDD [ 58 ]
—
1.00
0.15
0.00
0.10
0.68
0.18
SDAT [ 39 ]
—
1.83
0.12
0.00
1.03
0.67
0.34
Appendix
Table 6: Standard deviations of transfer accuracy (%) on Office-31 over three runs. Pairwise denotes per-source–target diffusion generation; Multi-target denotes one source-guided generator reused across targets.
Method
Diffusion Gen. Protocol
Ar → Cl
Ar → Pr
Ar → Rw
Cl → Ar
Cl → Pr
Cl → Rw
Pr → Ar
Pr → Cl
Pr → Rw
Rw → Ar
Rw → Cl
Rw → Pr
ERM
—
0.25
0.26
0.42
0.17
0.20
0.15
0.07
0.17
0.05
0.34
0.33
0.01
DANN [ 8 ]
—
0.44
0.72
0.38
0.02
0.30
0.39
0.58
0.47
0.59
0.84
0.14
0.51
CDAN [ 25 ]
—
0.25
0.62
0.22
0.37
0.58
0.30
0.57
0.36
0.16
0.33
0.23
0.35
AFN [ 54 ]
—
0.16
0.30
0.06
0.23
0.31
0.14
0.32
0.15
0.02
0.19
0.18
0.22
MDD [ 58 ]
—
0.51
0.32
0.06
0.24
0.73
0.41
0.36
0.53
0.24
0.03
0.09
0.11
SDAT [ 39 ]
—
0.51
0.44
0.24
0.13
0.41
0.01
1.46
0.40
0.11
0.46
0.19
0.29
Appendix
Table 7: Standard deviations of transfer accuracy (%) on Office-Home over three runs. Pairwise denotes per-source–target diffusion generation; Multi-target denotes one source-guided generator reused across targets.
Method
Diffusion Gen. Protocol
C → P
C → R
C → S
P → C
P → R
P → S
R → C
R → P
R → S
S → C
S → P
S → R
ERM
—
0.34
0.72
0.51
0.28
0.89
0.45
0.63
0.21
0.77
0.55
0.82
0.39
DANN [ 8 ]
—
0.41
0.88
0.22
0.56
0.74
0.31
0.68
0.49
0.85
0.27
0.61
0.53
AFN [ 54 ]
—
0.73
0.29
0.84
0.46
0.62
0.38
0.71
0.25
0.59
0.87
0.42
0.65
CDAN [ 25 ]
—
0.24
0.67
0.35
0.81
0.52
0.78
0.43
0.26
0.69
0.57
0.83
0.31
MDD [ 58 ]
—
0.58
0.82
0.47
0.64
0.21
0.76
0.32
0.55
0.89
0.41
0.68
0.23
SDAT [ 39 ]
—
0.37
0.61
0.85
0.29
0.44
0.73
0.56
0.81
0.34
0.67
0.25
0.79
Appendix
Table 8: Standard deviations of transfer accuracy (%) on miniDomainNet over three runs. Pairwise denotes per-source–target diffusion generation; Multi-target denotes one source-guided generator reused across targets.
Method
Ar → Cl
Ar → Pr
Ar → Rw
Cl → Ar
Cl → Pr
Cl → Rw
Pr → Ar
Pr → Cl
Pr → Rw
Rw → Ar
Rw → Cl
Rw → Pr
Avg.
Δ Avg.
MCC+MUSE w/o Lorth
64.23
80.38
81.92
72.56
81.39
80.63
71.16
63.46
82.92
74.28
65.67
84.09
75.22
-1.31
MCC+MUSE
64.31
83.44
82.37
73.14
84.44
83.50
72.90
63.02
83.47
75.37
66.58
85.82
76.53
0.00
ELS+MUSE w/o Lorth
64.91
80.67
82.61
72.03
81.51
81.27
71.61
63.51
83.71
74.17
65.86
85.11
75.58
-1.99
ELS+MUSE
66.51
84.09
82.62
74.14
83.74
82.96
74.61
64.73
85.52
75.61
68.81
87.45
77.57
0.00
SSRT+MUSE w/o Lorth
77.73
90.27
91.23
85.54
90.83
91.65
86.16
79.93
92.33
87.35
80.14
91.32
87.04
-1.11
SSRT+MUSE
79.09
90.43
92.44
87.81
92.66
92.69
87.44
79.42
93.45
88.52
80.79
93.01
88.15
0.00
Appendix
Table 9: Ablation on the orthogonality regularizer on Office-Home. Δ Avg. is computed relative to the corresponding full MUSE model under the same downstream learner.
Method
Ar → Cl
Ar → Pr
Ar → Rw
Cl → Ar
Cl → Pr
Cl → Rw
Pr → Ar
Pr → Cl
Pr → Rw
Rw → Ar
Rw → Cl
Rw → Pr
Avg.
Δ Avg.
MCC+MUSE w/ target-updated shared branch
64.05
79.01
81.54
70.09
79.91
79.73
72.02
61.97
83.77
73.92
63.98
84.68
74.56
-1.97
MCC+MUSE
64.31
83.44
82.37
73.14
84.44
83.50
72.90
63.02
83.47
75.37
66.58
85.82
76.53
0.00
ELS+MUSE w/ target-updated shared branch
64.77
79.54
82.35
71.78
80.49
80.22
71.26
62.43
83.25
74.25
64.35
85.01
74.98
-2.59
ELS+MUSE
66.51
84.09
82.62
74.14
83.74
82.96
74.61
64.73
85.52
75.61
68.81
87.45
77.57
0.00
SSRT+MUSE w/ target-updated shared branch
77.46
90.81
91.05
87.27
90.38
91.65
86.21
78.97
92.56
87.31
80.27
91.52
87.12
-1.03
SSRT+MUSE
79.09
90.43
92.44
87.81
92.66
92.69
87.44
79.42
93.45
88.52
80.79
93.01
88.15
0.00
Appendix
Table 10: Ablation on freezing the shared branch during target-conditioned updates on Office-Home. Δ Avg. is computed relative to the corresponding full MUSE model under the same downstream learner.
Method
Ar → Cl
Ar → Pr
Ar → Rw
Cl → Ar
Cl → Pr
Cl → Rw
Pr → Ar
Pr → Cl
Pr → Rw
Rw → Ar
Rw → Cl
Rw → Pr
Avg.
Δ Avg.
MCC+MUSE w/o adaptive sampling
63.19
83.75
81.61
71.66
83.41
81.85
73.08
62.10
82.10
74.53
67.02
84.61
75.74
-0.79
MCC+MUSE
64.31
83.44
82.37
73.14
84.44
83.50
72.90
63.02
83.47
75.37
66.58
85.82
76.53
0.00
ELS+MUSE w/o adaptive sampling
65.68
82.82
81.98
74.51
82.63
82.20
73.23
64.14
84.06
74.69
69.03
86.41
76.78
-0.79
ELS+MUSE
66.51
84.09
82.62
74.14
83.74
82.96
74.61
64.73
85.52
75.61
68.81
87.45
77.57
0.00
SSRT+MUSE w/o adaptive sampling
77.75
89.51
90.68
86.70
91.18
91.85
85.51
79.73
92.20
86.85
79.81
91.47
86.94
-1.21
SSRT+MUSE
79.09
90.43
92.44
87.81
92.66
92.69
87.44
79.42
93.45
88.52
80.79
93.01
88.15
0.00
Appendix
Table 11: Ablation on adaptive target-domain sampling on Office-Home. Δ Avg. is computed relative to the corresponding full MUSE model under the same downstream learner.
Generation mode
Time per image (s) ↓
DDIM inversion source-to-target generation
4.23
Pure-noise class-conditional generation
2.08
Appendix
Table 12: Per-image generation time of the two bridge-generation modes.
Method
Ar → Cl
Ar → Pr
Ar → Rw
Cl → Ar
Cl → Pr
Cl → Rw
Pr → Ar
Pr → Cl
Pr → Rw
Rw → Ar
Rw → Cl
Rw → Pr
Avg.
Δ Avg.
MCC+MUSE, Binv
61.26
78.95
78.68
69.02
79.16
78.13
69.31
61.76
79.12
73.51
65.23
84.22
73.20
-3.33
MCC+MUSE, Bcls
60.12
77.72
80.92
71.21
81.95
80.17
70.12
59.12
81.04
72.48
62.53
80.58
73.16
-3.37
MCC+MUSE, Binv∪Bcls
64.31
83.44
82.37
73.14
84.44
83.50
72.90
63.02
83.47
75.37
66.58
85.82
76.53
0.00
ELS+MUSE, Binv
62.96
81.25
79.34
69.84
79.83
79.34
69.02
61.32
80.85
74.21
66.51
85.22
74.14
-3.43
ELS+MUSE, Bcls
61.39
80.26
81.52
71.15
81.69
80.51
70.37
59.34
82.11
73.18
63.18
82.91
73.97
-3.60
ELS+MUSE, Binv∪Bcls
66.51
84.09
82.62
74.14
83.74
82.96
74.61
64.73
85.52
75.61
68.81
87.45
77.57
0.00
Appendix
Table 13: Ablation on the two components of the generated bridge set on Office-Home. Δ Avg. is computed relative to the corresponding full bridge set under the same downstream learner.
Table 14: Additional analysis of SDXL-prior data generation on Office-Home. Prompt-only denotes SDXL generation without source–target diffusion fine-tuning; Pairwise and Multi-target follow the main experimental protocol.
Figure 3: t-SNE visualizations of feature distributions for representative categories from Office-31, Office-Home, and miniDomainNet. We compare source-domain samples, target-domain samples, DDIM-inversion adapted source samples, and generated target-domain samples.
Figure 4: Qualitative visualization of DDIM-inversion-based source-to-target generation on Office-31. Rows correspond to the 6 transfer tasks A → W, D → W, W → D, A → D, D → A, and W → A. Each row shows 5 classes. Within each source–generated pair, the left image is the source-domain input and the right image is the corresponding target-style image generated by the target-specific MUSE branch.
Figure 5: Qualitative visualization of class-conditional samples generated from Gaussian noise on Office-31. Rows correspond to the 6 transfer tasks A → W, D → W, W → D, A → D, D → A, and W → A. Each row shows generated target-style samples from five different classes.
Figure 6: Qualitative visualization of DDIM-inversion-based source-to-target generation on Office-Home. Rows correspond to the 12 transfer tasks Ar → Cl, Ar → Pr, Ar → Rw, Cl → Ar, Cl → Pr, Cl → Rw, Pr → Ar, Pr → Cl, Pr → Rw, Rw → Ar, Rw → Cl, and Rw → Pr. Each row shows 5 classes. Within each pair, the left image is the source-domain input and the right image is the generated target-style image.
Figure 7: Qualitative visualization of class-conditional samples generated from Gaussian noise on Office-Home. Rows correspond to the 12 transfer tasks Ar → Cl, Ar → Pr, Ar → Rw, Cl → Ar, Cl → Pr, Cl → Rw, Pr → Ar, Pr → Cl, Pr → Rw, Rw → Ar, Rw → Cl, and Rw → Pr. Each row shows generated target-style samples from 5 different classes.
Figure 8: Qualitative visualization of DDIM-inversion-based source-to-target generation on miniDomainNet. Rows correspond to the 12 transfer tasks C → P, C → R, C → S, P → C, P → R, P → S, R → C, R → P, R → S, S → C, S → P, and S → R. Each row shows 10 classes. Within each pair, the left image is the source-domain input and the right image is the corresponding target-style image generated by the target-specific MUSE branch.
Figure 9: Qualitative visualization of class-conditional samples generated from Gaussian noise on miniDomainNet. Rows correspond to the 12 transfer tasks C → P, C → R, C → S, P → C, P → R, P → S, R → C, R → P, R → S, S → C, S → P, and S → R. Each row shows generated target-style samples from 10 different classes.