Revisiting Diffusion Fine-Tuning for Unsupervised Domain Adaptation
Authors: Xuan Qi, Yi Wei, Daniele Berardini, Vito Paolo Pastore, Vittorio Murino
Organizations: AI for Good (AIGO), Istituto Italiano di Tecnologia, Genoa, Italy · DITEN, University of Genoa, Genoa, Italy · National Key Laboratory of Novel Software Technology, Nanjing University, China · School of Intelligence Science and Technology, Nanjing University, China · MaLGa, DIBRIS, University of Genoa, Genoa, Italy · Department of Computer Science, University of Verona, Verona, Italy
Diffusion-based unsupervised domain adaptation (UDA) improves cross-domain transfer by generating target-specific synthetic data for downstream adaptation. Existing methods are largely designed for single-target adaptation: when a model trained on one labeled source domain must be adapted to multiple unlabeled target domains, they typically require separate diffusion fine-tuning for each source--target pair, causing training, storage, and deployment costs to grow with the number of targets. In this paper, we study multi-target data generation for diffusion-based UDA, where a single source-guided diffusion fine-tuning process is reused to generate target-specific synthetic data for multiple target domains. We propose MUSE (Multi-target UDA-oriented Synthesis with Efficient diffusion fine-tuning), a decoupled adaptation framework that separates source-supervised semantic adaptation from target-specific style adaptation. MUSE uses a shared semantic branch updated by labeled source data and target-private style branches specialized to individual target domains, enabling target-specific generation while avoiding repeated source-guided fine-tuning for each target. Experiments on standard UDA benchmarks show that MUSE achieves a stronger accuracy--efficiency trade-off than repeated per-target diffusion adaptation, reducing diffusion fine-tuning cost while improving average target-domain accuracy. The project page is available at https://xuanqi99.github.io/MUSE/.
Figures & tables
Figure 1: Overview of the proposed MUSE pipeline. MUSE performs branch-decoupled diffusion adaptation with a source-supervised shared semantic branch and target-private style branches, then activates the corresponding target branch to generate target-specific bridge samples for downstream UDA.
Figure 2: Bridge generation on miniDomainNet. (a) DDIM inversion translates a Real-domain lion into Clipart, Painting, and Sketch targets, i.e., R→{C,P,S} . (b) Gaussian-noise class-conditional generation synthesizes broccoli samples for Painting, Real, and Sketch targets under C→{P,R,S} .
Method
Diffusion Gen. Protocol
A → W
D → W
W → D
A → D
D → A
W → A
Avg.
ERM
—
77.07
96.60
99.20
81.08
64.11
64.01
80.35
DANN [ 8 ]
—
89.85
97.95
99.90
83.26
73.28
73.75
86.33
CDAN [ 25 ]
—
92.42
98.62
100.00
91.44
74.61
72.80
88.32
AFN [ 54 ]
—
91.82
98.77
100.00
95.12
72.43
70.71
88.14
MDD [ 58 ]
—
93.55
98.66
100.00
93.92
75.29
73.95
89.23
SDAT [ 39 ]
—
91.32
98.83
100.00
95.25
76.97
73.19
89.26
Table 1: Transfer accuracy (%) on Office-31. Avg. is over six tasks. Pairwise denotes per-source–target diffusion generation; Multi-target denotes one source-guided generator reused across targets.
Method
Diffusion Gen. Protocol
Ar → Cl
Ar → Pr
Ar → Rw
Cl → Ar
Cl → Pr
Cl → Rw
Pr → Ar
Pr → Cl
Pr → Rw
Rw → Ar
Rw → Cl
Rw → Pr
Avg.
ERM
—
44.06
67.12
74.26
53.26
61.96
64.54
51.91
38.90
72.94
64.51
43.84
75.39
59.39
DANN [ 8 ]
—
52.53
62.57
73.20
56.89
67.02
68.34
58.37
54.14
78.31
70.78
60.76
80.57
65.29
CDAN [ 25 ]
—
54.21
72.18
78.29
61.97
71.43
72.39
62.96
55.68
80.68
74.71
61.22
83.68
69.12
AFN [ 54 ]
—
52.58
72.42
76.96
64.90
71.14
72.91
64.08
51.29
77.83
72.21
57.46
82.09
67.99
MDD [ 58 ]
—
56.37
75.53
79.17
62.95
73.21
73.55
62.56
54.86
79.49
73.84
61.45
84.06
69.75
SDAT [ 39 ]
—
58.20
77.46
81.35
66.06
76.45
76.41
63.70
56.69
82.49
76.02
62.09
85.24
71.85
Table 2: Transfer accuracy (%) on Office-Home. Avg. is over twelve tasks. Pairwise denotes per-source–target diffusion generation; Multi-target denotes one source-guided generator reused across targets.
Method
Diffusion Gen. Protocol
C → P
C → R
C → S
P → C
P → R
P → S
R → C
R → P
R → S
S → C
S → P
S → R
Avg.
ERM
—
39.48
53.27
42.93
49.55
68.18
41.99
49.30
55.52
37.35
54.60
45.33
53.08
49.22
DANN [ 8 ]
—
45.94
56.12
49.40
50.72
65.61
50.07
55.15
60.55
49.95
58.54
54.64
58.99
54.64
AFN [ 54 ]
—
49.23
60.11
51.11
55.60
70.59
51.78
55.84
60.41
47.46
60.69
56.36
62.28
56.79
CDAN [ 25 ]
—
47.99
58.50
51.17
56.36
68.71
53.01
61.15
62.85
53.44
60.89
55.90
60.88
57.57
MDD [ 58 ]
—
48.53
61.75
52.32
59.74
70.62
55.43
62.18
62.22
54.04
63.07
58.55
64.50
59.41
SDAT [ 39 ]
—
50.97
62.42
53.91
60.57
69.97
55.85
64.39
64.83
55.86
64.07
59.43
64.28
60.55
Table 3: Transfer accuracy (%) on miniDomainNet. Avg. is over twelve tasks. Pairwise denotes per-source–target diffusion generation; Multi-target denotes one source-guided generator reused across targets.
Dataset
# Domains
Terra Time (h) ↓
MUSE Time (h) ↓
Speedup ↑
Office-31
3
96.78
51.74
1.87 ×
Office-Home
4
196.36
80.79
2.43 ×
miniDomainNet
4
197.68
81.04
2.44 ×
Table 4: SDXL fine-tuning time on one A100 80GB GPU. The maximum number of fine-tuning steps is 10000 for both Terra and MUSE. # Domains denotes the total number of domains in the benchmark.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Method
M
Peak allocated VRAM per process (GiB)
Aggregate VRAM across M concurrent processes (GiB)
Cumulative fine-tuning time (h)
Office-31
Independent fine-tuning (Terra)
2
32.433
64.865
96.78
Office-31
MUSE
2
27.849
27.849
51.74
Office-Home
Independent fine-tuning (Terra)
3
32.433
97.298
196.36
Office-Home
MUSE
3
27.835
27.835
80.79
miniDomainNet
Independent fine-tuning (Terra)
3
32.433
97.298
197.68
miniDomainNet
MUSE
3
27.835
27.835
81.04
Appendix
Table 5: Memory and training-time comparison for multi-target diffusion adaptation. Peak allocated VRAM is measured per training process. Aggregate VRAM for independent fine-tuning corresponds to concurrently executing the M target-specific processes and is obtained by summing their unrounded per-process peaks. Cumulative fine-tuning time is measured over the complete benchmark protocol.
Jtar=λmbLtgt(mb)+λorthLorth+λcoreLcore
Appendix
Algorithm 1 MUSE Diffusion Adaptation
B(m)=Binv(m)∪Bcls(m)
Appendix
Algorithm 2 Target-Specific Bridge Generation and Downstream Adaptation
Method
Diffusion Gen. Protocol
A → W
D → W
W → D
A → D
D → A
W → A
ERM
—
0.11
0.00
0.00
1.22
0.15
0.11
DANN [ 8 ]
—
1.34
0.06
0.08
0.68
0.65
0.39
CDAN [ 25 ]
—
1.75
0.18
0.00
1.19
0.79
0.45
AFN [ 54 ]
—
0.63
0.07
0.00
0.53
0.50
0.32
MDD [ 58 ]
—
1.00
0.15
0.00
0.10
0.68
0.18
SDAT [ 39 ]
—
1.83
0.12
0.00
1.03
0.67
0.34
Appendix
Table 6: Standard deviations of transfer accuracy (%) on Office-31 over three runs. Pairwise denotes per-source–target diffusion generation; Multi-target denotes one source-guided generator reused across targets.
Method
Diffusion Gen. Protocol
Ar → Cl
Ar → Pr
Ar → Rw
Cl → Ar
Cl → Pr
Cl → Rw
Pr → Ar
Pr → Cl
Pr → Rw
Rw → Ar
Rw → Cl
Rw → Pr
ERM
—
0.25
0.26
0.42
0.17
0.20
0.15
0.07
0.17
0.05
0.34
0.33
0.01
DANN [ 8 ]
—
0.44
0.72
0.38
0.02
0.30
0.39
0.58
0.47
0.59
0.84
0.14
0.51
CDAN [ 25 ]
—
0.25
0.62
0.22
0.37
0.58
0.30
0.57
0.36
0.16
0.33
0.23
0.35
AFN [ 54 ]
—
0.16
0.30
0.06
0.23
0.31
0.14
0.32
0.15
0.02
0.19
0.18
0.22
MDD [ 58 ]
—
0.51
0.32
0.06
0.24
0.73
0.41
0.36
0.53
0.24
0.03
0.09
0.11
SDAT [ 39 ]
—
0.51
0.44
0.24
0.13
0.41
0.01
1.46
0.40
0.11
0.46
0.19
0.29
Appendix
Table 7: Standard deviations of transfer accuracy (%) on Office-Home over three runs. Pairwise denotes per-source–target diffusion generation; Multi-target denotes one source-guided generator reused across targets.
Method
Diffusion Gen. Protocol
C → P
C → R
C → S
P → C
P → R
P → S
R → C
R → P
R → S
S → C
S → P
S → R
ERM
—
0.34
0.72
0.51
0.28
0.89
0.45
0.63
0.21
0.77
0.55
0.82
0.39
DANN [ 8 ]
—
0.41
0.88
0.22
0.56
0.74
0.31
0.68
0.49
0.85
0.27
0.61
0.53
AFN [ 54 ]
—
0.73
0.29
0.84
0.46
0.62
0.38
0.71
0.25
0.59
0.87
0.42
0.65
CDAN [ 25 ]
—
0.24
0.67
0.35
0.81
0.52
0.78
0.43
0.26
0.69
0.57
0.83
0.31
MDD [ 58 ]
—
0.58
0.82
0.47
0.64
0.21
0.76
0.32
0.55
0.89
0.41
0.68
0.23
SDAT [ 39 ]
—
0.37
0.61
0.85
0.29
0.44
0.73
0.56
0.81
0.34
0.67
0.25
0.79
Appendix
Table 8: Standard deviations of transfer accuracy (%) on miniDomainNet over three runs. Pairwise denotes per-source–target diffusion generation; Multi-target denotes one source-guided generator reused across targets.
Method
Ar → Cl
Ar → Pr
Ar → Rw
Cl → Ar
Cl → Pr
Cl → Rw
Pr → Ar
Pr → Cl
Pr → Rw
Rw → Ar
Rw → Cl
Rw → Pr
Avg.
Δ Avg.
MCC+MUSE w/o Lorth
64.23
80.38
81.92
72.56
81.39
80.63
71.16
63.46
82.92
74.28
65.67
84.09
75.22
-1.31
MCC+MUSE
64.31
83.44
82.37
73.14
84.44
83.50
72.90
63.02
83.47
75.37
66.58
85.82
76.53
0.00
ELS+MUSE w/o Lorth
64.91
80.67
82.61
72.03
81.51
81.27
71.61
63.51
83.71
74.17
65.86
85.11
75.58
-1.99
ELS+MUSE
66.51
84.09
82.62
74.14
83.74
82.96
74.61
64.73
85.52
75.61
68.81
87.45
77.57
0.00
SSRT+MUSE w/o Lorth
77.73
90.27
91.23
85.54
90.83
91.65
86.16
79.93
92.33
87.35
80.14
91.32
87.04
-1.11
SSRT+MUSE
79.09
90.43
92.44
87.81
92.66
92.69
87.44
79.42
93.45
88.52
80.79
93.01
88.15
0.00
Appendix
Table 9: Ablation on the orthogonality regularizer on Office-Home. Δ Avg. is computed relative to the corresponding full MUSE model under the same downstream learner.
Method
Ar → Cl
Ar → Pr
Ar → Rw
Cl → Ar
Cl → Pr
Cl → Rw
Pr → Ar
Pr → Cl
Pr → Rw
Rw → Ar
Rw → Cl
Rw → Pr
Avg.
Δ Avg.
MCC+MUSE w/ target-updated shared branch
64.05
79.01
81.54
70.09
79.91
79.73
72.02
61.97
83.77
73.92
63.98
84.68
74.56
-1.97
MCC+MUSE
64.31
83.44
82.37
73.14
84.44
83.50
72.90
63.02
83.47
75.37
66.58
85.82
76.53
0.00
ELS+MUSE w/ target-updated shared branch
64.77
79.54
82.35
71.78
80.49
80.22
71.26
62.43
83.25
74.25
64.35
85.01
74.98
-2.59
ELS+MUSE
66.51
84.09
82.62
74.14
83.74
82.96
74.61
64.73
85.52
75.61
68.81
87.45
77.57
0.00
SSRT+MUSE w/ target-updated shared branch
77.46
90.81
91.05
87.27
90.38
91.65
86.21
78.97
92.56
87.31
80.27
91.52
87.12
-1.03
SSRT+MUSE
79.09
90.43
92.44
87.81
92.66
92.69
87.44
79.42
93.45
88.52
80.79
93.01
88.15
0.00
Appendix
Table 10: Ablation on freezing the shared branch during target-conditioned updates on Office-Home. Δ Avg. is computed relative to the corresponding full MUSE model under the same downstream learner.
Method
Ar → Cl
Ar → Pr
Ar → Rw
Cl → Ar
Cl → Pr
Cl → Rw
Pr → Ar
Pr → Cl
Pr → Rw
Rw → Ar
Rw → Cl
Rw → Pr
Avg.
Δ Avg.
MCC+MUSE w/o adaptive sampling
63.19
83.75
81.61
71.66
83.41
81.85
73.08
62.10
82.10
74.53
67.02
84.61
75.74
-0.79
MCC+MUSE
64.31
83.44
82.37
73.14
84.44
83.50
72.90
63.02
83.47
75.37
66.58
85.82
76.53
0.00
ELS+MUSE w/o adaptive sampling
65.68
82.82
81.98
74.51
82.63
82.20
73.23
64.14
84.06
74.69
69.03
86.41
76.78
-0.79
ELS+MUSE
66.51
84.09
82.62
74.14
83.74
82.96
74.61
64.73
85.52
75.61
68.81
87.45
77.57
0.00
SSRT+MUSE w/o adaptive sampling
77.75
89.51
90.68
86.70
91.18
91.85
85.51
79.73
92.20
86.85
79.81
91.47
86.94
-1.21
SSRT+MUSE
79.09
90.43
92.44
87.81
92.66
92.69
87.44
79.42
93.45
88.52
80.79
93.01
88.15
0.00
Appendix
Table 11: Ablation on adaptive target-domain sampling on Office-Home. Δ Avg. is computed relative to the corresponding full MUSE model under the same downstream learner.
Generation mode
Time per image (s) ↓
DDIM inversion source-to-target generation
4.23
Pure-noise class-conditional generation
2.08
Appendix
Table 12: Per-image generation time of the two bridge-generation modes.
Method
Ar → Cl
Ar → Pr
Ar → Rw
Cl → Ar
Cl → Pr
Cl → Rw
Pr → Ar
Pr → Cl
Pr → Rw
Rw → Ar
Rw → Cl
Rw → Pr
Avg.
Δ Avg.
MCC+MUSE, Binv
61.26
78.95
78.68
69.02
79.16
78.13
69.31
61.76
79.12
73.51
65.23
84.22
73.20
-3.33
MCC+MUSE, Bcls
60.12
77.72
80.92
71.21
81.95
80.17
70.12
59.12
81.04
72.48
62.53
80.58
73.16
-3.37
MCC+MUSE, Binv∪Bcls
64.31
83.44
82.37
73.14
84.44
83.50
72.90
63.02
83.47
75.37
66.58
85.82
76.53
0.00
ELS+MUSE, Binv
62.96
81.25
79.34
69.84
79.83
79.34
69.02
61.32
80.85
74.21
66.51
85.22
74.14
-3.43
ELS+MUSE, Bcls
61.39
80.26
81.52
71.15
81.69
80.51
70.37
59.34
82.11
73.18
63.18
82.91
73.97
-3.60
ELS+MUSE, Binv∪Bcls
66.51
84.09
82.62
74.14
83.74
82.96
74.61
64.73
85.52
75.61
68.81
87.45
77.57
0.00
Appendix
Table 13: Ablation on the two components of the generated bridge set on Office-Home. Δ Avg. is computed relative to the corresponding full bridge set under the same downstream learner.
Table 14: Additional analysis of SDXL-prior data generation on Office-Home. Prompt-only denotes SDXL generation without source–target diffusion fine-tuning; Pairwise and Multi-target follow the main experimental protocol.
Figure 3: t-SNE visualizations of feature distributions for representative categories from Office-31, Office-Home, and miniDomainNet. We compare source-domain samples, target-domain samples, DDIM-inversion adapted source samples, and generated target-domain samples.
Figure 4: Qualitative visualization of DDIM-inversion-based source-to-target generation on Office-31. Rows correspond to the 6 transfer tasks A → W, D → W, W → D, A → D, D → A, and W → A. Each row shows 5 classes. Within each source–generated pair, the left image is the source-domain input and the right image is the corresponding target-style image generated by the target-specific MUSE branch.
Figure 5: Qualitative visualization of class-conditional samples generated from Gaussian noise on Office-31. Rows correspond to the 6 transfer tasks A → W, D → W, W → D, A → D, D → A, and W → A. Each row shows generated target-style samples from five different classes.
Figure 6: Qualitative visualization of DDIM-inversion-based source-to-target generation on Office-Home. Rows correspond to the 12 transfer tasks Ar → Cl, Ar → Pr, Ar → Rw, Cl → Ar, Cl → Pr, Cl → Rw, Pr → Ar, Pr → Cl, Pr → Rw, Rw → Ar, Rw → Cl, and Rw → Pr. Each row shows 5 classes. Within each pair, the left image is the source-domain input and the right image is the generated target-style image.
Figure 7: Qualitative visualization of class-conditional samples generated from Gaussian noise on Office-Home. Rows correspond to the 12 transfer tasks Ar → Cl, Ar → Pr, Ar → Rw, Cl → Ar, Cl → Pr, Cl → Rw, Pr → Ar, Pr → Cl, Pr → Rw, Rw → Ar, Rw → Cl, and Rw → Pr. Each row shows generated target-style samples from 5 different classes.
Figure 8: Qualitative visualization of DDIM-inversion-based source-to-target generation on miniDomainNet. Rows correspond to the 12 transfer tasks C → P, C → R, C → S, P → C, P → R, P → S, R → C, R → P, R → S, S → C, S → P, and S → R. Each row shows 10 classes. Within each pair, the left image is the source-domain input and the right image is the corresponding target-style image generated by the target-specific MUSE branch.
Figure 9: Qualitative visualization of class-conditional samples generated from Gaussian noise on miniDomainNet. Rows correspond to the 12 transfer tasks C → P, C → R, C → S, P → C, P → R, P → S, R → C, R → P, R → S, S → C, S → P, and S → R. Each row shows generated target-style samples from 10 different classes.
Unsupervised domain adaptation (UDA) aims to learn a target-domain classifier from labeled source data and unlabeled target data under distribution shift. Recent diffusion-based UDA methods approach this problem by synthesizing labeled target-style images and training on the resulting synthetic data. However, their performance depends heavily on the conditioning design: class prompts provide only coarse guidance, while domain adaptation modules mainly control appearance, which may leave target-style synthesis insufficiently specified. We propose VT-DUDA, a visual-token conditioning framework for diffusion-guided UDA. Instead of relying only on text prompts, VT-DUDA uses source images to provide additional instance-level visual context for target-style synthesis. Specifically, VT-DUDA maps each source image to a compact sequence of visual tokens and forms a hybrid conditioning context by concatenating these tokens with the corresponding text embeddings along the cross-attention context dimension of a latent diffusion model. This provides instance-dependent conditioning beyond text alone, while synthesis is performed with the target-domain adapter branch. Because guidance is represented explicitly as a token sequence, the same interface also permits inference-time manipulation of the conditioning signal through token selection and token-strength adjustment. The proposed method preserves the standard diffusion objective and can be integrated into existing adapter-based diffusion frameworks without modifying the backbone. Across Office-31, Office-Home, and VisDA-2017, VT-DUDA improves average target-domain accuracy over strong discriminative and diffusion-based UDA baselines. The results suggest that, in generation-based UDA, a stronger conditioning interface can improve the downstream usefulness of synthetic target-style data.
Xuan Qi, Daniele Berardini, Dario Serez +2
AI for Good (AIGO), Istituto Italiano di Tecnologia · DITEN, University of Genoa · MaLGa, DIBRIS, University of Genoa +1
Fine-tuning large diffusion models for new domains or styles involves a trade-off: improving target-specific generation often degrades the pretrained model's broad generative capability. Existing full and parameter-efficient fine-tuning methods typically handle this trade-off only implicitly. In this work, we propose a novel source-prior-driven selective adaptation method to efficiently fine-tune diffusion models, achieving a favorable trade-off. Our method relies on two key observations: (1) the loss of general generative capability is highly inconsistent across pretrained parameters, and (2) parameters that have a relatively small impact on the model's general generative capability remain structurally inconsistent across layers and parameter types. Motivated by these observations, we first learn a static mask to explicitly identify parameters better suited for downstream adaptation, and then construct structured update strategies for the selected subset. Experiments show that our method achieves a better adaptation-retention trade-off than existing strong baselines.
Yi Xiong, Yuan-Yuan Cheng, Xiao-Ming Fu
University of Science and Technology of China, China
Semi-supervised domain adaptation (SSDA) seeks to achieve accurate predictions in a target domain with limited labeled target data by exploiting abundant source and unlabeled target data. We study this problem under structural causal models (SCMs), which provide a statistical framework to describe distribution shifts between source and target domains as interventions in the data-generating process rather than ad hoc changes in model parameters. The central phenomenon is that, under low-dimensional interventions, source and unlabeled target data can help identify the high-dimensional shared structure, leaving only a low-dimensional target-specific correction to be learned from limited labeled target data. We formalize this principle for three canonical intervention models and propose the corresponding SSDA methods FT-DIP, FT-OLS-Src and FT-CIP. Under each intervention model, we demonstrate how extending an unsupervised domain adaptation (UDA) method to SSDA can achieve minimax-optimal target performance with limited target labels, with the labeled-target sample complexity scaling with the intervention dimension rather than the ambient dimension. When the distribution shift is underspecified, we propose the Multi-Adaptive-Start Fine-Tuning (MASFT) algorithm, which fine-tunes from multiple adaptive starts and selects among them using a small target validation set, incurring only logarithmic overhead in the number of starts. We validate the effectiveness of our proposed methods through simulated and real data experiments.
Wooseok Ha, Yuansi Chen
Department of Mathematical Sciences, KAIST · Department of Mathematics, ETH Zürich