Organizations: The Graduate University for Advanced Studies, SOKENDAI, Tokyo, Japan · The Institute of Statistical Mathematics, Tokyo, Japan · East China Normal University, Shanghai, China · Hunan Normal University, Hunan, China · Riken AIP, Tokyo, Japan · The University of Tokyo, Tokyo, Japan · Hasso Plattner Institute, Brandenburg, Germany
Conditional independence (CI) is a fundamental concept in statistics and machine learning. Recent advances in conditional generative modeling provide flexible tools for generative-model-based CI tests, which rely on an estimated conditional distribution to generate randomized samples. However, errors in estimating this distribution accumulate in existing Type I error bounds, and consistency of the generative estimator alone does not guarantee asymptotic Type I error control. To address this limitation, we formulate conditional generative modeling as a domain adaptation problem and leverage auxiliary data from multiple source domains to improve estimation in the target CI testing domain. We propose Domain-Adapted Diffusion (DA-Diff), a multi-source domain adaptation framework for conditional diffusion models based on weighted empirical risk minimization over both target and source domains. We establish the convergence rate of DA-Diff and show how transferable source data can improve target-domain estimation through an increased effective sample size while controlling transfer bias. Building on DA-Diff, we further propose Domain-Adapted Conditional Independence Testing (DA-CIT) and show that its Type I error satisfies P(p≤α)≤α+o(1). Experiments demonstrate that DA-Diff improved conditional generation quality compared with transfer-learning diffusion baselines, while DA-CIT provides strong Type I error control and competitive power.
Figures & tables
Method
W2 ( K=2 )
W2 ( K=5 )
W2 ( K=10 )
W2 ( K=20 )
DA-Diff (Ours)
0.194 (0.030)
0.190 (0.052)
0.216 (0.060)
0.196 (0.033)
Target-only
0.343(0.079)
0.343(0.079)
0.343(0.079)
0.343(0.079)
Cat-all
0.644(0.134)
0.495(0.114)
0.372(0.109)
0.330(0.079)
Finetune
0.384(0.063)
0.283(0.056)
0.193 (0.043)
0.168 (0.036)
DoG ( 2025 )
0.311 † (0.057)
0.240 (0.068)
0.218 † (0.062)
0.222 † (0.069)
TGDP ( 2024 )
0.663(0.151)
0.466(0.110)
0.372(0.089)
0.331(0.077)
Table 1: Performance of proposed DA-Diff, Target-only, Cat-all, Finetune, DoG, TGDP, UOWQ, and Robust under the data-generating mechanism in Appendix D.1 . Experiments are repeated over 30 seeds; we report the W2 distance along with one standard deviation (std). The best, second-best, and third-best results are highlighted in bold, underlined, and † , respectively.
Figure 1: Type I error and power of each CI testing under different dz .
Label
Role
Dataset
Stimulation / intervention condition
sample size
0
Target
cd3cd28
CD3 + CD28
853
1
Source
b2camp
β2 cAMP
707
2
Source
cd3cd28_aktinhib
CD3 + CD28 + Akt inhibitor
911
3
Source
cd3cd28_g0076
CD3 + CD28 + Gö6976
723
4
Source
cd3cd28_ly
CD3 + CD28 + LY294002
848
5
Source
cd3cd28_psitect
CD3 + CD28 + Psitectorigenin
810
Table 2: The 14 domains used in the Sachs experiment. The cd3cd28 condition is used as the target domain, while the remaining 13 conditions are treated as source domains. The domain labels are consistent with those used in our implementation and visualization.
Figure 2: t-SNE visualization of the 14 domains in the Sachs dataset. Each color represents a different stimulation or intervention condition. The cd3cd28 dataset is used as the target domain, and the remaining 13 datasets are used as source domains.
Method
Precision
Recall
F-score
DA-CIT (Ours)
0.738 †
0.960
0.834
CRT ∗ ( 2026c )
0.764
0.260
0.388
CDCIT ( 2025a )
0.671 / 0.692
0.980 / 0.540
0.796 † / 0.606
KCIT ( 2011 )
0.725 / 0.750
0.900 † / 0.060
0.803 / 0.111
RBPT ( 2023 )
0.620 / 0.625
0.980 / 0.600
0.759 / 0.612
ECCIT-GCM ( 2026 )
0.764 / 0.733
0.520 / 0.220
0.619 / 0.338
Table 3: Performance of different CI testing methods on the Sachs dataset. Each entry reports Target-only / Cat-all . DA-CIT and CRT ∗ report a single result because both methods directly leverage external source data, making the Target-only/Cat-all distinction unnecessary. The best, second-best, and third-best results are highlighted in bold, underlined, and † , respectively.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Left panel: The source densities are sufficiently covered by the target density (Cdrmax1≤k≤Kpk(z)≤p0(z)) , while their weighted mixture pw∖0 also covers the target density (p0(z)≤CDRpw∖0(z)) ; hence both density-ratio conditions can hold uniformly. Right panel: One source is located far from the target, while the remaining source does not sufficiently cover the target distribution. Consequently, the weighted source density fails to uniformly dominate p0 , violating the transferability condition.
Sample size
Method
W2 ( K=2 )
W2 ( K=5 )
W2 ( K=10 )
W2 ( K=20 )
1000
DA-Diff (Ours)
0.142 (0.039)
0.140 (0.037)
0.133 (0.030)
0.134 (0.027)
Target-only
0.376(0.070)
0.376(0.070)
0.376(0.070)
0.376(0.070)
Cat-all
0.440(0.112)
0.371(0.087)
0.334(0.088)
0.308(0.063)
Finetune
0.216 † (0.056)
0.203 (0.057)
0.173 (0.045)
0.164 (0.040)
DoG ( 2025 )
0.285(0.103)
0.231(0.066)
0.216(0.064)
0.181(0.046)
TGDP ( 2024 )
0.562(0.135)
0.413(0.116)
0.389(0.104)
0.328(0.074)
Appendix
Table 4: Performance of the proposed DA-Diff, Target-only, Cat-all, Finetune, DoG ( Zhong et al., 2025 ) , TGDP ( Ouyang et al., 2024 ) , and Robust ( Konstantinov and Lampert, 2019 ) under the data-generating mechanism in Section 4.1 , with sample size n0=⋯=nK=1000,1500,2000 . We report the W2 distance along with one standard deviation (std).
Figure 4: Sensitivity study of DA-Diff on the Gaussian mixture data. We vary the number of time bins BT , the weight/tuning update interval update_every , the resolution of the Gλ , and the number of warmup epochs ew , and evaluate the performance using the W2 distance.
Figure 5: Sensitivity results of DA-CIT on the post-nonlinear data with dz=20 . We vary the number of source domains, the proportion of near and far sources, the target and source sample sizes, and the alternative strength h1 .
Figure 6: Ablation results on the post-nonlinear data with dz=20 under the Cat-all setting. We vary the number of source domains, the proportion of near and far sources, the target and source sample sizes, and the alternative strength h1 .
Figure 7: ECDFs of the p -values under H0 for the CI testing methods considered in Figure 1 . The dashed diagonal line represents the CDF of U(0,1) , and pKS denotes the p -value of the KS test based on 100 null p -values.
Figure 8: ECDFs of the DA-CIT p -values under H0 for the ablation settings considered in Figures 5 and 6 , with dz=20 . The dashed diagonal line represents the CDF of U(0,1) , and pKS denotes the p -value of the KS test based on 100 null p -values.
p=1+B1+∑b=1B1{T(b)≥T}.
Appendix
Algorithm 2 Domain-adapted conditional independence test (DA-CIT)
Figure 9: Histograms of the target and source variable X in the Gaussian mixture experiment. The orange histograms represent the target-domain data, while the blue histograms represent the source-domain data. The left panel compares the target domain with the first source domain, whose shift strength is d(1)=0.01 , whereas the right panel compares the target domain with the third source domain, whose shift strength is d(3)=0.2 . As the shift strength increases to d(3)=0.2 , a clear distributional discrepancy between the target and source domains can be observed.
Figure 10: Validation loss curves of DA-Diff and the target-only diffusion model during training.
Figure 11: Evolution of the MM-QP solution π and the corresponding w1,…,wK and WN recovered by ( 34 ) during the training of DA-Diff.
Figure 12: Hyperparameters ρ , Creg , and λ selected by the gradient-alignment criterion during the training of DA-Diff.
Figure 13: Generated samples in the Gaussian mixture experiment. From left to right, the methods are Target only, Cat-all, and DA-Diff. Their corresponding W2 distances are 0.040 , 0.047 , and 0.023 , respectively.
Figure 14: Source proportions α1,…,αK , source weights w1,…,wK , and the total weighted source sample size WN solved by UOWQ ( Zhang et al., 2026a ) .
Figure 15: Wall-clock runtime of the CI testing methods considered in Figure 1 across different conditioning dimensions dz . We report the runtime under both the Target-only and Concatenate-all settings. The vertical axis is shown on a logarithmic scale.
Figure 16: Detailed runtime decomposition of DA-CIT under the ablation settings considered in Figures 5 and 6 . The total runtime is decomposed into source pretraining, MM-QP optimization, adaptation excluding MM-QP, CRT sampling, and statistic computation.
aDepartment of Applied Mathematics, The Hong Kong Polytechnic University, Hong Kong, China · bSchool of Statistics and Data Science, Nankai University, Tianjin, China · cSchool of Statistics, East China Normal University, Shanghai, China +1