Randomized controlled trials (RCTs) identify treatment effects without confounding but are often small, whereas observational studies (OBS) are large but may be confounded. Many estimators combining a small RCT with a large OBS have been developed for the average treatment effect (ATE) and the conditional ATE (CATE). However, existing ATE estimators either make assumptions on the OBS or do not borrow enough power from them. The CATE has been studied less than the ATE. Existing CATE methods either assume the OBS are unconfounded, rely on a model of the confounding function, or accept bias in exchange for lower variance. We therefore propose a framework that, without special assumptions on the OBS, fuses the OBS and the RCT by preserving the unbiasedness of RCT-based estimation while borrowing power from the large OBS to boost precision. Applying this principle, we build an ATE estimator, AIPW-Fusion, with closed-form weights and confidence intervals, and two CATE learners, DR-Fusion and R-Fusion. Experiments corroborate our findings.
Figures & tables
Method
Borrows OBS outcomes
Unbiased under any OBS confounding
OBS regressions in the trial’s AIPW score
OBS covariates used to reduce the trial variance
[YD20] Yang and Ding (2020)
✓
×
×
×
[L23] Lee et al. (2023)
×
–
×
×
[R23] Rosenman et al. (2023)
✓
×
×
×
[G25] Gao et al. (2025)
✓
×
×
×
[GB23] Gagnon-Bartsch et al. (2023)
✓
✓
✓
×
[DB25] De Bartolomeis et al. (2025)
✓
✓
✓
×
Table 1: ATE fusion methods. A dash means that the property does not apply. The second column is evaluated for a trial and an OBS with different assignment mechanisms. [DB25] borrows predictions from foundation models trained on external data rather than from an OBS. [D24] averages its predictions over a target covariate sample and, in its appendix, also uses the augmented regression in a doubly robust score.
Method
Borrows when the OBS is confounded
No model of the OBS bias
Flexible CATE class
Orthogonal loss or score
Keeps the trial’s target under any OBS confounding
[K18] Kallus et al. (2018)
✓
×
✓
×
×
[CC21] Cheng and Cai (2021)
×
✓
✓
✓
×
[WY22] Wu and Yang (2022)
✓
×
✓
✓
×
[L22] Li et al. (2022)
×
✓
✓
✓
×
[Y23] Yang et al. (2023)
×
✓
×
✓
×
[Y25] Yang et al. (2025)
✓
×
×
✓
×
Table 2: CATE fusion methods. In the last column, × means that the trial’s target is kept only if the OBS are unconfounded [L22], only if the OBS bias follows the method’s model [K18, WY22, Y25, H22], or only up to a bias traded for variance under weak confounding [CC21, Y23, G23].
Figure 1: Causal diagrams; U is unmeasured.
τ0(x)≜ER[Y(1)−Y(0)∣X=x]
CATE (target)
θ0≜ER[τ0(X)]
ATE (target)
e(x)≜PR(A=1∣X=x)
RCT propensity
r0≜dPRX/dPOX ; r
density ratio; estimate
DSnuis , DStune , DSeval
samples by role
n , N
RCT, OBS eval. sizes
PR,n , PO,N
eval.-sample means
Table 3: Notation for the ATE (Secs. 2 and 3 ).
Figure 2: ATE fusion with AIPWF (Def. 1 ). (a) Where λ and ω enter θ . (b) Contours of Eq. ( 5 ). (c) Ratio calibration (Def. 2 ): bars are the imbalances POnuis[rfk]−PRnuis[fk] .
Algorithm 1 AIPWF
R0(t)≜ER[{t(X)−τ0(X)}2]
CATE risk (Sec. 4.1 )
mdl∈{DRF,RF}
learner
mλ≜eμλ(⋅,1)+(1−e)μλ(⋅,0)
marginal regression
κmdl
unit weight (Def. 3 )
Zmdlλ
pseudo-outcome (Def. 3 )
PStune
average over DStune
Rmdl(t;λ,ω,r)
proxy risk (Eq. ( 17 ))
Table 4: Notation for the CATE (Sec. 4 ); the notation of Table 3 carries over.
Figure 3: CATE fusion with DRF and RF. (a) Where λ and ω enter the loss (Def. 3 ). (b) Algorithm 2 . (c) Thm. 4 : shading is the CATE risk; the selected candidate is within 2εeval(α) of the best (dashed) when r=r0 .
Figure 4: ATE MSE and its ratio to trial-only AIPW (top); CATE risk ratio to trial-only DR-learner (bottom); 95% bootstrap intervals. Columns 1–4: versus n ; synthetic with r0=1 , then synthetic, IHDP, and ACIC 2016 under a covariate shift. Column 5: shifted synthetic at n=3,000 , versus confounding bias /sdRτ0(X) .
Synthetic
IHDP
ACIC 2016
n
Method
r0=1
weak
strong
r0=1
weak
strong
r0=1
weak
strong
300
PPI ++
0.95
1.01
1.01
1.00
1.04
1.05
1.00
1.02
1.02
Fusion
0.55
0.60
0.60
0.46
0.47
0.48
0.45
0.46
0.47
Shrinkage
1.17
1.18
1.18
1.21
1.21
1.22
1.25
1.26
1.26
Pretest
1.17
1.16
1.20
1.71
1.70
1.76
2.07
2.09
2.16
Naive pool
17.41
16.92
16.24
11.05
10.92
10.69
9.16
8.93
8.56
Table 5: ATE mean squared error relative to trial-only AIPW, by trial size n and covariate law of the OBS: common ( r0=1 ), weak shift (effective sample fraction 0.70), and strong shift (0.40, as in Fig. 4 ).
Synthetic
IHDP
ACIC 2016
n
Method
r0=1
weak
strong
r0=1
weak
strong
r0=1
weak
strong
300
DRF
0.58
0.59
0.59
0.56
0.57
0.57
0.54
0.54
0.56
RF
0.58
0.58
0.59
0.56
0.56
0.57
0.53
0.53
0.56
PPI ++ DRF
0.81
0.80
0.81
0.79
0.79
0.79
0.82
0.80
0.83
PPI ++ RF
0.74
0.74
0.74
0.70
0.70
0.70
0.78
0.76
0.79
2-step DR
0.96
0.95
0.94
1.06
1.07
1.08
1.06
1.10
1.14
Table 6: CATE risk relative to the trial-only DR-learner; otherwise as Table 5 . RCT (R) is the trial-only R-learner, and IR the integrative R -learner.
Figure 5: Bias sweep on the synthetic design under the strong shift at n=300 , 900 , and 3,000 , against the confounding bias in units of sdRτ0(X) . Top: ATE MSE relative to trial-only AIPW; bottom: CATE risk relative to the trial-only DR-learner. 95% bootstrap intervals.
Figure 6: Effect of the observational sample size N (split 60/20/20) under the covariate shift, for the synthetic design and IHDP at n=300 and 900 . Top: ATE MSE relative to RCT only; bottom: CATE risk relative to the trial-only DR-learner. 95% bootstrap intervals.
Synthetic
IHDP
ACIC 2016
n
Method
r0=1
weak
strong
r0=1
weak
strong
r0=1
weak
strong
300
λ only
0.60
0.60
0.60
0.47
0.46
0.46
0.46
0.46
0.46
PPI ++
0.95
1.01
1.01
1.00
1.04
1.05
1.00
1.02
1.02
Fusion
0.55
0.60
0.60
0.46
0.47
0.48
0.45
0.46
0.47
Fusion, ω=D/B
0.55
0.62
0.62
0.46
0.49
0.50
0.45
0.47
0.48
900
λ only
0.78
0.78
0.78
0.83
0.84
0.83
0.42
0.42
0.43
Table 7: The two channels of fusion. ATE MSE relative to trial-only AIPW of λ only ( ω=0 ), PPI ++ ( λ=0 , ω=Proj[0,1](D/B) ), fusion with ω of Eq. ( 11 ) or with ω=Proj[0,1](D/B) .
Synthetic
IHDP
ACIC 2016
n
Method
r0=1
weak
strong
r0=1
weak
strong
r0=1
weak
strong
300
Fusion / trial-only (2/3)
0.70
0.76
0.76
0.63
0.64
0.65
0.40
0.41
0.41
Trial-only (1/3) / (2/3)
1.27
1.27
1.27
1.37
1.37
1.37
0.89
0.89
0.89
DRF / DR (2/3)
0.67
0.67
0.68
0.67
0.68
0.69
0.56
0.57
0.59
RF / DR (2/3)
0.66
0.67
0.67
0.66
0.67
0.68
0.56
0.56
0.58
RF / R (2/3)
0.69
0.70
0.70
0.70
0.70
0.71
0.58
0.58
0.61
Table 8: Ratios to the practitioner’s trial-only estimators (3-fold cross-fitting, nuisances fitted on two thirds of the trial, no tuning sample), marked (2/3). The rotated trial-only estimators of Tables 5 and 6 , whose nuisances use one third, are marked (1/3).
Synthetic
IHDP
ACIC 2016
n
Method
r0=1
weak
strong
r0=1
weak
strong
r0=1
weak
strong
300
Trial loss, sieve with g^ (DR)
0.84
0.83
0.83
0.84
0.84
0.85
0.90
0.92
0.93
Trial loss, sieve with g^ (R)
0.79
0.78
0.78
0.75
0.76
0.76
0.88
0.90
0.91
DRF / same sieve (DR)
0.69
0.70
0.72
0.67
0.68
0.68
0.60
0.59
0.60
RF / same sieve (R)
0.74
0.74
0.76
0.74
0.74
0.74
0.61
0.59
0.61
900
Trial loss, sieve with g^ (DR)
0.68
0.67
0.67
0.63
0.64
0.65
0.76
0.78
0.79
Table 9: CATE with the trial loss alone on the sieve with g : its risk relative to the trial-only DR-learner, and the risk of DRF and RF relative to it.
Design
n
Method
Bias
SD
RMSE
Cov.
MSE ratio
Synthetic
300
RCT only
0.007
0.302
0.302
0.95
1.00
Fusion
0.002
0.232
0.232
0.94
0.59
Fusion, uncalibrated r^CLS
0.123
0.233
0.264
0.90
0.76
900
RCT only
-0.007
0.139
0.139
0.95
1.00
Fusion
-0.005
0.123
0.123
0.96
0.78
Fusion, uncalibrated r^CLS
0.105
0.124
0.163
0.87
1.36
Table 10: Misspecified shift (tilt exp{a(h2−1)/2} , effective sample fraction 0.40): fusion with the calibrated ratio rBAL and with the uncalibrated logistic ratio rCLS and ω=Proj[0,1](D/B) . MSE ratio: relative to RCT only.
Law
n
Method
Bias
SD
RMSE
Cov.
MSE ratio
r0=1
300
RCT only
0.008
0.307
0.307
0.96
1.00
PPI ++
0.007
0.281
0.281
0.96
0.84
λ only
0.013
0.237
0.238
0.95
0.60
Fusion
0.012
0.205
0.206
0.95
0.45
Fusion, ω=D/B
0.012
0.205
0.206
0.95
0.45
900
RCT only
0.000
0.149
0.149
0.95
1.00
Table 11: Synthetic design with large effect heterogeneity (outcome noise 0.2, VarRτ0(X)=5.88 , confounding bias 0.75sdRτ0(X) ), under r0=1 and the strong shift. MSE ratio: relative to RCT only.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
Where
Sec. 3 : ATE
A,B,C,D
variance coefficients
Thm. 1
λ⋆,ω⋆
oracle coefficients
Thm. 1
V(λ,ω)
scaled variance nVar{θr0(λ,ω)}
Cor. 1.1
ϵA,ϵB
coefficient thresholds
Eq. ( 11 )
A,B,C,D
tuning-sample estimates of A,…,D
Sec. 3.3
Appendix
Table 12: Section-specific notation. Notation used throughout the paper is in Tables 3 and 4 .
Design
n
Method
Bias
SD
RMSE
SE
Cov.
Width
Synthetic, r0=1
300
RCT only
0.007
0.302
0.302
0.298
0.95
1.169
PPI++
0.007
0.294
0.294
0.294
0.95
1.152
Fusion
0.003
0.225
0.225
0.219
0.94
0.857
Shrinkage
0.077
0.318
0.327
–
–
–
Pretest
0.014
0.326
0.326
–
–
–
Naive pool
1.255
0.117
1.260
–
–
–
Appendix
Table 13: ATE over 1,000 replications on the synthetic design: bias, standard deviation, and root mean squared error, mean estimated standard error (SE), coverage, and mean width of the 95% Wald interval. Shrinkage, Pretest, and the naive pool have no interval.
Design
n
Method
Bias
SD
RMSE
SE
Cov.
Width
IHDP
300
RCT only
-0.007
0.255
0.255
0.250
0.96
0.981
PPI++
-0.008
0.261
0.261
0.253
0.96
0.991
Fusion
-0.002
0.177
0.177
0.182
0.95
0.713
Shrinkage
0.066
0.274
0.281
–
–
–
Pretest
0.034
0.337
0.339
–
–
–
Naive pool
0.831
0.073
0.835
–
–
–
Appendix
Table 14: As Table 13 , for IHDP and ACIC 2016.
Figure 7: Absolute errors for the designs of Fig. 4 , columns 1–4. Top: ATE MSE; bottom: CATE risk. 95% bootstrap intervals.