Randomized controlled trials (RCTs) identify treatment effects without confounding but are often small, whereas observational studies (OBS) are large but may be confounded. Many estimators combining a small RCT with a large OBS have been developed for the average treatment effect (ATE) and the conditional ATE (CATE). However, existing ATE estimators either make assumptions on the OBS or do not borrow enough power from them. The CATE has been studied less than the ATE. Existing CATE methods either assume the OBS are unconfounded, rely on a model of the confounding function, or accept bias in exchange for lower variance. We therefore propose a framework that, without special assumptions on the OBS, fuses the OBS and the RCT by preserving the unbiasedness of RCT-based estimation while borrowing power from the large OBS to boost precision. Applying this principle, we build an ATE estimator, AIPW-Fusion, with closed-form weights and confidence intervals, and two CATE learners, DR-Fusion and R-Fusion. Experiments corroborate our findings.
Figures & tables
Method
Borrows OBS outcomes
Unbiased under any OBS confounding
OBS regressions in the trial’s AIPW score
OBS covariates used to reduce the trial variance
[YD20] Yang and Ding (2020)
✓
×
×
×
[L23] Lee et al. (2023)
×
–
×
×
[R23] Rosenman et al. (2023)
✓
×
×
×
[G25] Gao et al. (2025)
✓
×
×
×
[GB23] Gagnon-Bartsch et al. (2023)
✓
✓
✓
×
[DB25] De Bartolomeis et al. (2025)
✓
✓
✓
×
Table 1: ATE fusion methods. A dash means that the property does not apply. The second column is evaluated for a trial and an OBS with different assignment mechanisms. [DB25] borrows predictions from foundation models trained on external data rather than from an OBS. [D24] averages its predictions over a target covariate sample and, in its appendix, also uses the augmented regression in a doubly robust score.
Method
Borrows when the OBS is confounded
No model of the OBS bias
Flexible CATE class
Orthogonal loss or score
Keeps the trial’s target under any OBS confounding
[K18] Kallus et al. (2018)
✓
×
✓
×
×
[CC21] Cheng and Cai (2021)
×
✓
✓
✓
×
[WY22] Wu and Yang (2022)
✓
×
✓
✓
×
[L22] Li et al. (2022)
×
✓
✓
✓
×
[Y23] Yang et al. (2023)
×
✓
×
✓
×
[Y25] Yang et al. (2025)
✓
×
×
✓
×
Table 2: CATE fusion methods. In the last column, × means that the trial’s target is kept only if the OBS are unconfounded [L22], only if the OBS bias follows the method’s model [K18, WY22, Y25, H22], or only up to a bias traded for variance under weak confounding [CC21, Y23, G23].
Figure 1: Causal diagrams; U is unmeasured.
τ0(x)≜ER[Y(1)−Y(0)∣X=x]
CATE (target)
θ0≜ER[τ0(X)]
ATE (target)
e(x)≜PR(A=1∣X=x)
RCT propensity
r0≜dPRX/dPOX ; r
density ratio; estimate
DSnuis , DStune , DSeval
samples by role
n , N
RCT, OBS eval. sizes
PR,n , PO,N
eval.-sample means
Table 3: Notation for the ATE (Secs. 2 and 3 ).
Figure 2: ATE fusion with AIPWF (Def. 1 ). (a) Where λ and ω enter θ . (b) Contours of Eq. ( 5 ). (c) Ratio calibration (Def. 2 ): bars are the imbalances POnuis[rfk]−PRnuis[fk] .
Algorithm 1 AIPWF
R0(t)≜ER[{t(X)−τ0(X)}2]
CATE risk (Sec. 4.1 )
mdl∈{DRF,RF}
learner
mλ≜eμλ(⋅,1)+(1−e)μλ(⋅,0)
marginal regression
κmdl
unit weight (Def. 3 )
Zmdlλ
pseudo-outcome (Def. 3 )
PStune
average over DStune
Rmdl(t;λ,ω,r)
proxy risk (Eq. ( 17 ))
Table 4: Notation for the CATE (Sec. 4 ); the notation of Table 3 carries over.
Figure 3: CATE fusion with DRF and RF. (a) Where λ and ω enter the loss (Def. 3 ). (b) Algorithm 2 . (c) Thm. 4 : shading is the CATE risk; the selected candidate is within 2εeval(α) of the best (dashed) when r=r0 .
Figure 4: ATE MSE and its ratio to trial-only AIPW (top); CATE risk ratio to trial-only DR-learner (bottom); 95% bootstrap intervals. Columns 1–4: versus n ; synthetic with r0=1 , then synthetic, IHDP, and ACIC 2016 under a covariate shift. Column 5: shifted synthetic at n=3,000 , versus confounding bias /sdRτ0(X) .
Synthetic
IHDP
ACIC 2016
n
Method
r0=1
weak
strong
r0=1
weak
strong
r0=1
weak
strong
300
PPI ++
0.95
1.01
1.01
1.00
1.04
1.05
1.00
1.02
1.02
Fusion
0.55
0.60
0.60
0.46
0.47
0.48
0.45
0.46
0.47
Shrinkage
1.17
1.18
1.18
1.21
1.21
1.22
1.25
1.26
1.26
Pretest
1.17
1.16
1.20
1.71
1.70
1.76
2.07
2.09
2.16
Naive pool
17.41
16.92
16.24
11.05
10.92
10.69
9.16
8.93
8.56
Table 5: ATE mean squared error relative to trial-only AIPW, by trial size n and covariate law of the OBS: common ( r0=1 ), weak shift (effective sample fraction 0.70), and strong shift (0.40, as in Fig. 4 ).
Synthetic
IHDP
ACIC 2016
n
Method
r0=1
weak
strong
r0=1
weak
strong
r0=1
weak
strong
300
DRF
0.58
0.59
0.59
0.56
0.57
0.57
0.54
0.54
0.56
RF
0.58
0.58
0.59
0.56
0.56
0.57
0.53
0.53
0.56
PPI ++ DRF
0.81
0.80
0.81
0.79
0.79
0.79
0.82
0.80
0.83
PPI ++ RF
0.74
0.74
0.74
0.70
0.70
0.70
0.78
0.76
0.79
2-step DR
0.96
0.95
0.94
1.06
1.07
1.08
1.06
1.10
1.14
Table 6: CATE risk relative to the trial-only DR-learner; otherwise as Table 5 . RCT (R) is the trial-only R-learner, and IR the integrative R -learner.
Figure 5: Bias sweep on the synthetic design under the strong shift at n=300 , 900 , and 3,000 , against the confounding bias in units of sdRτ0(X) . Top: ATE MSE relative to trial-only AIPW; bottom: CATE risk relative to the trial-only DR-learner. 95% bootstrap intervals.
Figure 6: Effect of the observational sample size N (split 60/20/20) under the covariate shift, for the synthetic design and IHDP at n=300 and 900 . Top: ATE MSE relative to RCT only; bottom: CATE risk relative to the trial-only DR-learner. 95% bootstrap intervals.
Synthetic
IHDP
ACIC 2016
n
Method
r0=1
weak
strong
r0=1
weak
strong
r0=1
weak
strong
300
λ only
0.60
0.60
0.60
0.47
0.46
0.46
0.46
0.46
0.46
PPI ++
0.95
1.01
1.01
1.00
1.04
1.05
1.00
1.02
1.02
Fusion
0.55
0.60
0.60
0.46
0.47
0.48
0.45
0.46
0.47
Fusion, ω=D/B
0.55
0.62
0.62
0.46
0.49
0.50
0.45
0.47
0.48
900
λ only
0.78
0.78
0.78
0.83
0.84
0.83
0.42
0.42
0.43
Table 7: The two channels of fusion. ATE MSE relative to trial-only AIPW of λ only ( ω=0 ), PPI ++ ( λ=0 , ω=Proj[0,1](D/B) ), fusion with ω of Eq. ( 11 ) or with ω=Proj[0,1](D/B) .
Synthetic
IHDP
ACIC 2016
n
Method
r0=1
weak
strong
r0=1
weak
strong
r0=1
weak
strong
300
Fusion / trial-only (2/3)
0.70
0.76
0.76
0.63
0.64
0.65
0.40
0.41
0.41
Trial-only (1/3) / (2/3)
1.27
1.27
1.27
1.37
1.37
1.37
0.89
0.89
0.89
DRF / DR (2/3)
0.67
0.67
0.68
0.67
0.68
0.69
0.56
0.57
0.59
RF / DR (2/3)
0.66
0.67
0.67
0.66
0.67
0.68
0.56
0.56
0.58
RF / R (2/3)
0.69
0.70
0.70
0.70
0.70
0.71
0.58
0.58
0.61
Table 8: Ratios to the practitioner’s trial-only estimators (3-fold cross-fitting, nuisances fitted on two thirds of the trial, no tuning sample), marked (2/3). The rotated trial-only estimators of Tables 5 and 6 , whose nuisances use one third, are marked (1/3).
Synthetic
IHDP
ACIC 2016
n
Method
r0=1
weak
strong
r0=1
weak
strong
r0=1
weak
strong
300
Trial loss, sieve with g^ (DR)
0.84
0.83
0.83
0.84
0.84
0.85
0.90
0.92
0.93
Trial loss, sieve with g^ (R)
0.79
0.78
0.78
0.75
0.76
0.76
0.88
0.90
0.91
DRF / same sieve (DR)
0.69
0.70
0.72
0.67
0.68
0.68
0.60
0.59
0.60
RF / same sieve (R)
0.74
0.74
0.76
0.74
0.74
0.74
0.61
0.59
0.61
900
Trial loss, sieve with g^ (DR)
0.68
0.67
0.67
0.63
0.64
0.65
0.76
0.78
0.79
Table 9: CATE with the trial loss alone on the sieve with g : its risk relative to the trial-only DR-learner, and the risk of DRF and RF relative to it.
Design
n
Method
Bias
SD
RMSE
Cov.
MSE ratio
Synthetic
300
RCT only
0.007
0.302
0.302
0.95
1.00
Fusion
0.002
0.232
0.232
0.94
0.59
Fusion, uncalibrated r^CLS
0.123
0.233
0.264
0.90
0.76
900
RCT only
-0.007
0.139
0.139
0.95
1.00
Fusion
-0.005
0.123
0.123
0.96
0.78
Fusion, uncalibrated r^CLS
0.105
0.124
0.163
0.87
1.36
Table 10: Misspecified shift (tilt exp{a(h2−1)/2} , effective sample fraction 0.40): fusion with the calibrated ratio rBAL and with the uncalibrated logistic ratio rCLS and ω=Proj[0,1](D/B) . MSE ratio: relative to RCT only.
Law
n
Method
Bias
SD
RMSE
Cov.
MSE ratio
r0=1
300
RCT only
0.008
0.307
0.307
0.96
1.00
PPI ++
0.007
0.281
0.281
0.96
0.84
λ only
0.013
0.237
0.238
0.95
0.60
Fusion
0.012
0.205
0.206
0.95
0.45
Fusion, ω=D/B
0.012
0.205
0.206
0.95
0.45
900
RCT only
0.000
0.149
0.149
0.95
1.00
Table 11: Synthetic design with large effect heterogeneity (outcome noise 0.2, VarRτ0(X)=5.88 , confounding bias 0.75sdRτ0(X) ), under r0=1 and the strong shift. MSE ratio: relative to RCT only.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
Where
Sec. 3 : ATE
A,B,C,D
variance coefficients
Thm. 1
λ⋆,ω⋆
oracle coefficients
Thm. 1
V(λ,ω)
scaled variance nVar{θr0(λ,ω)}
Cor. 1.1
ϵA,ϵB
coefficient thresholds
Eq. ( 11 )
A,B,C,D
tuning-sample estimates of A,…,D
Sec. 3.3
Appendix
Table 12: Section-specific notation. Notation used throughout the paper is in Tables 3 and 4 .
Design
n
Method
Bias
SD
RMSE
SE
Cov.
Width
Synthetic, r0=1
300
RCT only
0.007
0.302
0.302
0.298
0.95
1.169
PPI++
0.007
0.294
0.294
0.294
0.95
1.152
Fusion
0.003
0.225
0.225
0.219
0.94
0.857
Shrinkage
0.077
0.318
0.327
–
–
–
Pretest
0.014
0.326
0.326
–
–
–
Naive pool
1.255
0.117
1.260
–
–
–
Appendix
Table 13: ATE over 1,000 replications on the synthetic design: bias, standard deviation, and root mean squared error, mean estimated standard error (SE), coverage, and mean width of the 95% Wald interval. Shrinkage, Pretest, and the naive pool have no interval.
Design
n
Method
Bias
SD
RMSE
SE
Cov.
Width
IHDP
300
RCT only
-0.007
0.255
0.255
0.250
0.96
0.981
PPI++
-0.008
0.261
0.261
0.253
0.96
0.991
Fusion
-0.002
0.177
0.177
0.182
0.95
0.713
Shrinkage
0.066
0.274
0.281
–
–
–
Pretest
0.034
0.337
0.339
–
–
–
Naive pool
0.831
0.073
0.835
–
–
–
Appendix
Table 14: As Table 13 , for IHDP and ACIC 2016.
Figure 7: Absolute errors for the designs of Fig. 4 , columns 1–4. Top: ATE MSE; bottom: CATE risk. 95% bootstrap intervals.
When treatment effects are naturally expressed as ratios -- as in medicine, pricing, and marketing -- the ratio-based CATE τ(x)=E[Y∣W=1,X=x]/E[Y∣W=0,X=x] is the appropriate estimand. Yet existing estimators either impose a log-linear parametric structure or apply generic regression without robustness guarantees for this functional. We introduce the Q-Learner, which decomposes τ(x) into a product of two odds ratios, reducing ratio-CATE estimation for binary outcomes to two propensity classification tasks. We further derive doubly robust augmentations for both S/T- and Q-style ratio learners and characterize their distinct robustness properties. In benchmarks on seven RCT datasets, the Q-Learner is the most consistently competitive method in low-conversion regimes, where its propensity-only construction sidesteps the imbalanced regression that hurts outcome-based estimators. On four observational datasets, where propensity must be estimated and confounding cannot be ruled out, the DR learners introduced here decisively come out on top, making them practitioners' natural default for confounded observational data.
Michael Fuchs, Dominik Kreiss
Actuarial Department Allianz Versicherungs-AG Munich, Germany
Randomized controlled trials (RCTs) are the gold standard for estimating treatment effects, yet they are often underpowered for detecting effect heterogeneity. Large observational studies (OS) can supplement RCTs for conditional average treatment effect (CATE) estimation, but a key barrier is covariate mismatch: the two sources measure different, only partially overlapping, covariates. We propose CALM (Calibrated ALignment under covariate Mismatch), which learns embeddings that map each source's features into a common representation space. OS outcome models are transferred to the RCT embedding space and calibrated using trial data, preserving causal identification from randomization. Finite-sample risk bounds decompose into alignment error, outcome-model complexity, and calibration complexity terms, making explicit when the learned embedding is accurate enough to reduce variance. We instantiate CALM in two forms: a closed-form linear version, CALM-Lin, and a neural representation-learning version, CALM-NN. Across 51 simulation settings, calibration-based linear methods are effectively tied in linear-CATE regimes, while CALM-NN wins all 22 nonlinear-CATE settings by wide margins. Moreover, on two real-data studies CALM-NN delivers the largest gains over the trial-only baseline.
Amir Asiaee, Samhita Pal
Department of Biostatistics, Vanderbilt University Medical Center, TN 37203, USA
Average treatment effect (ATE) estimation in observational studies is a fundamental statistical tool used frequently in social science, medicine, and other fields. These fields often work with sensitive data where privacy protections are important, so a differentially private mechanism for ATE estimation is highly desirable. Here we present two propensity score-based algorithms for ATE estimation on observational data, one improving the inverse probability weighting (IPW) method used in prior work, and the other using blocking on the propensity score (BPS). Both show lower error and less bias than prior work, with the BPS-based algorithm frequently reducing error by 75% or more compared to prior work.
Duncan Stewardson, Grayson W. White, Adam Groce
Department of Computer Science Reed College Portland, OR, 97202 · Department of Mathematics and Statistics Reed College Portland, OR, 97202