Estimating heterogeneous treatment effects from observational data is difficult because the most appropriate inductive bias varies with overlap, treatment imbalance, prognostic structure, and sample size. We introduce the Geometry-Diverse Anchor-Correction Expert Ensemble (GeoACE), a five-expert framework that combines a common anchor-correction estimator with complementary overlap-aware and outcome-guided geometries. Its task-level ensemble weights are learned only from internal validation predictions, frozen before test evaluation, and then applied to experts refitted on the complete development sample. The fifth expert, O-Phi-ACE, constructs an outcome-free, overlap-aware statistical projection from covariates and treatment assignment and replaces the anchor input with this lower-dimensional geometry. We evaluate GeoACE against 11 comparators on eight benchmark protocols. Adding O-Phi-ACE reduced mean sqrt(PEHE) relative to the four-expert ensemble on all seven benchmarks with individual-effect truth, winning 998 of 1,225 paired tasks; the change on JOBS policy risk was negligible. The five-expert ensemble ranked first on IHDP100, IHDPA, and IHDPB and second on NEWS, differing from the NEWS leader by 0.13%. Across the seven sqrt(PEHE) benchmarks it obtained the lowest observed average rank (3.714), although the omnibus Friedman and Iman-Davenport tests were not significant (p=0.328 and p=0.330). Using the same five frozen experts, inverse-DR weighting was consistently better than winner-take-all selection, convex DR fitting, R-stacking, and causal Q-aggregation in benchmark-balanced analyses, but was statistically indistinguishable from equal weighting and DR ridge shrinkage. The evidence therefore supports geometry-diverse expert libraries and leakage-free aggregation as a robustness strategy, not universal superiority of either GeoACE or one weighting rule.
Figures & tables
Expert
Anchor input
Representation network input
Factual-loss weighting
Distinctive mechanism
ACE
Xstd
Xstd
Arm-frequency
Unmodified structured anchor used as the reference expert
OW-ACE
Xstd
Xstd
Overlap
Emphasizes observations for which both treatments are empirically plausible
G Φ -ACE
ΦG(Xstd)
Xstd
Arm-frequency
Pooled outcome utility, covariance-scaled mean-separation penalty, and ZCA whitening
A Φ -ACE
ΦA(Xstd)
Xstd
Arm-frequency
Arm-specific outcome utility with mean, covariance, and propensity-direction penalties
O Φ -ACE
ΦO(Xstd)
Xstd
Arm-frequency
Outcome-free overlap utility with mean, covariance, and propensity-direction penalties; projected coordinates replace the anchor input
Table 1: Compact comparison of the five causal experts. All experts use the same anchor–correction architecture, with separately fitted parameters. Each correction head also receives its expert-specific anchor prior r(j)(x) .
Figure 1: Prediction pathways of the five GeoACE experts and their validation-weighted aggregation. The structured-anchor quantities and expert-specific differences are defined in the accompanying text.
Input:
Development data Dfit∪Dval , untouched test covariates, fixed experts E ( K=5 ), and prespecified seeds b .
On Dfit fit preprocessing and, for each (j,b) , the expert’s anchor, geometry when applicable, and correction network; select its training duration mjb using Dval .
Average each expert’s validation potential-outcome predictions over seeds; construct M0,M1 and H=M1−M0 as in Equation 24 .
Form the validation DR pseudo-outcome from the fixed ACE nuisance provider ( Equation 25 ); compute each expert’s validation DR risk and the simplex weights wDR ( Equations 26 and 27 ).
Freeze {mjb} and wDR ; refit each expert, including its preprocessing, anchor and geometry, on the complete development data for exactly mjb steps.
Average each refitted expert’s test predictions over seeds; apply the frozen weights to both potential outcomes and take their difference ( Equations 11 and 12 ).
Output:
Locked test potential-outcome and CATE predictions; evaluate once after prediction.
Algorithm 1: GeoACE for one benchmark task (primary inverse-DR selector).
Benchmark
n
p
Tasks
Fit/val/test
Data source
Evaluation information
IHDPA
747
25
100
449/224/74
TEDVAE archive [ 45 ]
Continuous; semi-synthetic ITE truth
IHDPB
747
25
100
449/224/74
TEDVAE archive [ 45 ]
Continuous; semi-synthetic ITE truth
IHDP100
747
25
100
470/202/75
Johansson archive [ 25 ]
Continuous; semi-synthetic ITE truth
NEWS
5,000
3,477
50
3,000/1,500/500
Johansson archive [ 25 ]
Continuous; semi-synthetic ITE truth
TWINS
11,984
50
10
7,190/3,595/1,199
Twin-birth benchmark [ 31 ]
Binary; paired potential outcomes
JOBS
3,212
17
10
1,799/771/642
Johansson archive [ 25 ]
Binary; ATT and policy evaluation
Table 2: Summary of the eight main benchmark protocols. Split sizes are reported as fitting/validation/test counts. Sources identify the exact data archives used in this study; further construction details appear in the supplementary material.
Benchmark
ACE
OW-ACE
G Φ -ACE
A Φ -ACE
O Φ -ACE
GeoACE
ACIC2016
2.0232
1.9706
1.9904
1.9742
1.9203
1.7518
ACIC2017
0.8951
0.8901
0.8721
0.8692
0.9527
0.7990
IHDP100
0.7332
0.7042
0.7106
0.7104
0.8139
0.6742
IHDPA
0.5569
0.5469
0.5558
0.5671
0.5858
0.5195
IHDPB
2.1948
2.2271
2.2215
2.2209
2.1968
2.0963
NEWS
1.7690
1.8055
1.7614
1.7428
1.6766
1.6707
Table 3: Test performance of the five individual experts and GeoACE. Lower is better. The metric is PEHE except for JOBS, where policy risk is reported. Bold denotes the best value in each row.
Benchmark
GeoACE
Best comparator
Method
Rank
ACIC2016
1.7518
1.1418
BART
8
ACIC2017
0.7990
0.4265
CF-DML
6
IHDP100
0.6742
0.9883
CFRNet-MMD
1
IHDPA
0.5195
0.6502
TEDVAE
1
IHDPB
2.0963
2.2298
TEDVAE
1
NEWS
1.6707
1.6685
PairNet
2
Table 4: GeoACE relative to the strongest comparator on each PEHE benchmark. Lower values and ranks are better.
Benchmark
Four experts
Five experts
Δ
Relative change
Five-expert wins
ACIC2016
1.8188
1.7518
-0.0670
-3.69%
326/385
ACIC2017
0.8196
0.7990
-0.0206
-2.51%
380/480
IHDP100
0.6854
0.6742
-0.0112
-1.63%
83/100
IHDPA
0.5351
0.5195
-0.0156
-2.92%
71/100
IHDPB
2.1371
2.0963
-0.0409
-1.91%
81/100
NEWS
1.7285
1.6707
-0.0579
-3.35%
49/50
Table 5: Paired ablation of the fifth expert. Δ is five-expert minus four-expert performance, so negative values favor the five-expert system.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Benchmark-specific leave-one-expert-out sensitivity under the primary inverse-DR-risk selector. Each cell is the paired mean increase in the primary loss after removing one expert, divided by the full-ensemble mean loss. Positive values therefore favor retaining the expert. JOBS uses policy risk; the remaining benchmarks use PEHE .
Rule
Relative loss (%)
95% interval
Equal-5
-0.12
[-0.36, 0.01]
Best-DR
6.97
[ 4.08, 9.63]
Convex-DR
3.69
[ 1.34, 5.94]
R-stacking
3.30
[ 0.80, 6.16]
Causal-Q
4.90
[ 2.27, 7.32]
Ridge-DR
-0.09
[-0.31, 0.10]
Appendix
Table 6: Benchmark-balanced relative loss of each control minus GeoACE. Positive values favor GeoACE. Intervals are descriptive 95% bootstrap intervals over the eight benchmark protocols.
Method
ACIC16
ACIC17
IHDP100
IHDPA
IHDPB
NEWS
TWINS
BART
1.1418
0.5174
2.2985
0.8929
2.7530
2.2304
0.3117
CFRNet-MMD
1.5669
0.9673
0.9883
0.9183
2.7432
1.7400
0.3184
CFRNet-WASS
3.9804
2.1141
1.7064
0.8430
2.7050
3.0557
0.3246
CF-DML
1.6658
0.4265
3.8323
1.3728
3.4368
2.4439
0.3113
PairNet
3.1202
1.5311
2.0492
0.8136
2.4197
1.6685
0.3292
S-learner
1.5341
0.6062
3.3104
1.1044
3.2180
1.9041
0.3124
Appendix
Table 7: Complete out-of-sample PEHE comparison. Lower is better; bold denotes the best result in each benchmark.
Seven-benchmark average rank
JOBS policy risk
Method
Rank
Method
Risk
GeoACE
3.714
SubgroupTE
0.22475
BART
5.000
TARNet
0.22910
TEDVAE
5.143
X-learner
0.22926
CFRNet-MMD
5.286
T-learner
0.23113
TARNet
6.143
S-learner
0.23738
Appendix
Table 8: Average ranks across the seven benchmarks in Table 7 and out-of-sample JOBS policy risk. Methods are ordered independently within each pair of columns; lower is better for both measures.