Estimating heterogeneous treatment effects from observational data is difficult because the most appropriate inductive bias varies with overlap, treatment imbalance, prognostic structure, and sample size. We introduce the Geometry-Diverse Anchor-Correction Expert Ensemble (GeoACE), a five-expert framework that combines a common anchor-correction estimator with complementary overlap-aware and outcome-guided geometries. Its task-level ensemble weights are learned only from internal validation predictions, frozen before test evaluation, and then applied to experts refitted on the complete development sample. The fifth expert, O-Phi-ACE, constructs an outcome-free, overlap-aware statistical projection from covariates and treatment assignment and replaces the anchor input with this lower-dimensional geometry. We evaluate GeoACE against 11 comparators on eight benchmark protocols. Adding O-Phi-ACE reduced mean sqrt(PEHE) relative to the four-expert ensemble on all seven benchmarks with individual-effect truth, winning 998 of 1,225 paired tasks; the change on JOBS policy risk was negligible. The five-expert ensemble ranked first on IHDP100, IHDPA, and IHDPB and second on NEWS, differing from the NEWS leader by 0.13%. Across the seven sqrt(PEHE) benchmarks it obtained the lowest observed average rank (3.714), although the omnibus Friedman and Iman-Davenport tests were not significant (p=0.328 and p=0.330). Using the same five frozen experts, inverse-DR weighting was consistently better than winner-take-all selection, convex DR fitting, R-stacking, and causal Q-aggregation in benchmark-balanced analyses, but was statistically indistinguishable from equal weighting and DR ridge shrinkage. The evidence therefore supports geometry-diverse expert libraries and leakage-free aggregation as a robustness strategy, not universal superiority of either GeoACE or one weighting rule.
Figures & tables
Expert
Anchor input
Representation network input
Factual-loss weighting
Distinctive mechanism
ACE
Xstd
Xstd
Arm-frequency
Unmodified structured anchor used as the reference expert
OW-ACE
Xstd
Xstd
Overlap
Emphasizes observations for which both treatments are empirically plausible
G Φ -ACE
ΦG(Xstd)
Xstd
Arm-frequency
Pooled outcome utility, covariance-scaled mean-separation penalty, and ZCA whitening
A Φ -ACE
ΦA(Xstd)
Xstd
Arm-frequency
Arm-specific outcome utility with mean, covariance, and propensity-direction penalties
O Φ -ACE
ΦO(Xstd)
Xstd
Arm-frequency
Outcome-free overlap utility with mean, covariance, and propensity-direction penalties; projected coordinates replace the anchor input
Table 1: Compact comparison of the five causal experts. All experts use the same anchor–correction architecture, with separately fitted parameters. Each correction head also receives its expert-specific anchor prior r(j)(x) .
Figure 1: Prediction pathways of the five GeoACE experts and their validation-weighted aggregation. The structured-anchor quantities and expert-specific differences are defined in the accompanying text.
Input:
Development data Dfit∪Dval , untouched test covariates, fixed experts E ( K=5 ), and prespecified seeds b .
On Dfit fit preprocessing and, for each (j,b) , the expert’s anchor, geometry when applicable, and correction network; select its training duration mjb using Dval .
Average each expert’s validation potential-outcome predictions over seeds; construct M0,M1 and H=M1−M0 as in Equation 24 .
Form the validation DR pseudo-outcome from the fixed ACE nuisance provider ( Equation 25 ); compute each expert’s validation DR risk and the simplex weights wDR ( Equations 26 and 27 ).
Freeze {mjb} and wDR ; refit each expert, including its preprocessing, anchor and geometry, on the complete development data for exactly mjb steps.
Average each refitted expert’s test predictions over seeds; apply the frozen weights to both potential outcomes and take their difference ( Equations 11 and 12 ).
Output:
Locked test potential-outcome and CATE predictions; evaluate once after prediction.
Algorithm 1: GeoACE for one benchmark task (primary inverse-DR selector).
Benchmark
n
p
Tasks
Fit/val/test
Data source
Evaluation information
IHDPA
747
25
100
449/224/74
TEDVAE archive [ 45 ]
Continuous; semi-synthetic ITE truth
IHDPB
747
25
100
449/224/74
TEDVAE archive [ 45 ]
Continuous; semi-synthetic ITE truth
IHDP100
747
25
100
470/202/75
Johansson archive [ 25 ]
Continuous; semi-synthetic ITE truth
NEWS
5,000
3,477
50
3,000/1,500/500
Johansson archive [ 25 ]
Continuous; semi-synthetic ITE truth
TWINS
11,984
50
10
7,190/3,595/1,199
Twin-birth benchmark [ 31 ]
Binary; paired potential outcomes
JOBS
3,212
17
10
1,799/771/642
Johansson archive [ 25 ]
Binary; ATT and policy evaluation
Table 2: Summary of the eight main benchmark protocols. Split sizes are reported as fitting/validation/test counts. Sources identify the exact data archives used in this study; further construction details appear in the supplementary material.
Benchmark
ACE
OW-ACE
G Φ -ACE
A Φ -ACE
O Φ -ACE
GeoACE
ACIC2016
2.0232
1.9706
1.9904
1.9742
1.9203
1.7518
ACIC2017
0.8951
0.8901
0.8721
0.8692
0.9527
0.7990
IHDP100
0.7332
0.7042
0.7106
0.7104
0.8139
0.6742
IHDPA
0.5569
0.5469
0.5558
0.5671
0.5858
0.5195
IHDPB
2.1948
2.2271
2.2215
2.2209
2.1968
2.0963
NEWS
1.7690
1.8055
1.7614
1.7428
1.6766
1.6707
Table 3: Test performance of the five individual experts and GeoACE. Lower is better. The metric is PEHE except for JOBS, where policy risk is reported. Bold denotes the best value in each row.
Benchmark
GeoACE
Best comparator
Method
Rank
ACIC2016
1.7518
1.1418
BART
8
ACIC2017
0.7990
0.4265
CF-DML
6
IHDP100
0.6742
0.9883
CFRNet-MMD
1
IHDPA
0.5195
0.6502
TEDVAE
1
IHDPB
2.0963
2.2298
TEDVAE
1
NEWS
1.6707
1.6685
PairNet
2
Table 4: GeoACE relative to the strongest comparator on each PEHE benchmark. Lower values and ranks are better.
Benchmark
Four experts
Five experts
Δ
Relative change
Five-expert wins
ACIC2016
1.8188
1.7518
-0.0670
-3.69%
326/385
ACIC2017
0.8196
0.7990
-0.0206
-2.51%
380/480
IHDP100
0.6854
0.6742
-0.0112
-1.63%
83/100
IHDPA
0.5351
0.5195
-0.0156
-2.92%
71/100
IHDPB
2.1371
2.0963
-0.0409
-1.91%
81/100
NEWS
1.7285
1.6707
-0.0579
-3.35%
49/50
Table 5: Paired ablation of the fifth expert. Δ is five-expert minus four-expert performance, so negative values favor the five-expert system.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Benchmark-specific leave-one-expert-out sensitivity under the primary inverse-DR-risk selector. Each cell is the paired mean increase in the primary loss after removing one expert, divided by the full-ensemble mean loss. Positive values therefore favor retaining the expert. JOBS uses policy risk; the remaining benchmarks use PEHE .
Rule
Relative loss (%)
95% interval
Equal-5
-0.12
[-0.36, 0.01]
Best-DR
6.97
[ 4.08, 9.63]
Convex-DR
3.69
[ 1.34, 5.94]
R-stacking
3.30
[ 0.80, 6.16]
Causal-Q
4.90
[ 2.27, 7.32]
Ridge-DR
-0.09
[-0.31, 0.10]
Appendix
Table 6: Benchmark-balanced relative loss of each control minus GeoACE. Positive values favor GeoACE. Intervals are descriptive 95% bootstrap intervals over the eight benchmark protocols.
Method
ACIC16
ACIC17
IHDP100
IHDPA
IHDPB
NEWS
TWINS
BART
1.1418
0.5174
2.2985
0.8929
2.7530
2.2304
0.3117
CFRNet-MMD
1.5669
0.9673
0.9883
0.9183
2.7432
1.7400
0.3184
CFRNet-WASS
3.9804
2.1141
1.7064
0.8430
2.7050
3.0557
0.3246
CF-DML
1.6658
0.4265
3.8323
1.3728
3.4368
2.4439
0.3113
PairNet
3.1202
1.5311
2.0492
0.8136
2.4197
1.6685
0.3292
S-learner
1.5341
0.6062
3.3104
1.1044
3.2180
1.9041
0.3124
Appendix
Table 7: Complete out-of-sample PEHE comparison. Lower is better; bold denotes the best result in each benchmark.
Seven-benchmark average rank
JOBS policy risk
Method
Rank
Method
Risk
GeoACE
3.714
SubgroupTE
0.22475
BART
5.000
TARNet
0.22910
TEDVAE
5.143
X-learner
0.22926
CFRNet-MMD
5.286
T-learner
0.23113
TARNet
6.143
S-learner
0.23738
Appendix
Table 8: Average ranks across the seven benchmarks in Table 7 and out-of-sample JOBS policy risk. Methods are ordered independently within each pair of columns; lower is better for both measures.
We study the problem of selecting the best heterogeneous treatment effect (HTE) estimator from a collection of candidates in settings where the treatment effect is fundamentally unobserved. We cast estimator selection as a multiple testing problem and introduce a ground-truth-free procedure based on a cross-fitted, exponentially weighted test statistic. A key component of our method is a two-way sample splitting scheme that decouples nuisance estimation from weight learning and ensures the stability required for valid inference. Leveraging a stability-based central limit theorem, we establish asymptotic familywise error rate control under mild regularity conditions. Empirically, our procedure provides reliable error control while substantially reducing false selections compared with commonly used methods across ACIC 2016, IHDP, and Twins benchmarks, demonstrating that our method is feasible and powerful even without ground-truth treatment effects.
Jiayi Guo, Zijun Gao
Peking University · Marshall School of Business, University of Southern California
Estimating heterogeneous treatment effects (CATE) requires simultaneously detecting effect modification and quantifying estimation uncertainty. Existing tree-based methods make an uneasy trade-off: significance-based approaches (Radcliffe and Surry 2011) identify subgroup interactions directly but lack valid inference; honest causal trees (Athey and Imbens 2016) deliver nominal confidence interval coverage but use outcome-agnostic splitting criteria that sacrifice interaction sensitivity. We introduce a hybrid algorithm that fuses significance-based splitting with honest sample-splitting and cross-validation. Our splitting criterion uses the squared t-statistic for the treatment × side interaction (t2), which is shown to be directly aligned with the honest EMSEτ criterion when the interaction is strong. Post-hoc honest cross-validation selects the cost-complexity penalty, giving a single principled estimator with nominal CI coverage at the leaf level. For forests, we retain bootstrap count vectors to enable an infinitesimal jackknife (IJ) variance estimate of Monte-Carlo convergence rather than formal pointwise inference. On the three synthetic designs from (Athey and Imbens 2016) the single tree achieves approximately 90% leaf-average CI coverage at the 90% nominal level across all three designs (200 replications each); on the Criteo, Hillstrom and Starbucks uplift datasets we match Qini coefficient performance of S-, T-learner and GRF baselines. An open-source Python package with reproducible seeds, sklearn-compatible API, and full test coverage accompanies this work (https://codeberg.org/hadjipantelis/rattus).
Pantelis Z. Hadjipantelis, Weng Man Chiang, Karthik Nagesh
A central goal of modern causal inference is estimating heterogeneous treatment effects to answer questions like "how does an intervention affect each unit," rather than only on average. We study this problem with panel-data where we observe n units across m times under unknown, non-uniform treatment assignments. The data in this setting is naturally represented as a matrix of all unit--time treatment effects. Estimating heterogeneous treatment effects can then be expressed as obtaining a good estimation of each row's average in this matrix. This allows us to formulate the problem as matrix completion, which can be solved under natural low-rankness assumptions. However, existing matrix-completion guarantees are not powerful enough to get meaningful bounds for the per-row guarantee required for estimating the heterogeneous treatment effect; roughly speaking, they are only useful for estimating average treatment effect bounds, as also illustrated in a recent line of work. We give a simple, computationally efficient estimator that, without knowledge of the propensities and under standard low-rankness and regularity assumptions, achieves a row-wise ℓ2 error of O~(n1+m2n). Technically, our analysis establishes the first sharp row-wise ℓ2-perturbation bound for low-rank approximation, complementing existing spectral-, Frobenius-, and entrywise perturbation theory.
Anay Mehrotra, Phuc Tran, Van H. Vu +1
Vin University · The University of Hong Kong · Yale University