Knowing how much a causal predictor could improve need not reveal the gain of the repair actually learned. We quantify this gap in a scalar Gaussian causal experiment with known intervention geometry: auxiliary data identify effect magnitude up to bounded contamination, while diagnostics identify direction. The target is the squared-loss gain of the realized trained repair relative to a fitted reference. Jointly optimizing the learner and assessor under uniform learning MSE η avoids the trivial solution of making no repair. At the usual 1/k learning scale, every feasible learner incurs a k−2 assessment floor, even when oracle potential is estimable at a faster rate. In the magnitude-rich regime, we characterize a sharp leading-log frontier: the assessment exponent is minℓk,2kηk/U to first relative order, where ℓk=log(1/(k2Ek)) and Ek is auxiliary precision. A diagnostic-abstention rule attains this exponent with unknown nuisance parameters. We also bound the critical allowance window and transfer the frontier to adaptive sampling by exact Gaussian simulation. Finite-grid experiments distinguish sign-tail suppression from total MSE and expose conservative finite-budget behavior. The result isolates how the assessment target changes information requirements in this experiment; it is not a general causal identifiability claim.
Figures & tables
Figure 1: Leading-log frontier, normalized by ℓk . The admissible assessment exponent saturates at auxiliary precision Ek . The curve depicts Theorem 2 , not an exact finite-sample risk. Theorem 3 shifts the first-order center Uℓk/2 to zc=Uℓk/(2−ℓk/k) . That correction can exceed the stated allowance window even when ℓk=o(k) ; it is not discarded in the refined result.
d(Q)=k8V∗logϵη8Q,T(Q)=max{0,Q−d(Q)}.
Algorithm 1: The same declared rule attains the leading exponent and the critical-window upper bound, with eventual uniform learning feasibility in their respective regimes. No unknown effect, variance, or sign is an algorithm input.
Budgets (m,n,k)
Population H
Fitted oracle A
Learned gain Ga
Decision class
Target only
Target only
Fixed clipped mean
(t,t2,t2)
t−1
t−2
t−2
(t3,t3,t)
t−3
t−3
t−2
Table 1: Same data, explicitly different decision classes. Clean data ( Δ=0 ); entries are assessment-MSE orders. The first two columns optimize only assessment (and policy if adaptive); the last fixes both the direct design and the clipped-mean learner. Joint learner selection is the separate problem in Theorem 1 .
Figure 2: Controlled mechanism checks. Left: identical report Q , different targets, at m=n=k3 , h=.3 ; bars are ±2 Monte Carlo SEs. Right: integrated wrong-sign MSE at k=224 , h2=1.1η>0 , and z=kη . The fixed-slack pair isolates the diagnostic gate; the fast schedule is Algorithm 1 . Tail suppression need not materially reduce total MSE.
Suite
Comparison
Cells
P90
Max.
Target scale
Corrected / Q
16/16
0.677
0.681
Diagnostic gate
Diagnostic / magnitude †
90/120
1
1
Robustness
Projected / raw fallback
16/24
1
1
Residual strength
Joint / mean
27/27
1.14
1.9
Fitted oracle
Selected / clipped unbiased
24/24
3.67
39.4
Pooled budgets
Joint / mean
18/18
1.59
1.83
Table 2: All seven original suites: cellwise total-MSE ratios (first report divided by second). P90 and maximum summarize the stated grid, not a minimax supremum. The Cells column gives positive-denominator cells / all cells; omitted cells have both observed MSEs zero. † Changes the learner and its target; every other row compares assessors of the same target.
Learner + report
Clean, rich
Stress
Lmax
R90
Lmax
R90
Clipped mean + Stein
0.174
6.12×103
0.178
8.69
Clipped mean + Q−V/k
0.174
17.6
0.178
8.27
Magnitude gate + IQ
0.5
7.6
10.7
16.5
Diagnostic gate + IQ
0.5
7.6
3.27
14.5
Fallback + projected report
0.122
21.2
0.291
8.37
Table 3: Finite-budget pipelines at kη=4zcclean : 192 clean, magnitude-rich cells and 768 stress cells (balanced, no-cheap, contamination-bound-only, and contaminated). Lmax is the grid maximum of learning MSE divided by η ; R90 is P90 of assessment MSE divided by EN+k−2 . Each pipeline assesses its own realized repair; only the first two rows share a learner. Values are estimates, not uniform feasibility or dominance certificates.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Analytic center-shift/window ratio for ℓ=k0.8 and U=2.5 . The condition ℓ=o(k) does not make the correction smaller than the allowance window.
k
h
k3RA(Q)
k2RG(Q)
k2RG(corr.)
k2RG(Stein)
16
.3
15.4(0.0893)
14.2(0.195)
9.67(0.149)
46.8(0.755)
16
.25/k
8.08(0.081)
13.9(0.176)
8.44(0.132)
32.4(0.505)
32
.3
17.4(0.11)
13.6(0.19)
9.21(0.145)
58.1(0.866)
32
.25/k
8.08(0.0804)
13.4(0.183)
8.36(0.141)
33(0.544)
64
.3
17.4(0.11)
13.3(0.191)
8.92(0.146)
83.1(1.18)
64
.25/k
8.15(0.0802)
12.9(0.184)
8.18(0.142)
32.7(0.566)
Appendix
Table 4: Full target-scale suite. Entries are scaled MSE (one Monte Carlo SE). The variance report is Q ; the corrected report is Q−V/k .
z
Gate
−log(k2RW)/z
E(a−θ)2/η
E(g−G)2/Ek
64
Magnitude
0.1413
6.60×10−9(7.54×10−13)
998(0.114)
64
Diagnostic, fixed
0.1413
6.60×10−9(7.54×10−13)
998(0.114)
64
Diagnostic, fast
0.1413
6.60×10−9(7.54×10−13)
998(0.114)
256
Magnitude
0.2334
3.28×10−12(4.59×10−14)
7.93(0.111)
256
Diagnostic, fixed
0.4146
3.28×10−12(4.59×10−14)
7.93(0.111)
256
Diagnostic, fast
0.4373
3.28×10−12(4.59×10−14)
7.93(0.111)
Appendix
Table 5: Conditional integration at k=224 , h2=1.1η , positive h . The exponent uses the log-domain mean of the wrong-sign MSE. Remaining columns use nuisance Monte Carlo means (one SE).
k
h
Δ
ζ
Joint
Mean
Variance
16
0.0
0.0
0.0
0.02(2.44×10−4)
0.072(0.00102)
0.0201(2.55×10−4)
16
0.0
0.1
0.0
0.0212(2.21×10−4)
0.0724(0.00104)
0.0206(2.57×10−4)
16
0.0
0.1
0.1
0.046(3.63×10−4)
0.0764(0.00107)
0.0526(4.33×10−4)
16
0.3
0.0
0.0
0.0172(1.47×10−4)
0.0811(0.0011)
0.0198(1.51×10−4)
16
0.3
0.1
0.0
0.0149(1.44×10−4)
0.0812(0.0011)
0.0197(1.54×10−4)
16
0.3
0.1
0.1
0.0297(2.57×10−4)
0.0876(0.00117)
0.0354(2.82×10−4)
Appendix
Table 6: Residual-strength suite, m=512 , n=4096 , p=1.5 . MSE (one SE); all three procedures share the same target and data. The mean method clips the corrected coefficient before applying the power.
m
k
h
Δ
Selected
Unbiased, clipped
Unbiased, raw
128
32
0.0
0.0
0.0282(4.18×10−4)
0.00968(3.00×10−4)
0.012(2.97×10−4)
128
32
0.0
0.1
0.0527(6.27×10−4)
0.011(3.41×10−4)
0.0135(3.39×10−4)
128
256
0.0
0.0
6.99×10−4(1.59×10−5)
5.95×10−4(1.41×10−5)
6.36×10−4(1.41×10−5)
128
256
0.0
0.1
7.24×10−4(1.58×10−5)
6.09×10−4(1.39×10−5)
6.55×10−4(1.39×10−5)
128
4096
0.0
0.0
3.14×10−5(6.88×10−7)
3.11×10−5(6.79×10−7)
3.11×10−5(6.79×10−7)
128
4096
0.0
0.1
3.25×10−5(6.77×10−7)
3.22×10−5(6.66×10−7)
3.23×10−5(6.66×10−7)
Appendix
Table 7: Fitted oracle-potential suite, n=4096 . MSE (one SE); Δ=ζ . Count selection chooses between the variance and mean plug-ins using EN≤1/k . The competing unbiased estimator is shown both raw and clipped.
n
k
h
Δ
Pooled joint
Mean
0
64
0.0
0.0
0.03(5.55×10−4)
0.0164(2.59×10−4)
0
64
0.0
0.1
0.0291(5.04×10−4)
0.0171(2.75×10−4)
0
64
0.3
0.0
0.0401(6.02×10−4)
0.031(4.42×10−4)
0
64
0.3
0.1
0.0427(6.05×10−4)
0.0325(4.67×10−4)
0
64
0.7
0.0
0.0615(4.57×10−4)
0.0653(4.84×10−4)
0
64
0.7
0.1
0.0615(4.58×10−4)
0.0671(4.90×10−4)
Appendix
Table 8: Pooled-budget suite, m=128 , p=1.5 , Δ=ζ . MSE (one SE). The pooled joint estimator uses diagnostic residuals even at n=0 .
c
h
k
Gain plug-in
Stein
Stein bias
Prn(G<0)
0.3
0.0
16
0.193(0.00311)
0.133(0.00237)
−8.83×10−4(0.00182)
0.783(0.00206)
0.3
0.0
128
0.004(6.77×10−5)
0.003(5.39×10−5)
4.59×10−4(2.74×10−4)
0.506(0.0025)
0.3
0.0
512
4.39×10−4(7.47×10−6)
3.76×10−4(6.39×10−6)
9.63×10−5(9.70×10−5)
0.294(0.00228)
0.3
0.3
16
0.264(0.00447)
0.197(0.00357)
0.00208(0.00222)
0.43(0.00248)
0.3
0.3
128
0.00986(1.48×10−4)
0.00888(1.26×10−4)
−0.00116(4.71×10−4)
0.0883(0.00142)
0.3
0.3
512
0.00192(2.29×10−5)
0.00186(2.10×10−5)
−1.14×10−4(2.16×10−4)
0.0265(8.02×10−4)
Appendix
Table 9: Clipping-boundary suite, m=128 , n=64 . Both reports assess the same clipped-mean repair. MSE and mean bias (one SE); the last column is the observed negative-gain fraction, not a confidence guarantee.
k
h
ζ
Base learning
Fallback learning
Base gain
Fallback gain
65536
0
0.0
6.99×10−15(7.06×10−17)
6.99×10−15(7.06×10−17)
0(0)
0(0)
65536
.95t
0.0
0.0195(1.66×10−10)
0.0195(1.66×10−10)
0(0)
0(0)
65536
1.05t
0.0
6.13×10−13(6.09×10−15)
6.13×10−13(6.09×10−15)
0.0239(1.84×10−10)
0.0239(1.84×10−10)
65536
.3
0.0
1.71×10−13(1.72×10−15)
1.71×10−13(1.72×10−15)
0.09(3.56×10−10)
0.09(3.56×10−10)
65536
0
0.05
0.05(1.73×10−9)
0.05(1.73×10−9)
−0.05(1.73×10−9)
−0.05(1.73×10−9)
65536
0
0.2
0.2(1.86×10−9)
0.2(1.86×10−9)
−0.2(1.86×10−9)
−0.2(1.86×10−9)
Appendix
Table 10: Robustness suite: learning MSE and realized mean gain (one SE). Here t2=η/64 , m=n=k3 , and η=8192log(k)/k . The base learner is magnitude-only; fallback changes the learner.
k
h
ζ
Base report/ learner
Raw / fallback
Projected / fallback
Stein / mean
65536
0
0.0
0(0)
0(0)
0(0)
7.39×10−9(1.88×10−10)
65536
.95t
0.0
0(0)
0(0)
0(0)
2.43×10−6(2.51×10−8)
65536
1.05t
0.0
5.85×10−14(5.82×10−16)
5.85×10−14(5.82×10−16)
5.85×10−14(5.82×10−16)
2.89×10−6(2.90×10−8)
65536
.3
0.0
6.17×10−14(6.21×10−16)
6.17×10−14(6.21×10−16)
6.17×10−14(6.21×10−16)
1.15×10−5(1.14×10−7)
65536
0
0.05
0.01(6.87×10−10)
0.01(6.87×10−10)
0.00179(9.18×10−7)
7.68×10−9(1.99×10−10)
65536
0
0.2
0.16(2.87×10−9)
0.16(2.87×10−9)
0.00719(3.81×10−6)
9.18×10−9(2.56×10−10)
Appendix
Table 11: Robustness suite: MSE (one SE). Raw and projected reports share the same fallback-learner target. The Stein column assesses a different direct-mean learner and is a pipeline comparator, not a same-target dominance comparison.
z
ρ
P(T>0,IQ=1)
−log(k2RW)/z
R/Ek
R/Rmag
16
1
0
−0.126
6.13×104
1
16
2
0
−0.126
6.13×104
1
16
4
0
−0.126
6.13×104
1
16
16
0
−0.126
6.13×104
1
64
1
0
0.141
8.99
1
64
2
0
0.141
8.99
1
Appendix
Table 12: Variance-bound ablation at k=16384 , m=n=k3 , h2=1.1η , c=.6 , τ=2 , and fast slack. The rule uses ρV∗ with V∗=4 ; the data and allowance are unchanged within a cell. P(T>0,IQ=1) is the nuisance-draw fraction with a positive eligible threshold. The last column compares total MSE with the matched magnitude-only rule. Negative normalized exponents at small z are permitted.
Causal queries are often only partially identifiable from observational data, and experiments that could tighten the resulting bounds are typically costly. We study the problem of selecting, prior to observing experimental outcomes, a cost-constrained subset of experiments that maximally tightens bounds on a target query. We formalize this as the max-potency problem, where epistemic potency measures the worst-case reduction in bound width guaranteed by an experiment, and show that this problem is NP-hard via a reduction from 0-1 knapsack. Building on the polynomial-programming framework of Duarte et al. (2023), we give a general procedure for evaluating epistemic potency in discrete settings. To control the super-exponential search space, we introduce two graphical pruning criteria that depend only on the causal graph and the query: a novel path-interception rule that exploits district structure to certify zero potency in linear time, and an identifiability check based on the ID algorithm. On Erdos-Renyi random graphs and 11 bnlearn benchmark networks, the two criteria together prune 50-88% of candidate experiments on average without solving a single polynomial program. For the general subset search, we show that ID-pruned experiments are combinatorially inert, yielding a super-exponential reduction in the number of subsets evaluated. We close with an end-to-end demonstration on observational NHANES data, selecting optimal experiments for estimating the effect of physical activity on diabetes.
Observational causal analyses increasingly pool records across sites, vendors, and collection systems, creating vulnerability to append-only attacks in which plausible records are strategically selected to alter a reported treatment effect. We develop a data-poisoning audit for augmented inverse-probability-weighted estimation. The analyst specifies a finite catalog of feasible records, an append budget, and nested source capacities, and the adversary selects a feasible subset to maximize movement in a prespecified direction. With preprocessing and nuisance fits held fixed, we propose a greedy scan that computes the exact finite-sample worst-case movement at every append budget. To account for nuisance refitting, we go on to derive a total-influence score combining each record's direct contribution with its effect through the propensity and outcome models. We further obtain a conservative finite-budget bound for the fully refitted estimate. Extensive simulations validate the exact result and show that total influence improves local refit prediction, while multisite and public-data analyses demonstrate material sensitivity at small append budgets. By translating adversarial data-composition risk into movement curves and critical budgets, the framework supports more reliable causal reporting and the design of source-level safeguards.
Kwangho Kim
Department of Statistics, Korea University, Seoul, Republic of Korea
Machine learning models often degrade when they are deployed on a target distribution that differs from the source distributions they were trained on. Recent work in causality-based domain generalization has shown how shared causal structure between domains can induce invariant predictors, e.g., models on a subset of features which have stable risk across structured domain shifts. However, the extent to which such population-level causal invariances can lead to gains in finite-sample settings remains underexplored. In particular, in practice we often have access to a few labeled target samples, a setting called supervised domain adaptation (sDA). In this paper, we explore when (full or partial) causal knowledge can provably improve supervised domain adaptation. As a first step, we study linear regression, where full or partial causal knowledge specifies a collection of invariant or possibly invariant feature subsets, each yielding a source-trained candidate predictor. We derive matching upper and lower bounds showing that finite-sample gains are governed by the target-risk margins separating the candidates, together with the finite-source estimation error. When these margins are sufficiently large relative to nQ, an adaptive aggregation procedure can match the best candidate predictor while avoiding negative transfer relative to target-only learning. On the other hand, when the margins are too small, no algorithm can reliably exploit the candidate collection to obtain faster finite-sample rates. We further connect these margins to structural shift magnitude in linear SCMs and validate the theory on real-world causal benchmarks.
Julia Kostin, Kasra Jalaldoust, Elias Bareinboim +2
Department of Computer Science, ETH Zurich · Causal Artificial Intelligence Lab, Columbia University · Department of Statistics, Columbia University