GeoDose-CP: Graph-Local Conformal Inference for Continuous-Treatment Earth Observation
Authors: Md Khalid Hasan Sakib, Dristi Datta, Manoranjan Paul, Davina White
Organizations: Department of Computer Science and Engineering, Uttara University, Dhaka, Bangladesh · School of Computing, Mathematics and Engineering, Charles Sturt University, Bathurst, NSW 2795, Australia · Australian Integrated Carbon (AiCarbon), Level 4, 191 Pulteney Street, Adelaide, SA 5000, Australia
Reliable intervention-oriented uncertainty quantification from Earth observation (EO) remains challenging when continuous treatment shifts, spatial dependence, limited support, and satellite-outcome uncertainty must be addressed simultaneously. Existing causal, conformal, and spatial approaches address parts of this problem, but their direct combination does not generally recover the appropriate interventional reference law because candidate reassignment jointly alters treatment likelihood, standardized residuals, and graph-dependent residual likelihood. This study presents GeoDose-CP, a support-aware conformal framework for localized stochastic potential outcomes under continuous or mixed continuous-atomic treatment. Its central methodological contribution is a graph-local target-orbit law that jointly represents intervention-induced treatment shift, the inverse outcome-scale Jacobian, and spatial residual dependence. The framework further provides exact weighted candidate inversion, a scalable sparse approximation with explicit discrepancy accounting, and refusal under inadequate support. Evaluation used controlled known-truth experiments, MineDoseBench, treatment-density sensitivity analysis, external conformal comparators, and a multi-mine New South Wales (NSW) study. In MineDoseBench, GeoDose-CP achieved mean selective coverage of 0.9692 across 27 configurations and a minimum local q0.05 of 0.8951; exact-sparse auditing produced nine inclusion disagreements over 2,700 targets. In the NSW study, the absence of an auditable longitudinal rehabilitation treatment rendered treatment-dependent inference nonoperational rather than forcing inference through a proxy exposure.
Figures & tables
Fig. 1: Operational workflow of GeoDose-CP. Spatial graph and EO data, together with frozen nuisance models and the supported-intervention conditions C1–C5, define the inputs to candidate evaluation. For each candidate response, G1 constructs the local candidate orbit and G2 evaluates the joint target-orbit law through the treatment likelihood, inverse-scale Jacobian, and graph-residual likelihood. Inference then proceeds through exact weighted inference (G3) or the sparse approximation (N2), followed by weighted p -value calculation and candidate inversion. Support and numerical checks, together with the applicable N3 assessment for the sparse route, determine whether the query returns a prediction set with diagnostics or an explicit refusal.
Method
Mean Sel. Cov.
Min. Sel. Cov.
Cases ≥0.90
Mean Local q0.05
Min. Local q0.05
Return Rate
M2
0.8855
0.8022
8/27
0.8265
0.7079
0.9500
M3
0.9856
0.8360
26/27
0.9725
0.7464
1.0000
M4
0.9805
0.9422
27/27
0.9574
0.8804
0.9500
M5
0.8989
0.8360
12/27
0.8385
0.7694
0.8945
M6
0.9692
0.9140
27/27
0.9488
0.8951
0.9500
TABLE I: Overall MineDoseBench Performance Across 27 Configurations
Fig. 2: MineDoseBench selective coverage, local robustness, and support-aware refusal across 27 configurations. (a) Configuration-wise selective coverage of full GeoDose-CP (M6). (b) Fifth-percentile target-local selective coverage for M2–M6. (c) Operational refusal rates for M2–M6. Dashed horizontal lines denote nominal 0.90 coverage. Results correspond to the RF-primary evaluation on the observed benchmark-response scale.
Fig. 3: Effect of increasing spatial dependence in S4. (a) RF-primary selective coverage of M2–M6 for ρ∈{0,0.2,0.4,0.6,0.8} ; error bars denote 95% replication-by-mine cluster-bootstrap confidence intervals. (b) Mean finite-domain interval/hull width on YMDB=[−1,1] for the same prespecified S4 target identities. For M3, M4, and M6, widths are obtained from full finite-domain candidate inversion; for M2 and M5, closed-form returned intervals are intersected with YMDB for this efficiency-only representation. Error bars denote 95% replication-by-mine cluster-bootstrap confidence intervals for mean width. Width is interpreted as efficiency only when the selective-coverage difference satisfies the prespecified 0.03 coverage-matching criterion; the ρ=0.8 M6–M5 comparison is therefore not interpreted as an efficiency comparison.
Fig. 4: Exact–sparse M6 discrepancy across 2,700 paired registered target identities. The empirical cumulative distribution function shows the absolute difference ∣Δp∣ between full-graph exact inference and the sparse m=64 route. Vertical reference lines mark the 0.90, 0.95, and 0.99 quantiles ( 0.01607 , 0.03467 , and 0.09739 , respectively). The median discrepancy is numerically zero and the maximum is 0.42945 . Only 9 of 2,700 targets (0.33%) changed inclusion status at α=0.10 , so close aggregate and decision-level agreement does not imply pointwise numerical equivalence.
Audit
Selective coverage
Return/support
Principal finding
Target-design shift
0.9644
90% return
Coverage remains above the nominal 0.90 level under nonidentity target-design shift.
Non-Gaussian spatial law
0.9667
90% return
Coverage remains above nominal under the non-Gaussian residual-dependence setting.
Measurement-error perturbations
0.966–0.972
100% return
Coverage remains above nominal across all five measurement-error perturbations.
Change of spatial support
90 m: 0.9640; 180 m: 0.9578
90 m: 100% return; 180 m: 90% return
Coverage remains above nominal at both independently anchored spatial supports, with lower return at 180 m.
Supported endpoint atoms
A=0 : 0.980; A=1 : 0.990
100% return at both atoms
Both genuine endpoint atoms remain supported under the mixed treatment measure.
Unsupported endpoints
—
100% refusal at A=0,1
The interior-only treatment law refuses unsupported endpoint queries rather than assigning artificial endpoint mass.
TABLE II: Stress, Support, and Applicability Audits for GeoDose-CP
Treatment law
Mean sel. cov.
Min. sel. cov.
Mean local q0.05
Min. local q0.05
Return rate
GORACLE
0.9830
0.9680
0.9661
0.9347
0.9571
GESTIMATED
0.9825
0.9644
0.9655
0.9267
0.9071
GMISSPECIFIED
0.9786
0.9537
0.9545
0.8882
0.9857
TABLE III: Treatment-Density Sensitivity for M6
Method
Mean
Min.
Cases
Mean local
Min. local
Native
Return
Infinite-set
sel. cov.
sel. cov.
≥0.90
q0.05
q0.05
applicability
∣ applicable
∣ returned
E1 continuous-treatment
0.8970
0.8120
3/6
0.8310
0.7193
0.8571
1.0000
0.3097
E2 weighted shift
0.9191
0.8580
5/7
0.8701
0.7940
1.0000
1.0000
0.2714
E3 spatial/local
0.9463
0.9180
7/7
0.9134
0.8817
1.0000
1.0000
0.0000
M6 GeoDose-CP
0.9825
0.9644
7/7
0.9655
0.9267
1.0000
0.9071
0.0000
TABLE IV: External-Comparator Performance on the Selected MineDoseBench Evaluation Subset
Fig. 5: External-comparator performance on the selected MineDoseBench evaluation subset. The three panels show configuration-level selective coverage, lower-tail local q0.05 , and return rate conditional on native applicability, respectively. E1, E2, and E3 denote the adapted continuous-treatment, weighted distribution-shift, and spatial/local conformal families; M6 denotes GeoDose-CP using the primary estimated treatment density. Dashed horizontal lines in the upper two panels denote nominal 0.90 coverage. E1 is natively inapplicable in the poor-overlap atomic-target configuration, so no conditional return-rate value is defined there; this is nonapplicability rather than refusal.
Fig. 6: Spatial design and empirical uncertainty diagnostics for the NSW demonstration at 90-m support. (a) Buffered partition for Mt Arthur Coal, HVO, and Bulga Complex, including nuisance-training, support, calibration, held-out test, buffer, secondary, and ineligible regions. (b) M3 conformity p-values for held-out RF-primary targets; lower values indicate weaker conformity with the fitted spatial reference. The p-values are empirical diagnostics, not causal evidence or theorem-certified coverage guarantees.
Fig. 7: Held-out NSW observed-product uncertainty at 90-m support under the primary RF track. (a) Pooled observed-product coverage for M1, M3, and the treatment-free graph-safe fallback GS over the held-out test frame; the dashed horizontal line denotes nominal 0.90 coverage. (b) Mean interval width for M1, M3, and GS evaluated on the same frozen 60-target outcome-blind subset used for finite-domain M3 inversion. The width comparison is therefore restricted to this common target subset and is not interpreted as population-level efficiency. M2, M4, and M6 are nonoperational because an authentic longitudinal rehabilitation treatment is unavailable in the NSW archive.
Estimating treatment effects from observational data requires choosing an adjustment set, but valid adjustment depends on an unknown causal graph. Graph misspecification can cause under-coverage, while graph-agnostic conformal wrappers may regain nominal coverage only through large padding. We introduce CausalGuard, a structure-weighted conformal framework that calibrates after aggregating graph-conditional doubly robust pseudo-outcomes. Candidate DAGs are proposed from an LLM-derived edge prior, pruned by conditional-independence tests, and reweighted by Bayesian Information Criterion. A composite nonconformity score then calibrates the posterior-weighted pseudo-outcome. CausalGuard provides distribution-free finite-sample marginal coverage for this aggregated pseudo-outcome; under causal identification, overlap, conditional-mean nuisance stability, and concentration on target-aligned valid adjustment strategies, its conditional mean converges to the true Conditional Average Treatment Effect. Across five benchmarks, CausalGuard attains mean coverage above the nominal 90% level for the directly evaluable target and reduces width when graph-agnostic conformal baselines require large padding. Stress tests show that CausalGuard suppresses invalid collider adjustment and remains stable under misspecified priors when the retained candidate set is data-supported.
Selective conformal prediction can yield substantially tighter uncertainty sets when we can identify calibration examples that are exchangeable with the test example. In interventional settings, such as perturbation experiments in genomics, exchangeability often holds only within subsets of interventions that leave a target variable "unaffected" (e.g., non-descendants of an intervened node in a causal graph). We study the practical regime where this invariance structure is unknown and must be estimated from data. Our main result quantifies how coverage degrades when the estimated safe calibration set accidentally includes interventions that affect the target, and gives a conservative correction when an upper bound on this error is available. Rather than learning a full causal graph, we learn only the intervention-target relationships needed to choose calibration interventions. We give algorithms for this partial learning task and evaluate them on synthetic structural equation models and Replogle K562 CRISPR-interference data, where the experiments illustrate synthetic gains from selective calibration and finite-sample tradeoffs on real perturbation screens.
Amir Asiaee, Kavey Aryan, James P. Long
Department of Biostatistics, Vanderbilt University Medical Center, TN 37203, USA · Department of Informatics, King’s College London, London, UK · Department of Biostatistics, MD Anderson Cancer Center, Houston, TX, USA
Causal inference in spatial domains faces two intertwined challenges: (1) unmeasured spatial factors, such as weather, air pollution, or mobility, that confound treatment and outcome, and (2) interference from nearby treatments that violate standard no-interference assumptions. While existing methods typically address one by assuming away the other, we show they are deeply connected: interference reveals structure in the latent confounder. Leveraging this insight, we propose the Spatial Deconfounder, a two-stage method that reconstructs a substitute confounder from local treatment vectors using a conditional variational autoencoder (C-VAE) with a spatial prior, then estimates causal effects with a flexible outcome model. We show that this enables nonparametric identification of direct and spillover effects under weak assumptions--without multiple treatment types or a known latent-field model. Empirically, we extend SpaCE, a benchmark suite for spatial confounding, to include treatment interference, and show that the Spatial Deconfounder consistently improves effect estimation across real-world environmental health and social science datasets. By turning local interference into a multi-cause proxy for latent spatial confounding, our framework advances robust causal inference for spatial data.
Ayush Khot, Miruna Oprescu, Maresa Schröder +2
University of Illinois at Urbana-Champaign · Cornell University, Cornell Tech · LMU Munich, Munich Center for Machine Learning (MCML) +1