Probabilistic Counterfactual Inference for Discrete Outcomes in Gaussian-Process Causal Models
Authors: Juliette Sinnott, Amir-Hossein Karimi, Mohammad Kohandel
Organizations: Department of Applied Mathematics, University of Waterloo, Waterloo, Ontario, N2L 3G1, Canada · Department of Electrical and Computer Engineering, University of Waterloo, Waterloo, Ontario, N2L 3G1, Canada · Vector Institute for Artificial Intelligence, Toronto, Ontario, M5G 0C6, Canada
Counterfactual inference in Gaussian-process structural causal models (GP-SCMs) has been developed primarily for continuous endogenous variables, limiting applicability to causal graphs that contain discrete child nodes with continuous parents. We introduce a unified probabilistic framework for counterfactual inference with heterogeneous variable types by pairing GP predictors with explicit exogenous noise mechanisms. For discrete outcomes, we derive exact conditional noise-abduction procedures using a uniform threshold for binary variables, a Gumbel-max race for nominal categories, and a latent Gaussian cut-point model for ordinal ones. In each case, we propagate abducted noise through interventions while accounting for posterior uncertainty in the GP latent functions, and prove that the resulting mechanisms reproduce the fitted model's observational and interventional distributions. On synthetic SCMs with known ground-truth counterfactuals, we evaluate estimation accuracy, consistency, and robustness to coupling misspecification. A key finding is that applying a categorical coupling to ordinal data inflates counterfactual error roughly threefold even when observational fit remains comparable, and that this error does not diminish with more data. As the training set grows, the fitted structural equation converges to the truth while the counterfactual error flattens onto a floor. In the reverse direction, forcing a false order onto nominal data instead degrades the fitted equation itself. The choice of coupling must therefore be justified on structural grounds rather than read off the fit.
Figures & tables
Figure 1: The four synthetic SCM structures used in Section 4 (Binary, Single-parent, Interaction and Chain), increasing in graph complexity and node type. Node fill color indicates each variable’s data type.
Categorical fit
Ordinal fit
SCM
Node type
Mediator
TVint
TVcf
TVint
TVcf
Binary
binary
0.021±0.010
0.021±0.014
0.018±0.009
0.015±0.007
Single-parent
categorical
0.030±0.010
0.031±0.015
0.184±0.031
0.173±0.033
Interaction
ordinal
0.078±0.016
0.231±0.100
0.045±0.011
0.068±0.040
learned cutpoints
ordinal
—
—
0.049±0.010
0.070±0.041
Chain, X3
categorical
categorical
0.070±0.012
0.091±0.017
0.223±0.021
0.221±0.062
Table 1: Counterfactual error TVcf and interventional error TVint across all four synthetic SCMs, under a categorical fit and under an ordinal fit with learned cutpoints; mean ± sd over 10 seeds at n=500 . Mediator applies only to the Chain SCM’s X3 : whether the upstream mediator X2 was fit as a categorical or ordinal node, isolating propagated error from X3 ’s own.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
arm
n=100
n=250
n=500
n=1000
Binary (probit)
0.040
0.022
0.021
0.015
Single-parent (Gumbel-max)
0.061
0.045
0.031
0.021
Interaction (ordinal) †
0.253
0.226
0.231
0.207
Chain, X3 (Gumbel-max)
0.119
0.084
0.091
0.071
mediator X2 (ordinal) †
0.279
0.249
0.274
0.242
Binary, refit
0.039
0.021
0.014
0.013
Appendix
Table 2: Counterfactual error TVcf against training-set size, mean over 10 seeds. Rows marked † use an estimator whose coupling does not match the true mechanism; their error flattens rather than falling with n .
estimator
Single-parent
Interaction
interventional (no abduction)
0.295±0.079
0.434±0.051
latent-only abduction
0.300±0.080
0.430±0.051
plug-in posterior mean
0.031±0.014
0.232±0.102
full estimator (ours)
0.031±0.015
0.231±0.100
oracle: true p , same coupling
0.005±0.001
0.205±0.092
Appendix
Table 3: Counterfactual error TVcf for five estimators on the same data, mean ± sd over 10 seeds at n=500 . The Single-parent SCM’s true mechanism is Gumbel-max, matching the coupling; the Interaction SCM’s is the ordinal cut-point model, so the coupling is misspecified.
The presence of confounding bias poses a key challenge in policy evaluation, as the target causal effects of actions are not identifiable (i.e., underdetermined) from observational data. On the other hand, existing confounding-robust evaluation strategies require detailed prior knowledge about the environment or apply only to discrete treatments and outcomes. This paper investigates causal effect evaluation over the continuous domain from confounded observations, while requiring only basic temporal ordering between the treatment and the outcome. We introduce a universal discretization of the exogenous domains that approximates the observational and interventional distributions of any causal model with arbitrary accuracy using a finite number of latent states. Building on this newfound universal approximation property, we develop a novel family of Causal Gaussian process (CGP) models that effectively approximate the observational and interventional distributions of any causal model with confounded observations.
Junzhe Zhang, Jingyuan Chen, Elias Bareinboim
Department of Electrical Engineering and Computer Science, Syracuse University · Department of Computer Science, Columbia University
Causal discovery, the problem of inferring the direction of causality, is generally ill-posed. We use the language of structural causal models (SCM) to show that assuming that the causal relations are acyclic and invariant across multiple environments (e.g., the way minimum wage affects employment rate is stable across different geographical regions), \textit{only} two auxiliary environments are sufficient to infer the causal graph for arbitrary nonlinear mechanisms. Moreover, we demonstrate that this implies identifiability of the SCM functional mechanisms: as a corollary, we show that \textit{two} auxiliary environments are sufficient to guarantee correct counterfactual inference. We empirically support our theoretical results on synthetic data.
We study causal structure learning from observational data in linear Gaussian structural causal models in the presence of directed cycles and an unknown number of exogenous latent confounders, bounded by a given maximum. We derive the covariance of the observed variables and introduce marginal quasi-equivalence, which characterizes when different causal models share a full-dimensional subset of the observational distributions they can generate. We formulate structure learning as minimization of the Gaussian negative log-likelihood with a logarithmically scaled complexity penalty that counts directed edges and latent variables. For a fixed number of observed variables and a fixed upper bound on latent variables, we establish consistency of global score minimizers up to marginal quasi-equivalence under algebraic faithfulness, structural minimality, and model-overlap assumptions. We parameterize the inclusion of directed edges and candidate latent variables using Bernoulli gates, whose continuous probabilities are optimized jointly with the structural coefficients. Averaging the penalized negative log-likelihood over these gates yields an objective with a closed-form differentiable complexity penalty. We prove that this expected objective has the same global infimum as the corresponding discrete structure-learning objective. Experimental results show that our approach achieves lower recovery error than previous methods in several experimental settings.
Sadegh Khorasani, Ali Najar, Saber Salehkaleybar +1
School of Computer and Communication Sciences, EPFL, Lausanne, Switzerland · Leiden Institute of Advanced Computer Science (LIACS), Leiden University, Leiden, The Netherlands · College of Management of Technology, EPFL, Lausanne, Switzerland