stat.MLAug 4, 2026

Causal Inference with Unstructured Outcomes

Authors: Kevin Christian WibisonoYixin Wang

Organizations: Department of Statistics University of Michigan, Ann Arbor

Abstract

Causal inference has traditionally centered on scalar outcomes: whether a patient recovers, how much a worker earns, or how many visits a website receives. Modern studies increasingly ask causal questions about outcomes with richer form, such as clinical notes, open-ended survey responses, and images. A hospital may want to know how an AI documentation tool changes the notes physicians write, or how a nurse training program alters what patients say in survey responses. For such outcomes, the usual average treatment effect is ill-defined: one cannot meaningfully subtract one text or image from another. To this end, we propose a causal query for unstructured outcomes. The key idea is to learn what features of the outcome are most causally affected by the treatment, which we call the maximally contrasting feature (MCF). To estimate the MCF, we learn a feature-scoring function that maps each outcome to a scalar and exposes the sharpest contrast between treated and control potential outcomes. We develop identification conditions and estimation algorithms for this query, and extend it to heterogeneous effects by allowing the feature-scoring function to depend on observed covariates. We also handle settings where both the treatment and the outcome are unstructured. Empirical studies on text and images show that the algorithm recovers salient aspects of an outcome changed by a treatment.

Explore similar work

Aug 1, 2026stat.ML

Causal Inference with Unstructured Treatments

Causal inference usually concerns a scalar treatment, yet in many problems the treatment is unstructured: a text, an image, or a sequence of clinical decisions. Consider an instructor writing a course description to attract more students: the treatment is the course description, and the outcome is enrollment. The standard target, the average treatment effect of fixing the treatment to one exact value versus another, runs into two problems. It cannot be estimated, because almost no exact description recurs across courses, leaving no comparable group from which to measure its effect; and it would be of little use even if it could, since no one wants every course to carry the same description. What the instructor actually wants to know is which features of a description raise enrollment, and which of those features can be acted on across many courses. To this end, we propose a causal query for unstructured treatments: the maximally influential feature (MIF), the feature of the treatment that most strongly influences the outcome. We formalize the MIF as a binary feature of the treatment, defined by a feature-scoring function, constrained so that both of its values stay well populated, and chosen to maximize the causal effect it induces. Turning the feature on shifts the distribution of treatments toward those that display it, turning it off shifts away, and the MIF effect contrasts the two average potential outcomes. We study identification conditions for the MIF, develop algorithms to estimate it, and make it actionable through a nudging algorithm that revises a treatment along the MIF into an outcome-improving version. We illustrate the MIF algorithm across applications in text, image, and dynamic treatment sequences.
Kevin Christian Wibisono, Yixin Wang
Jun 25, 2026cs.LG

A Causal Foundation Model for Structure and Outcome Prediction

We introduce TabPFN-CFM, a causal foundation model that can handle multiple causal problems. TabPFN-CFM predicts both causal structure and outcomes from observational data, supports queries on all three levels of Pearl's Causal Hierarchy and uses known graph structure when available to improve predictions. TabPFN-CFM is trained on synthetic datasets, and generalises to real datasets, demonstrating improved performance over both structural and outcome prediction baselines.
Max Zhu, Martino Mansoldo, Ching-Hao Wang +1
Sep 15, 2026stat.ME

Information Set Emulation: Causal Certificates for AI Derived EHR Features

AI and large language models can recover clinically meaningful features from electronic health records (EHRs), but predictive usefulness does not establish admissibility for causal inference. We introduce information set emulation: an AI typed lift attaches source evidence, clinical and recording times, decision-time availability, representation version, proposed causal roles, and unresolved ambiguity to extracted features under a locked target trial. Causal certificates record auditable evidence for those roles. Features with unresolved downstream roles are routed to compatible reporting or separate analyses. Typed evidence defines an observational fiber of causal worlds consistent with the observed law. The locked scalar estimand maps this fiber to a compatible image whose squared Chebyshev radius equals the residual minimax mean squared error when the image is nonempty and compact. This classical identity provides a target-specific measure of information ambiguity. The contribution is its integration with a joint EHR observation map and an auditable certificate architecture. Under explicit exchangeability, positivity, and nuisance-consistency conditions, we give identification and cross-fitted augmented inverse probability weighted estimation, distinguishing empirical and population targets. An EHR compression-drift identity separates the roles of frame presence, treatment assignment, and outcome observation. Artificial simulations and a common-law finite-world example illustrate estimation failures and information-radius reduction. Synthetic Phase 0 notes demonstrate audit diagnostics; a separate role-specific analysis spread illustrates routing and is not an exact fiber radius. All experiments are synthetic. The framework specifies when reconstructed information can support a point claim and when compatible reporting is required.
Takes Fujita, Nobutaka Hattori