Many real-world machine learning tasks are anti-causal: they require inferring latent causes from observed effects. In practice, we often face multiple related tasks where the structural dependencies are a hybrid of task-invariant and task-specific mechanisms. We propose Multi-Task Anti-Causal learning (MTAC), a framework for estimating causes from outcomes and confounders by explicitly exploiting such cross-task invariances. MTAC learns a structural equation model (SEM) that factorizes the outcome-generation process into (i) a task-invariant mechanism and (ii) task-specific mechanisms via a shared backbone with task-specific deviations. Building on the learned forward model, MTAC performs maximum A posteriori (MAP) based inference to reconstruct causes by jointly optimizing latent mechanism variables and cause magnitudes under the learned structural model. We evaluate MTAC on the application of urban event reconstruction from resident reports, spanning three tasks: parking violations, abandoned properties, and unsanitary conditions. On real-world data collected from Manhattan and Newark, MTAC consistently improves reconstruction accuracy over strong baselines, achieving up to 33.04% MAE reduction and demonstrating the benefits of learning transferable mechanisms across tasks.
Figures & tables
Figure 1: Multi-task causal graph with shared causal mechanism.
Figure 2: MTAC Framework. The white nodes represent parameterized deterministic neural networks, gray nodes represents drawing samples from the respective distribution. The left panel shows the multi-task structural equation model, in which the mechanism variables W and confounders Z are shared across tasks while causes Xk and outcomes Yk are task-dependent. The orange box details the multi-task mechanism module for a single mediator W1 . The right panel illustrates the MAP inference procedure.
Figure 3: Multi-task SEM for mechanism variable Wi . The norm of first-layer weights decides whether input variables are parents.
Model
Parking Violation
Abandoned Property
Unsanitary Condition
MAE
MSE
MAE
MSE
MAE
MSE
MTAC
0.2971
0.1616
0.4163
0.1910
0.3755
0.2501
CEVAE
0.3235
0.1925
0.4867
0.2607
0.4091
0.2884
TEDVAE
0.4315
0.2636
0.4932
0.2583
0.4188
0.2700
BSM-UR
0.4437
0.2976
0.5438
0.3175
0.3982
0.3026
PLE
0.3337
0.2174
0.4533
0.2351
0.4101
0.2697
Table 1: MAE and MSE estimation for MTAC and baselines across three tasks.
Parking Violation
Abandoned Property
Unsanitary Condition
Model
MAE
MSE
MAE
MSE
MAE
MSE
Multi-task
0.2971
0.1616
0.4163
0.1910
0.3755
0.2501
Single-task
0.3693
0.2947
0.4974
0.2517
0.4486
0.2989
Table 2: Performance comparison for MTAC between multi-task training and single-task training.
UC+PV → AP
UC+AP → PV
PV+AP → UC
Model
MAE
MSE
MAE
MSE
MAE
MSE
Full fine-tuned
0.4507
0.2313
0.3475
0.2502
0.3977
0.2656
Deviation-only fine-tuned
0.4631
0.2297
0.3407
0.2586
0.4106
0.2694
Zero-shot
0.7702
0.4323
0.7699
0.5472
0.8965
0.7518
Single-task trained
0.4974
0.2517
0.3693
0.2947
0.4486
0.2989
Table 3: Prediction Error of transferred MTAC. The transferred models are trained on two tasks and transferred to the remaining task. PV: parking violation; AP: abandoned property; UC: unsanitary condition.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Factors
Finance
mean income, unemployment rate, mortgage ratio, poverty rate, housing cost
Education Attainment
% less than high school, % high school, % bachelor or higher
Race & Culture
% Hispanic, % white, % black, % Asian
Access to Technology
% has computer, % has smartphone, % has internet
Social Environment
population, % multifamily, % owner occupied, % renter occupied, median year built, % room occupation ≥0.5 , mobility rate
Appendix
Table 4: An overview of the socioeconomic status factors.
Model
Parking Violation
Abandoned Property
Unsanitary Condition
MAE
MSE
MAE
MSE
MAE
MSE
Complete MTAC
0.2971
0.1616
0.4163
0.1910
0.3755
0.2501
w/o MAP
0.5841
0.6977
0.7841
0.6006
0.5957
0.7426
w/o Causality
0.3916
0.1811
0.4967
0.2842
0.4287
0.2938
Appendix
Table 5: Ablation study. For MTAC without MAP, the cause is estimated directly from the forward causal model.
Figure 4: Sensitivity of reconstruction performance on model parameters.
Causal discovery, the problem of inferring the direction of causality, is generally ill-posed. We use the language of structural causal models (SCM) to show that assuming that the causal relations are acyclic and invariant across multiple environments (e.g., the way minimum wage affects employment rate is stable across different geographical regions), \textit{only} two auxiliary environments are sufficient to infer the causal graph for arbitrary nonlinear mechanisms. Moreover, we demonstrate that this implies identifiability of the SCM functional mechanisms: as a corollary, we show that \textit{two} auxiliary environments are sufficient to guarantee correct counterfactual inference. We empirically support our theoretical results on synthetic data.
Event-centric intelligent analytical systems heavily depend on explicit causal event knowledge for risk early warning, decision-making support and narrative comprehension. Nevertheless, existing instance-level causal pairs suffer severe generalization deficits on low-frequency long-tail and unseen event combinations. To address this limitation, this work proposes Abstract Event Causal Rule (AECR), a novel relation-level causal abstraction paradigm that transforms concrete cause-effect pairs into generalized abstract causal logic while retaining their intrinsic causal relationships. We design a multi-agent Concrete-to-Abstract Causal Induction (CACI) system coupled with similarity-constrained clustering to distill trustworthy AECRs from noisy raw causal data, based on which two complete AECR knowledge bases are built. To validate the practical utility of abstract causal knowledge, we propose an Abstract Rule-Guided Causal Attention Encoder (AR-GCAE), which injects the retrieved AECRs into the causality Graph Event Prediction (CGEP) benchmark task via rule-guided attention layers and gated representation fusion. Quantitative experimental results reveal that applying AECRs substantially strengthens the generalization capacity of event causal reasoning and brings consistent performance improvements to event prediction, with the most prominent gains observed on rare and unseen event samples.
Ziwei Zheng, Peiqiong Chen, Bang Wang
School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan, China
Causal discovery from observational data remains challenging due to the need to recover directed structure and latent confounding without interventions. We propose FoundCause, an amortized causal discovery model trained entirely on synthetic data that maps datasets directly to causal graphs in a single forward pass. By learning from large collections of simulated structural causal models, FoundCause captures transferable statistical patterns that generalize beyond individual datasets. The architecture incorporates several key inductive biases for causal discovery. It uses a permutation-invariant transformer encoder with alternating attention over samples and variables to jointly model cross-variable dependence and per-variable distributions. Pairwise statistical features derived from classical asymmetry measures are injected through statistics-conditioned attention, guiding the model toward known causal signals. A factorized decoder separates edge existence from direction, while a triangular refinement module enables reasoning over higher-order causal motifs such as chains and colliders. In addition, a dedicated confounder module based on learnable latent tokens explicitly models hidden common causes, and the model explicitly handles missing data via its masked input representation. To our knowledge, FoundCause is the first amortized causal discovery approach to explicitly model latent confounding. FoundCause outperforms 11 classical non-amortized methods (e.g., PC, GES, NOTEARS-style optimization) and 4 amortized causal discovery methods on 15 real-world datasets, achieving +9.6% improvement in F1, +1.2% in AUROC, and an 18.9% reduction in structural Hamming distance relative to the strongest non-amortized methods, while performing inference in a single forward pass.
Patrick Blöbaum, Krishnakumar Balasubramanian, Shiva Prasad Kasiviswanathan
Amazon Web Services · Department of Statistics, University of California, Davis