Causal AI is a branch of Artificial Intelligence which helps understand and reason about cause and effect relationships, not just patterns or correlations. Causal discovery aims to infer elements of the underlying causal structure--often represented as a directed graph--from observational and, when available, interventional data. While causal discovery is the fundamental step for moving beyond mere associations toward genuine understanding, and thus the basic building block of causal AI, it becomes intrinsically difficult when causal relations must be inferred from single observations. In such situations, standard causal discovery methods cannot be used and one has to identify causal relations from limited amount of information. This is typically the case for, e.g., sequences of events produced by different alarms which need to be analyzed on the fly to detect abnormal phenomena, which are usually rare. We show in this study that it is possible to leverage the predictive power of Large Language Models (LLMs) to infer causal relations between only two sequences of events. This approach, which is validated on both synthetic and real data, provides better results than standard causal discovery algorithms on several time series data, even though these data were converted into smaller, single observed sequences.
Figures & tables
Length range (# of symbols) (∣X∣,∣Y∣)
Class
16–41
32–83
128–166
256–332
X→Y
376
156
101
99
Y→X
372
139
99
111
Table 1 . Number of synthetic sequence pairs (X,Y) generated for each causal direction ( X→Y or Y→X ) and length range (number of symbols per sequence). Sequence symbols are chosen randomly from the set of all lower cased English alphabet and decimal numbers ({a…z,0…9}) .
Subset
Length Range (∣X∣,∣Y∣)
Class
# Pairs
MoM
14–191
X→Y
10
Storm Ingestion
52–170
X→Y
9
Web Activity
46–1000
X→Y
14
Antivirus Activity
19–256
X→Y
16
Table 2 . Statistics of discretized real-world datasets: number of sequence pairs (X,Y) (all with X→Y causal direction) for each subset, including length ranges (∣X∣,∣Y∣) . The event sequence symbols are randomly chosen from the set {a,b,c,d,e,f,A,B,C,D,E,F}.
Model / Setting
μ -F1
Standard prompting baseline
(a) Zero-Shot (No encoding, r=0 )
Llama-3.2-3B †
0.4868
Llama-3.1-8B †
0.5016
Llama-3.3-70B †
0.5110
(b) Zero-Shot (Encoded, r=0 )
Llama-3.2-3B †
0.4901
Table 3 . Average test performance ( μ -F1) on Synthetic data across all sequence lengths. † indicates the Instruct version of the model. Results compare three settings: (a) classical zero-shot LLM prompting, (b) zero-shot with one random encoding ( Encoded ), and (c) zero-shot with one random encoding and one replication ( r=1 ); with our method ( K=1111,r=1 ).
Length Range (∣X∣,∣Y∣)
μ -F1
( r=0 )
( r=1 )
16-41
0.5856
0.7888
32-83
0.5153
0.9322
128-166
0.9450
0.9900
256-332
0.9571
0.9667
Avg.
0.7477
0.9194
Table 4 . Effect of sequence lengths and replication ( r∈{0,1} ) on performance of our method (with Llama-3.2-3B, K=1111 ) on synthetic data.
Figure 1 . Curves showing performance variation of our method (with Llama-3.2-3B) as K varies between 0 (raw observed sequence) and 1111, for different sequence lengths.
Our Method
r
Length
# of Encodings ( K )
0
1
51
311
1111
with
r=0
16-41
0.5548
0.5495
0.5468
0.5508
0.5388
GPT-2
32-83
0.5186
0.5458
0.5017
0.4678
0.4475
128-166
0.4550
0.6450
0.7100
0.7000
0.7200
256-332
0.5238
0.6619
0.6286
0.6952
0.6619
Avg.
0.5131
0.6005
0.5968
0.6035
0.5920
Table 5 . Detailed performance ( μ -F1 scores) on synthetic data using GPT-2 and Llama-3.2-1B backbones. † = sequence truncated to 256 symbol length for GPT-2 context limit.
Table 6 . Performance ( μ -F1, mean ± standard deviation) of Algorithm 1 on real data, after 10 balanced pair-and-label randomization. Zero-shot baseline uses discretized sequences with random encoding for r∈ 0,1 across 50 iterations. Our method (with K=1111) is run over four random seeds. † indicates instruct version of the respective LLM
Causal discovery is a cornerstone of scientific reasoning, yet whether large language models can perform it reliably remains an open question. Recent benchmarks show that even fine-tuned models plateau on simple causal graphs and degrade as complexity grows, but why they fail has not been established. We prove the failure is fundamental: supervised fine-tuning, direct preference optimization, and in-context learning all produce predictors that cannot distinguish between causal graphs generating similar observational data, and any attempt to do so requires the model's internal representations to grow unboundedly, violating the very conditions under which these methods work. We formalize this as a kernel obstruction theorem, establishing that the limitation is intrinsic to the learning paradigm, \emph{not any particular model or dataset}. We propose Agentic Causal Bayesian Optimization (A-CBO), wherein a frozen language model serves as an interventional oracle answering targeted queries about intervention effects, while an external Bayesian loop concentrates beliefs over candidate graphs in logarithmically many rounds. Because the decision operates outside the space where the obstruction applies, A-CBO provably converges while the underlying model remains unchanged. On Corr2Cause, A-CBO matches fine-tuned baselines without any training. On Extended Corr2Cause, a new benchmark scaling to 24 variables with 18K test samples, A-CBO significantly outperforms both fine-tuning and preference optimization, with the advantage growing
Amartya Roy, Sonali Parbhoo
SIRE, IIT Delhi and Robert Bosch GmbH, India · Imperial College London
We introduce CausaLab, a scalable environment for evaluating interactive causal discovery by LLM agents. Unlike prior evaluations, CausaLab evaluates both whether an agent can solve a problem using causal evidence and whether its answer is grounded in a faithful recovered causal mechanism. Each episode places an agent in a synthetic laboratory: it receives prior measurement records, intervenes on a manipulator crystal, and predicts the resonance frequency of a held-out reactor crystal governed by the same mechanism. The hidden data-generating process is a randomly sampled structural causal model (SCM), so success requires recovering both a causal graph and structural equations rather than recalling prior knowledge. Experiments show a persistent gap between prediction and mechanism recovery: in the purely observational 6-node setting, GPT-5.2-high reaches 92% task accuracy but only 0.471 all-edge F1. Mixed observation-intervention strategies improve structural fidelity, while pure intervention remains difficult even for strong agents. We identify premature stopping as a major weakness and show that consistency verification mitigates it. CausaLab therefore separates predictive success from causal understanding and exposes current LLM agents' limits as experimental causal reasoners.
Junlin Yang, Dylan Zhang, Xiangchen Song +7
Tsinghua University · University of Illinois Urbana-Champaign · Carnegie Mellon University +2
Textual event records, such as alarm logs, have become an increasingly common data source in engineering and manufacturing systems. Beyond identifying correlations or recurring patterns, engineers are often interested in understanding which types of events causally trigger or influence other events during system operation. Textual event descriptions may contain semantic clues about such causal relationships, and recent large language models (LLMs) provide a promising tool for extracting these signals. However, relying solely on LLM-encoded textual information is insufficient for accurate causal discovery, since semantic patterns do not directly reveal causal mechanisms and may confuse causation with correlation or frequent sequential patterns. To address these challenges, we propose \textbf{LMT}, a Bayesian causal discovery framework for engineering event data that jointly leverages textual descriptions and timestamps. Specifically, LMT first uses LLMs to extract semantic causal signals from event descriptions and constructs a prior distribution over causal graphs among event types or event clusters. It then incorporates temporal evidence through a Poisson-process-based likelihood, allowing the LLM-informed prior to be refined by timestamp-based statistical evidence. By integrating the textual and temporal information, LMT produces a causal graph that is both interpretable and data-supported. Simulation studies show that the proposed framework is effective across different settings and is especially advantageous in small-sample alarm-event scenarios.
Xiaofeng Xiao, Jianhong Chen, Qiuzhuang Sun +2
Department of Mechanical & Industrial Engineering, Northeastern University, Boston, MA, USA · College of Integrative Studies, Singapore Management University, Singapore · Department of Industrial Engineering and Management Sciences, Department of Mechanical Engineering, Northwestern University, IL, USA