Root cause analysis (RCA) is a critical problem in many real-world scenarios. RCA enables the identification of faulty or failing mechanisms in a system by comparing anomalous observations with corresponding reference (i.e., regular) observations. However, existing approaches rely either on heuristic methods or on conditional independence tests with a strong unconfoundedness assumption, and thus fail to exploit other complicated distributional constraints in the presence of latent variables. To relax these assumptions, we model the underlying system as a causal model and the anomalous system as a change in the structural functions of the same causal model. Specifically, to handle unobserved confounders, we establish an implicit connection between distributional constraint testing and root cause analysis. To adapt our approach to data generated from arbitrary causal models, we employ the deep causal model (DCM) framework, in which we design the causal model using neural networks. Finally, we illustrate how our method, RCA-DCM, can utilize different levels of partial graphical knowledge to perform RCA. We evaluate RCA-DCM against state-of-the-art baselines on simulated datasets, a physics-based causal chamber and two micro-service applications. RCA-DCM improves top-1 accuracy over the strongest baseline on both Sock Shop (0.880 vs. 0.752) and Online Boutique (0.776 vs. 0.712), and when the true root cause in the causal chamber is unobserved and acts as a latent confounder, it recovers the exact root-cause set more often than any competing method (perfect recovery rate (PRR) 0.846 vs. 0.731).
Figures & tables
Figure 1: Baselines fail when latent confounders are present and non-root-cause variables experience larger distributional shifts than the root causes; Given F→T , deciding whether Y is a root cause ( F→Y ) is equivalent to deciding whether F is a valid instrument for the effect of T on Y .
Figure 2: (a) No confounding. (b)–(c) Bow graphs with an unobserved confounder between X and Y (dashed ↔ ).
Figure 3: Workflow example. In the true augmented graph the fault shifts X and Z , so R∗={X,Z} , and latent confounders act on both X↔Z and Z↔Y . The induced pair P∗=(Pn∗,Pa∗) falls outside the shaded Valid-IV region of the input space: it violates the instrumental inequalities for the effect of X on Z , which rules out F being excluded from Z (exclusion condition) and hence forces the edge F→Z . Each candidate V is tested by freezing fV while the mechanisms on V∖{V} shift freely. The violation makes V∖{Z} infeasible, so ℓ∗(V∖{Z})>0 certifies that Z lies in every R∈SolRC(G∗) and hence in R . For V=Y the score attains zero, which is inconclusive: it is consistent with V∖{Y}∈SolRC(G∗) , and we draw no conclusion from it.
Figure 4
Figure 5: Perfect recovery rate (PRR) with 95% confidence intervals. (a) increasing latent confounding, (b) heterogeneous anomalies where a non-cause can out-shift a true cause, (c) real causal-chamber data with the root cause observed vs. hidden. RCA-DCM remains strong across all three settings; comparisons use a paired test, so overlapping intervals do not imply the absence of a difference.
Sock Shop
Online Boutique
Method
T1
T3
T5
T1
T3
T5
DCM
0.88
0.99
0.99
0.78
0.90
0.98
NSigma
0.75
0.97
0.99
0.70
0.91
0.96
BARO
0.74
0.98
0.99
0.62
0.90
0.94
CIRCA
0.65
0.96
0.99
0.52
0.89
0.98
RCD
0.38
0.54
0.58
0.49
0.59
0.62
Table 1: Average Top- k accuracy over the five fault types (best in bold).
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Ns
Na
DCM
BARO
RCD
#Instances
1000
1000
0.63
0.39
0.32
78
1000
500
0.58
0.48
0.48
95
500
500
0.47
0.47
0.31
51
1000
100
0.27
0.35
0.24
131
100
100
0.23
0.24
0.02
62
Appendix
Table 2: Perfect recovery rate (PRR) under varying normal ( Ns ) and anomalous ( Na ) sample sizes.
CPU
MEM
DISK
DELAY
LOSS
Average
Dataset
Method
T1
T3
T5
T1
T3
T5
T1
T3
T5
T1
T3
T5
T1
T3
T5
T1
T3
T5
Sock Shop
DCM
1.00
1.00
1.00
0.96
0.96
0.96
0.76
1.00
1.00
0.92
1.00
1.00
0.76
1.00
1.00
0.88
0.99
0.99
NSigma
1.00
1.00
1.00
0.96
0.96
0.96
0.64
0.96
1.00
0.60
0.92
1.00
0.56
1.00
1.00
0.75
0.97
0.99
BARO
1.00
1.00
1.00
0.48
0.96
0.96
0.68
0.96
1.00
0.92
1.00
1.00
0.64
1.00
1.00
0.74
0.98
0.99
CIRCA
0.96
1.00
1.00
0.56
0.96
0.96
0.68
0.84
1.00
0.48
1.00
1.00
0.56
1.00
1.00
0.65
0.96
0.99
RCD
0.76
0.92
0.96
0.28
0.48
0.48
0.32
0.52
0.52
0.28
0.48
0.52
0.28
0.32
0.44
0.38
0.54
0.58
Appendix
Table 3: Per-fault Top-1/Top-3/Top-5 accuracy. Best value per column within each dataset is in bold.
Method
Typical time / dataset
BARO
∼0.01 s
RCD
∼0.4 s (median)
DCM (ours)
∼41 s (median)
Appendix
Table 4: Per-dataset wall-clock time in the synthetic setting ( 12 variables, 4 latents, Ns=Na=1000 ).
Figure 6: Non-rc Y has higher shift in median than RC X .
Signal Processing Laboratory, São Carlos School of Engineering, University of São Paulo, São Paulo, Brazil · Guaratinguetá School of Engineering, São Paulo State University, Guaratinguetá, Brazil