Causal Lag Structure Discovery in Confounded Time Series via Orthogonalized Adaptive Estimation
Authors: Hong Kiat Tan, Isaac-Neil Zanoria, James Chen, Haoyang Lyu, Mihai Cucuringu
Organizations: Department of Mathematics, University of California, Los Angeles · Department of Electrical and Computer Engineering, University of California, Los Angeles · Oxford-Man Institute of Quantitative Finance, University of Oxford · Department of Statistics, University of Oxford
Finding which variables cause which others in multivariate time series, and at what lags, is central to science and policy, yet existing methods force a choice between flexible confounder adjustment, data-driven lag selection, and inference that controls the false discovery rate (FDR). ORACLE-VARX does all three in one pipeline. First, double/debiased machine learning (DML) removes nonlinear confounder effects from the outcomes and the lagged series. Second, adaptive causal lag estimation (ACLE) picks the lag order at each time step by sequential significance tests, tracking regime changes. Third, entry-wise z-tests with Benjamini--Hochberg correction select directed edges at a target FDR. We prove that in each rolling window, the debiased coefficients are asymptotically normal around a window-averaged target, so their z-tests are asymptotically valid. On a synthetic benchmark with time-varying structure and nonlinear confounding, ORACLE-VARX (LightGBM) tracks the true lag order best (RMSE 0.96 vs 1.1--1.5), has edge FDR 0.047, close to PCMCI (0.045) and below VAR (0.129) and VAR-LiNGAM (0.187), and forecasts better than all three. On nine U.S. sector ETFs with macroeconomic confounders, it yields interpretable causal graphs whose lag order rises in high-volatility regimes.
Figures & tables
Figure 1: An example of a time-unrolled causal graph of a 3-variable system. Solid blue arrows show the lag-1 cycle (coefficient 0.60 ). Dashed red and dotted green arrows show lag-2 and lag-3 self-effects a(t) and b(t) that decay to zero (Section 5.1 ). Gray dashed arrows indicate nonlinear confounder effects.
Figure 2: Sequential lag testing in ACLE. Left: the model with lags 1,…,p−1 . Right: testing H0:Ap=0 on the lag- p block. If rejected at level α , lag p is retained and we proceed to p+1 ; otherwise we stop and set pα=p−1 .
Method
Lag Sel.
Conf. Adj.
DML
VAR
RMSE
None
–
VARX
RMSE
Linear
–
ACLE-VAR
ACLE
None
–
ACLE-VARX
ACLE
Linear
–
OR-VARX
RMSE
Nonlinear
✓
ORACLE-VARX
ACLE
Nonlinear
✓
Table 1: Hierarchy of methods. Each row adds one capability relative to the row above. “RMSE” lag selection picks the lag order minimizing rolling validation RMSE ( 9 ) over {1,…,pmax} ; “ACLE” uses significance-based sequential testing with validation-based α -tuning (Section 3.3 ).
Method
Lag
Coeff.
Fcst.
FDR
Power
RMSE
MAE
MAE
↑
No confounders
VAR
1.475
0.060
0.110
0.129
0.655
ACLE-VAR
1.006
0.066
0.109
0.147
0.661
All confounders, no DML
VARX
1.239
0.063
0.109
0.076
0.650
Table 2: Synthetic benchmark results (All confounders observed unless noted). ET: Extra Trees; LGBM: LightGBM; TP: TabPFN. FDR: BH at q=0.05 . Bold: best per column; lower is better except for Power.
Figure 3: Estimated vs. true edge coefficients over time for ORACLE-VARX (Extra Trees, all confounders). Each panel shows one edge; the black line is the true coefficient Ak∗[i,j](t) and the colored line is the rolling estimate A^k[i,j](t) . Vertical dashed lines mark regime boundaries ( t=1000 , t=2000 ).
Figure 4: Lag tracking: estimated p^t (blue) vs. true p∗(t) (black step function) over time. Left : ACLE-VAR (no confounders). Right : ORACLE-VARX (Extra Trees, all confounders). The lag RMSE of every method is in Table 4 .
Figure 5: ETF application (ORACLE-VARX, Extra Trees, all10 confounders). Blue step function: adaptively selected lag order p^t over time (2004–2020). Orange dashed line: 21-day realized volatility of the S&P 500 (right axis). The lag order spikes during the 2008 financial crisis and COVID-2020, rising with the volatility.
Method
vix
macro5
all10
VAR
1.08
ACLE-VAR
1.32
VARX
1.11
0.43
0.75
ACLE-VARX
1.07
0.81
0.86
OR-VARX, ET
0.71
0.71
0.67
ORACLE-VARX, ET
0.83
0.85
0.68
Table 3: ETF Sharpe ratios (weighted strategy, 2004–2020) by confounder set; VAR and ACLE-VAR use none.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Three scenarios illustrating the role of causal sufficiency in Proposition A.13 . (a) When all confounders are observed ( Wt ), DML correctly controls for them and the estimated coefficient A^k[i,j] reflects the true causal relationship. (b) An unobserved common cause Ut (dashed circle) of Y(j) and Y(i) creates a spurious association: DML estimates A^k[i,j]=0 even though no direct causal link exists. (c) An unobserved mediator Ut on the sole causal pathway from Y(j) to Y(i) : the indirect effect Y(j)→U→Y(i) is not lost. It appears as a nonzero A^k[i,j] at the lag where the two steps add up, so the edge is real but indirect. Causal sufficiency (Definition 5 ) rules out scenario (b) by requiring that all common causes are observed; scenario (c) does not violate it, because a mediator is not a common cause.
Method
Obs.
Nonzero MAE ↓
Fcst. MAE ↓
Lag RMSE ↓
FDR ↓
Power ↑
F1 ↑
Non-DML methods
VAR
none
0.0605
0.1100
1.475
0.129
0.655
0.734
ACLE-VAR
none
0.0655
0.1094
1.006
0.147
0.661
0.730
VARX
all
0.0627
0.1092
1.239
0.076
0.650
0.755
ACLE-VARX
all
0.0691
0.1081
1.099
0.080
0.651
0.754
VARX
partial-2
0.0604
0.1098
1.292
0.096
0.647
0.745
Appendix
Table 4: Complete synthetic benchmark results. All metrics are averaged over the rolling evaluation window. Obs.: observability level. FDR, Power and F1 use BH at level q=0.05 . VARX treats the observed confounders as extra series; its lag p^t is chosen on the equations of the n target series only, and only their edges are scored. Bold indicates the best value in each column.
Method
Obs.
Nonzero MAE ↓
Fcst. MAE ↓
Lag RMSE ↓
FDR ↓
Power ↑
F1 ↑
PCMCI + OLS refit
all
0.0646
0.1081
1.214
0.045
0.685
0.792
PCMCI + OLS refit
partial-2
0.0664
0.1089
1.223
0.052
0.672
0.780
PCMCI + OLS refit
partial-1
0.0664
0.1083
1.211
0.052
0.672
0.780
PCMCI + OLS refit
none
0.0657
0.1091
1.136
0.066
0.673
0.775
VAR-LiNGAM
all
0.0704
0.1074
1.132
0.187
0.641
0.703
VAR-LiNGAM
partial-2
0.0710
0.1082
1.132
0.196
0.641
0.697
Appendix
Table 5: Causal-discovery baselines on the synthetic benchmark, with ORACLE-VARX at full observability for reference. Same days and metrics as Table 4 .
Method
First-Stage Learner
Ann. Return (%)
Sharpe Ratio
VAR
—
5.93
1.08
ACLE-VAR
—
7.34
1.32
VARX
—
4.22
0.75
ACLE-VARX
—
4.80
0.86
OR-VARX
Extra Trees
3.51
0.67
ORACLE-VARX
Extra Trees
3.55
0.68
Appendix
Table 6: ETF portfolio performance (weighted strategy, market-adjusted). All DML methods use the all10 confounder preset (10 macroeconomic variables). Annualized return and Sharpe ratio are reported over the common test period (2004–2020).
Method
Learner
Naive
Weighted
Top 50%
Top 25%
Top 75%
VAR
—
4.32 / 1.04
5.93 / 1.08
6.16 / 0.95
5.12 / 1.01
8.41 / 0.86
ACLE-VAR
—
4.64 / 1.15
7.34 / 1.32
8.54 / 1.31
6.73 / 1.31
10.14 / 1.04
VARX
OLS
3.35 / 0.83
6.00 / 1.11
7.16 / 1.10
5.59 / 1.09
8.98 / 0.92
ACLE-VARX
OLS
3.09 / 0.78
5.83 / 1.07
6.47 / 1.00
5.85 / 1.14
8.76 / 0.90
OR-VARX
LightGBM
1.94 / 0.51
2.55 / 0.50
2.63 / 0.45
1.99 / 0.43
2.04 / 0.26
OR-VARX
XGBoost
0.98 / 0.27
1.43 / 0.29
2.07 / 0.36
1.79 / 0.38
3.21 / 0.38
Appendix
Table 7: Full strategy performance: vix confounder preset. Each cell shows Ann. Return (%) / Sharpe Ratio (market-adjusted, 2004–2020). VAR and ACLE-VAR use no confounders and are included as baselines. Best Sharpe per column in bold .
Method
Learner
Naive
Weighted
Top 50%
Top 25%
Top 75%
VAR
—
4.32 / 1.04
5.93 / 1.08
6.16 / 0.95
5.12 / 1.01
8.41 / 0.86
ACLE-VAR
—
4.64 / 1.15
7.34 / 1.32
8.54 / 1.31
6.73 / 1.31
10.14 / 1.04
VARX
OLS
0.87 / 0.23
2.34 / 0.43
2.51 / 0.40
1.51 / 0.32
2.99 / 0.35
ACLE-VARX
OLS
1.99 / 0.48
4.66 / 0.81
5.25 / 0.79
3.56 / 0.69
7.87 / 0.82
OR-VARX
LightGBM
0.11 / 0.05
0.47 / 0.11
− 0.94 / − 0.12
− 0.36 / − 0.05
0.52 / 0.10
OR-VARX
XGBoost
− 0.21 / − 0.03
− 0.83 / − 0.12
− 0.95 / − 0.12
− 0.16 / − 0.01
− 2.96 / − 0.27
Appendix
Table 8: Full strategy performance: macro5 confounder preset. Each cell shows Ann. Return (%) / Sharpe Ratio (market-adjusted, 2004–2020). Best Sharpe per column in bold .
Method
Learner
Naive
Weighted
Top 50%
Top 25%
Top 75%
VAR
—
4.32 / 1.04
5.93 / 1.08
6.16 / 0.95
5.12 / 1.01
8.41 / 0.86
ACLE-VAR
—
4.64 / 1.15
7.34 / 1.32
8.54 / 1.31
6.73 / 1.31
10.14 / 1.04
VARX
OLS
1.87 / 0.46
4.22 / 0.75
5.03 / 0.77
3.45 / 0.66
7.27 / 0.77
ACLE-VARX
OLS
1.89 / 0.47
4.80 / 0.86
5.68 / 0.86
3.38 / 0.65
5.98 / 0.66
OR-VARX
LightGBM
− 0.44 / − 0.09
0.13 / 0.05
− 0.33 / − 0.02
− 0.46 / − 0.07
1.44 / 0.19
OR-VARX
XGBoost
0.79 / 0.21
0.29 / 0.08
0.64 / 0.13
0.15 / 0.06
− 2.95 / − 0.27
Appendix
Table 9: Full strategy performance: all10 confounder preset. Each cell shows Ann. Return (%) / Sharpe Ratio (market-adjusted, 2004–2020). Best Sharpe per column in bold .
Figure 7: Cumulative market-adjusted returns for three representative configurations. (a) The ACLE-VAR baseline achieves the highest Sharpe ratios overall (up to 1.32). (b) ORACLE-VARX with Extra Trees and the all10 preset—the configuration used for the lag-volatility analysis in Figure 5 —produces stable, positive returns across all strategies. (c) The macro5 preset with Extra Trees yields the highest DML Sharpe ratio (1.15, Top 50% strategy) among all DML experiments.
Panel A: Experiment Configuration
Parameter
Synthetic
ETF
pmax
5
10
α -grid
{0.01,0.05,0.10,0.15,0.20,0.25,0.30}
OLS window
200 days
504 days
Tree training window
200 days
504 days
Validation days
20
21
Appendix
Table 10: Experiment configuration and first-stage learner hyperparameters.
Synthetic
ETF
Windows
2,595
4,261
Blocks F
140
228
Fits per block S
60
585
Ours ( F⋅S )
8,400
133,380
(A) fresh split per window
155,700 (18.5 × )
2,492,685 (18.7 × )
(B) same estimator per window
1,704,960 (203 × )
62,198,370 (466 × )
Appendix
Table 11: First-stage fits per run (one learner, one observation level). Ratios are relative to our implementation.
Learner
Ours
Per window
(A)
(B)
Extra Trees
6.1 min
0.14 s
1.9 h
20.8 h
Random Forest
9.4 min
0.22 s
2.4 h
31.8 h
LightGBM
13.6 min
0.31 s
3.7 h
45.9 h
XGBoost
17.1 min
0.40 s
6.1 h
57.9 h
Appendix
Table 12: Wall-clock time of one synthetic run (Apple M5 laptop, CPU).
Synthetic
ETF ( all10 )
Method
Window
Run
Window
Run
PCMCI (ParCorr)
0.16 s
7 min
4.4–9.1 s
5–11 h
PCMCI (CMIknn)
6.2–6.6 min
268–284 h
> 1 h
> 4,000 h
VAR-LiNGAM
0.02 s
1 min
9–15 s
11–17 h
Ours, Extra Trees
0.14 s
6.1 min
–
≈ 1.6 h ∗
Ours, TabPFN (A100)
–
48–57 s
–
88 min
Appendix
Table 13: Time of the causal-discovery baselines on the Apple M5 laptop CPU (TabPFN: one A100 GPU). Baselines: one window at a time in one process, scaled to a full run (2,595 synthetic and 4,261 ETF windows); windows are independent, so these runs split evenly over cores. Ours uses 9 threads. CMIknn is PCMCI with the nonparametric conditional-mutual-information test.
We propose CEDAR (Causal Edge Discovery for Autoregressive Processes), a constraint-based method for lagged causal edge discovery in sparse autoregressive time series. CEDAR screens candidate cross-variable lags using AR(1)-residualized, U-centered distance correlation, then applies two targeted conditional-independence tests per significant cross-variable lag candidate and accepts at most one lag per ordered pair. A stable MCI pruning step removes indirect edges, and optional deterministic C-nodes adjust for specified trend-like nonstationarity. In sparse regimes where few lags survive screening, CEDAR requires O(d2) CI tests after screening while retaining edge-level interpretability. CEDAR is most effective when data are scarce and variables exhibit lag-1 self-dynamics; methods with richer conditioning sets become preferable as T grows or when higher-order autoregressive or simultaneous multi-lag effects are common.
Causal discovery aims to recover causal relationships from observed data. In various fields, exploring causal relationships among variables remains an important topic, but this task becomes challenging due to the existence of latent confounders. Ignoring such confounders can lead to false associations and incorrect edge directions. In this paper, we study the linear structural equation model with latent confounders. We propose an algorithm that iteratively identifies terminal (observed) nodes and reconstructs the directed acyclic graph of the observed variables. To do this, we recover the precision matrix of the observed variables as a sparse plus low-rank matrix: a sparse matrix captures the conditional dependencies among observed variables, while a low-rank matrix captures the combined influence of a few latent confounders. We establish that for p observed variables, r latent confounders and s edges, our procedure correctly identifies the directed causal relationship among observed variables, for n≳max{slogp,rp} samples. Experimental results validate our theoretical contributions.
Granger causality characterizes directed predictive dependencies in multivariate time series, but recovering such dependencies becomes challenging in the presence of latent confounding. Cross-environment invariance provides a natural source of information in heterogeneous settings, yet invariance alone can be insufficient: when latent-to-observed mechanisms remain stable, hidden confounders can induce predictive dependencies that are just as invariant as genuine Granger-causal relations. We show that interventions provide an additional source of identifying information by inducing structured variation in observed mechanisms, while stable latent pathways need not exhibit the same cross-environment changes. In practice, however, neither the intervened environments nor the affected mechanisms are known. We propose GRACE, a framework for learning Granger causality under latent confounding from intervention-induced heterogeneity. GRACE decomposes multivariate dynamics into a shared Granger mechanism, sparse environment-specific deviations that capture edge-level interventions, and a latent component that accounts for confounding. Under a linear generative model, we show that GRACE can recover which environments intervene on a given edge when the edge is perturbed in at least one but fewer than half of the environments and the intervention effect is sufficiently large to survive sparsity shrinkage; the recovered intervention pattern then provides a certificate for the corresponding Granger causal edge. Experiments on synthetic and real-world time series demonstrate improved Granger causal structure recovery under latent confounding and unknown interventions.
Ziyi Zhang, Shaogang Ren, Xiaoning Qian +1
Texas A&M University · University of Tennessee at Chattanooga · Brookhaven National Laboratory