Identifiability Guarantees for Drivers and Dynamics of Delayed Physical Systems
Authors: Julien Boussard, Antoine Debouchage, Théo Saulus
Organizations: School of Computer Science, McGill University, Canada · Mila - Quebec AI Institute, Canada · LaMMe, Université Évry Paris-Saclay, France · DIRO, Université de Montréal, Canada
A wide range of methods have been proposed, including physics-informed neural networks, which are powerful but do not guarantee identifiability of the dynamics, symbolic regression, which requires a set of precomputed operations, and causal discovery, which is more principled but usually relies on strong assumptions that physical systems may violate. In this work, we develop a theory-grounded method and prove that under a set of permissive assumptions, the structural drivers and drift of stochastic delayed differential equations are identifiable. Our method outperforms others on a benchmark for driver identifiability, and on a second benchmark to evaluate physical consistency of the learned dynamics.
Figures & tables
Experiment
Simple
Coupled
Climate
Avg. rank
Default
Confounder
Noise
Default
Noise
Confounder
Time-lag
Standardize
MAOOAM
ENSO
SHD
PCMCI+
41.04
23.02
47.64
224.80
183.90
324.63
327.72
228.32
80.00
529.36
7.30
F-PCMCI
35.30
21.07
45.09
192.90
149.60
195.74
350.61
201.79
130.00
530.27
5.80
NGC
28.91
19.96
28.84
840.95
842.55
670.53
793.67
840.26
31.00
337.09
7.90
CUTS+
48.11
24.04
61.32
152.00
150.50
272.68
247.22
310.63
130.00
608.73
7.50
RCD
61.85
26.74
61.38
157.05
155.65
136.53
201.11
159.84
130.00
665.36
6.80
Table 1: L0 -NDDE outperforms competing methods on the CausalDynamics benchmark. SHD (lower is better) and AUROC / AUPRC (higher is better) are reported for L0 -NDDE (Ours), AGL, C-NODE and the benchmark causal discovery methods. The best values are bold, the second best values are underlined. ∗ indicates statistical significance between AGL or C-NODE and L0 -NDDE, calculated with a Welch’s t-test (p-value 0.05). The benchmark only provides the mean values for the causal discovery methods, and we thus do not run statistical significance tests for these methods. For AUROC and AUPRC, statistical significance results always match, so we report only one ∗ for each pair. The last column reports the average rank for each method and metric. For readability, methods that never score best are not shown here, but in the Appendix ( Table 4 )
Method
NRMSE ( ↓ )
VLPT ( ↑ )
FLEE ( ↓ )
DKY ( ↓ )
LSD ( ↓ )
Dstsp ( ↓ )
SHD ( ↓ )
Med.
Rank
Med.
Rank
Med.
Rank
Med.
Rank
Med.
Rank
Med.
Rank
Med.
Rank
AGL
1.15
2.69
2.85
3.14
0.25
3.47
0.88 ∗
3.82
14.71 ∗
2.82
3e 6 ∗
4.37
0.44 ∗
4.46
PathReg
1.19
2.49
4.33
3.24
0.25
3.78
0.89 ∗
3.9
16.99
3.02
3e 6 ∗
4.62
0.22
2.70
C-NODE
1.53
4.05
4.61
3.39
0.20
3.00
0.72 ∗
3.17
37.79 ∗
4.08
34.60 ∗
2.86
0.17
2.67
NODE
1.20
2.94
5.12
2.63
0.16
2.51
0.076
2.17
13.99
2.57
4.29
1.62
0.22
2.70
L0 -NDDE
1.19
2.83
6.54
2.60
0.16
2.25
0.075
1.94
14.75
2.50
2.50
1.54
0.11
2.45
Table 2: L0 -NDDE outperforms other feature selection methods. L0 -NDDE ranks best in five of the metrics. The difference with NODE is non significant, as NODE also performs very well, but L0 -NDDE consistently ranks better than the other methods performing input feature selection, especially on attractor dimension reconstruction ( DKY ) and convergence in distribution ( Dstsp ). ∗ indicates statistical significance with the L0 -NDDE metric, obtained using a permutation test (p-value 0.05). For each metric, best results are bolded and second best are underlined.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1: Dependency map of the theoretical results. Yellow boxes are assumptions, blue boxes are lemmas and propositions, and green boxes are theorems. Arrows denote when an assumption or a result is used in another proof. A use already implied by a chain of arrows is not drawn, unless it plays a separate role (labelled arrows). The proofs are grouped vertically into the three groups of proofs in Appendix A , contributing the finite-sample guarantee at the bottom.
Parameter
Value
Architecture, functions fj
Number of layers
2
Number of units per layer
8
Optimization parameters
Gradient penalty coefficient
100
Warm-up iterations
100
Appendix
Table 3: Final hyperparameters used for L0 -NDDE, C-NODE, AGL and No grad. pen. The last 4 lines show the penalty coefficient for each method, as it is the only parameter tuned for each the method. Note that No grad. pen. does not use gradient penalty.
Experiment
Simple
Coupled
Climate
Avg. rank
Default
Confounder
Noise
Default
Noise
Confounder
Time-lag
Standardize
MAOOAM
ENSO
SHD
PCMCI+
41.04
23.02
47.64
224.80
183.90
324.63
327.72
228.32
80.00
529.36
7.30
F-PCMCI
35.30
21.07
45.09
192.90
149.60
195.74
350.61
201.79
130.00
530.27
5.80
VARLiNGAM
35.69
22.04
42.84
311.45
248.90
159.63
449.33
349.63
130.00
453.00
8.00
DYNOTEARS
52.37
21.74
51.68
181.50
180.95
248.32
261.44
243.84
94.00
589.36
6.60
NGC
28.91
19.96
28.84
840.95
842.55
670.53
793.67
840.26
31.00
337.09
7.90
Appendix
Table 4: L0 -NDDE outperforms competing methods most of the time on the CausalDynamics benchmark. This table is similar to Table 1 with additional methods being shown.
Experiment
Simple
Coupled
Climate
Default
Confounder
Noise
Default
Noise
Confounder
Time-lag
Standardize
MAOOAM
ENSO
SHD
No Grad. Pen.
26.95
18.21
24.47
215.08
142.67
145.08
233.58
220.68
49.00
341.27
L0 -NDDE (Ours)
26.42
16.25
23.67
205.35
144.20
143.33
222.34
212.11
40.00
325.54
AUROC /
No Grad. Pen.
.64 / .80
.58 / .65
.61 / .79
.63 / .42
.64 / .44
.67 / .47
.59 / .41
.65 / .44
.85 / .96
.52 / .68
AUPRC
L0 -NDDE (Ours)
.64 / .80
.58 / .66
.63 / .80
.64 / .43
.65 / .44
.67 / .47
.60 / .43
.66 / .45
.86 / .96
.58 / .75
Appendix
Table 5: L0 -NDDE is slightly stronger with gradient penalty. Gradient penalty helps avoid overfitting spurious correlations and leads to slightly better performance on the CausalDynamics benchmark.
Optimization parameters
L0 -NDDE
NODE
C-NODE
AGL
PathReg
Gradient penalty coefficient
0
100
0
1000
1000
niter_warmup
0
0
100
50
200
n_iter_after_convergence
200
0
500
1000
500
Sparsity penalty coefficients
0.05
0
0.05
50
1
Appendix
Table 6: Final hyperparameters used for L0 -NDDE, C-NODE, AGL and PathReg
Figure 2: Example predicted trajectories (red) vs. true trajectories (black) on 26 Dysts systems, for L0 -NDDE (left), NODE (middle), and C-NODE (right).
Figure 3: Lower training error and lower SHD are both associated with higher performance. We perform an analysis to better understand why L0 -NDDE and NODE outperform other methods on the Dysts datasets. We compute, for all methods and all datasets, the training error (Train MSE) and SHD. We then plot the 6 metrics versus the SHD (upper half) and versus the training error (lower half). We see that DKY and Dstsp increase, while VLPT decreases with SHD and training error. These relationships are statistically significant according to a Kendall- τ test, a standard rank-correlation test for investigating associations between two sets of values, without assuming Gaussianity of the samples or needing large samples. The relationship is also significant between LSD and training error. This indicates that a lower training error is associated with better reconstruction, indicating why NODE performs well, as the absence of input feature regularization allows it to overfit to the training data; and why L0 -NDDE outperforms other methods, as it achieves lower SHD. Moreover, the L0 -penalty has the advantage that it does not perform weight decay, but explicitly turns features on or off, allowing the rest of the network to approximate the dynamics very well given the set of selected features.
Learning governing dynamics from data is a common goal across the sciences, yet it is only well-posed when the underlying mechanisms are identifiable. In practice, many data-driven methods implicitly assume identifiability; when this assumption fails, estimated models can yield spurious predictions and invalid mechanistic conclusions. Classical identifiability guarantees for controlled linear time-invariant (LTI) systems provide sufficient conditions -- controllability and persistent excitation -- but leave open whether identifiability holds when these conditions fail, and which parts of the system remain identifiable without full identifiability. We show that the experimental setup, i.e., the realized initial state and control input, dictates a fundamental limit on the information recoverable from the observed trajectory. We develop a geometric characterization of this limit and derive a closed-form description of all systems consistent with the experimental setup. Crucially, we prove that even when the full system is not identifiable, the restricted dynamics on the subspace reachable by the experiment remain uniquely determined.
Aybüke Ulusarslan, Niki Kilbertus, Nora Schneider
1Technical University of Munich · 2Helmholtz Munich · 3Munich Center for Machine Learning (MCML)
Causal representation learning for time series has developed strong identifiability results in discrete-time latent causal models, but identifiability in continuous-time latent stochastic differential equation (SDE) models remains largely open. We address this gap using environment-induced shifts in diffusion covariance. We study additive-noise latent SDEs observed through an unknown nonlinear diffeomorphism, with shared drift but environment-specific diffusion covariance. We show that two diagonal diffusion regimes with pairwise distinct coordinate-wise variance ratios identify the latent coordinates up to permutation, coordinate-wise scaling, and a possible constant shift, without any sparsity assumption on the drift. We first prove this result for linear Ornstein-Uhlenbeck systems and then extend it to general additive-noise latent SDEs. Under mild smoothness, the instantaneous drift-Jacobian causal graph is identifiable up to the same permutation. We propose a two-stage estimator for latent disentanglement and optional graph recovery; experiments on synthetic systems confirm the predicted identifiability boundary, and an application to Hardanger Bridge monitoring data illustrates the approach on real sensor trajectories.
Yuanyuan Wang, Wenjie Wang, Haoxuan Li +2
Mohamed bin Zayed University of Artificial Intelligence · The University of Melbourne · Peking University +1
We study identifiability in continuous-time linear stationary stochastic differential equations with a known causal structure. Unlike existing approaches, we relax the assumption of a known diffusion matrix, thereby respecting the model's intrinsic scale invariance. Therefore, rather than recovering drift coefficients themselves, we introduce edge-sign identifiability: for a given causal structure, we ask whether the sign of a given drift entry is uniquely determined across all observational covariance matrices induced by parametrisations compatible with that structure. This leads to a trichotomy of edge-sign identifiability: identifiable, non-identifiable, and partially identifiable. This trichotomy introduces the new notion of partial identifiability to the literature, which we show is a genuine category in our setting. Under a notion of faithfulness, we derive criteria to identify membership of each category for general graphs. Applying our criteria to specific causal structures, both analogous to classical causal settings (e.g., instrumental variables) and novel cyclic settings, we determine their edge-sign identifiability and, in some cases, obtain explicit expressions for the sign of a target edge in terms of the observational covariance matrix.