Identifiability Guarantees for Drivers and Dynamics of Delayed Physical Systems
Authors: Julien Boussard, Antoine Debouchage, Théo Saulus
Organizations: School of Computer Science, McGill University, Canada · Mila - Quebec AI Institute, Canada · LaMMe, Université Évry Paris-Saclay, France · DIRO, Université de Montréal, Canada
A wide range of methods have been proposed, including physics-informed neural networks, which are powerful but do not guarantee identifiability of the dynamics, symbolic regression, which requires a set of precomputed operations, and causal discovery, which is more principled but usually relies on strong assumptions that physical systems may violate. In this work, we develop a theory-grounded method and prove that under a set of permissive assumptions, the structural drivers and drift of stochastic delayed differential equations are identifiable. Our method outperforms others on a benchmark for driver identifiability, and on a second benchmark to evaluate physical consistency of the learned dynamics.
Figures & tables
Experiment
Simple
Coupled
Climate
Avg. rank
Default
Confounder
Noise
Default
Noise
Confounder
Time-lag
Standardize
MAOOAM
ENSO
SHD
PCMCI+
41.04
23.02
47.64
224.80
183.90
324.63
327.72
228.32
80.00
529.36
7.30
F-PCMCI
35.30
21.07
45.09
192.90
149.60
195.74
350.61
201.79
130.00
530.27
5.80
NGC
28.91
19.96
28.84
840.95
842.55
670.53
793.67
840.26
31.00
337.09
7.90
CUTS+
48.11
24.04
61.32
152.00
150.50
272.68
247.22
310.63
130.00
608.73
7.50
RCD
61.85
26.74
61.38
157.05
155.65
136.53
201.11
159.84
130.00
665.36
6.80
Table 1: L0 -NDDE outperforms competing methods on the CausalDynamics benchmark. SHD (lower is better) and AUROC / AUPRC (higher is better) are reported for L0 -NDDE (Ours), AGL, C-NODE and the benchmark causal discovery methods. The best values are bold, the second best values are underlined. ∗ indicates statistical significance between AGL or C-NODE and L0 -NDDE, calculated with a Welch’s t-test (p-value 0.05). The benchmark only provides the mean values for the causal discovery methods, and we thus do not run statistical significance tests for these methods. For AUROC and AUPRC, statistical significance results always match, so we report only one ∗ for each pair. The last column reports the average rank for each method and metric. For readability, methods that never score best are not shown here, but in the Appendix ( Table 4 )
Method
NRMSE ( ↓ )
VLPT ( ↑ )
FLEE ( ↓ )
DKY ( ↓ )
LSD ( ↓ )
Dstsp ( ↓ )
SHD ( ↓ )
Med.
Rank
Med.
Rank
Med.
Rank
Med.
Rank
Med.
Rank
Med.
Rank
Med.
Rank
AGL
1.15
2.69
2.85
3.14
0.25
3.47
0.88 ∗
3.82
14.71 ∗
2.82
3e 6 ∗
4.37
0.44 ∗
4.46
PathReg
1.19
2.49
4.33
3.24
0.25
3.78
0.89 ∗
3.9
16.99
3.02
3e 6 ∗
4.62
0.22
2.70
C-NODE
1.53
4.05
4.61
3.39
0.20
3.00
0.72 ∗
3.17
37.79 ∗
4.08
34.60 ∗
2.86
0.17
2.67
NODE
1.20
2.94
5.12
2.63
0.16
2.51
0.076
2.17
13.99
2.57
4.29
1.62
0.22
2.70
L0 -NDDE
1.19
2.83
6.54
2.60
0.16
2.25
0.075
1.94
14.75
2.50
2.50
1.54
0.11
2.45
Table 2: L0 -NDDE outperforms other feature selection methods. L0 -NDDE ranks best in five of the metrics. The difference with NODE is non significant, as NODE also performs very well, but L0 -NDDE consistently ranks better than the other methods performing input feature selection, especially on attractor dimension reconstruction ( DKY ) and convergence in distribution ( Dstsp ). ∗ indicates statistical significance with the L0 -NDDE metric, obtained using a permutation test (p-value 0.05). For each metric, best results are bolded and second best are underlined.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1: Dependency map of the theoretical results. Yellow boxes are assumptions, blue boxes are lemmas and propositions, and green boxes are theorems. Arrows denote when an assumption or a result is used in another proof. A use already implied by a chain of arrows is not drawn, unless it plays a separate role (labelled arrows). The proofs are grouped vertically into the three groups of proofs in Appendix A , contributing the finite-sample guarantee at the bottom.
Parameter
Value
Architecture, functions fj
Number of layers
2
Number of units per layer
8
Optimization parameters
Gradient penalty coefficient
100
Warm-up iterations
100
Appendix
Table 3: Final hyperparameters used for L0 -NDDE, C-NODE, AGL and No grad. pen. The last 4 lines show the penalty coefficient for each method, as it is the only parameter tuned for each the method. Note that No grad. pen. does not use gradient penalty.
Experiment
Simple
Coupled
Climate
Avg. rank
Default
Confounder
Noise
Default
Noise
Confounder
Time-lag
Standardize
MAOOAM
ENSO
SHD
PCMCI+
41.04
23.02
47.64
224.80
183.90
324.63
327.72
228.32
80.00
529.36
7.30
F-PCMCI
35.30
21.07
45.09
192.90
149.60
195.74
350.61
201.79
130.00
530.27
5.80
VARLiNGAM
35.69
22.04
42.84
311.45
248.90
159.63
449.33
349.63
130.00
453.00
8.00
DYNOTEARS
52.37
21.74
51.68
181.50
180.95
248.32
261.44
243.84
94.00
589.36
6.60
NGC
28.91
19.96
28.84
840.95
842.55
670.53
793.67
840.26
31.00
337.09
7.90
Appendix
Table 4: L0 -NDDE outperforms competing methods most of the time on the CausalDynamics benchmark. This table is similar to Table 1 with additional methods being shown.
Experiment
Simple
Coupled
Climate
Default
Confounder
Noise
Default
Noise
Confounder
Time-lag
Standardize
MAOOAM
ENSO
SHD
No Grad. Pen.
26.95
18.21
24.47
215.08
142.67
145.08
233.58
220.68
49.00
341.27
L0 -NDDE (Ours)
26.42
16.25
23.67
205.35
144.20
143.33
222.34
212.11
40.00
325.54
AUROC /
No Grad. Pen.
.64 / .80
.58 / .65
.61 / .79
.63 / .42
.64 / .44
.67 / .47
.59 / .41
.65 / .44
.85 / .96
.52 / .68
AUPRC
L0 -NDDE (Ours)
.64 / .80
.58 / .66
.63 / .80
.64 / .43
.65 / .44
.67 / .47
.60 / .43
.66 / .45
.86 / .96
.58 / .75
Appendix
Table 5: L0 -NDDE is slightly stronger with gradient penalty. Gradient penalty helps avoid overfitting spurious correlations and leads to slightly better performance on the CausalDynamics benchmark.
Optimization parameters
L0 -NDDE
NODE
C-NODE
AGL
PathReg
Gradient penalty coefficient
0
100
0
1000
1000
niter_warmup
0
0
100
50
200
n_iter_after_convergence
200
0
500
1000
500
Sparsity penalty coefficients
0.05
0
0.05
50
1
Appendix
Table 6: Final hyperparameters used for L0 -NDDE, C-NODE, AGL and PathReg
Figure 2: Example predicted trajectories (red) vs. true trajectories (black) on 26 Dysts systems, for L0 -NDDE (left), NODE (middle), and C-NODE (right).
Figure 3: Lower training error and lower SHD are both associated with higher performance. We perform an analysis to better understand why L0 -NDDE and NODE outperform other methods on the Dysts datasets. We compute, for all methods and all datasets, the training error (Train MSE) and SHD. We then plot the 6 metrics versus the SHD (upper half) and versus the training error (lower half). We see that DKY and Dstsp increase, while VLPT decreases with SHD and training error. These relationships are statistically significant according to a Kendall- τ test, a standard rank-correlation test for investigating associations between two sets of values, without assuming Gaussianity of the samples or needing large samples. The relationship is also significant between LSD and training error. This indicates that a lower training error is associated with better reconstruction, indicating why NODE performs well, as the absence of input feature regularization allows it to overfit to the training data; and why L0 -NDDE outperforms other methods, as it achieves lower SHD. Moreover, the L0 -penalty has the advantage that it does not perform weight decay, but explicitly turns features on or off, allowing the rest of the network to approximate the dynamics very well given the set of selected features.