Partial observation and delay-coordinate reconstruction are rooted in deterministic dynamical-systems theory, whereas there are many systems with intrinsic stochasticity. We propose a conditional-moment interpretation of delay-coordinate reconstruction for stochastic systems, in which a delay vector is used to reconstruct conditional moments of a future distribution rather than a unique future sample path. Two complementary arguments motivate this viewpoint. First, the probability density of a stochastic differential equation obeys a deterministic Fokker-Planck equation and, under certain assumptions, is represented by an infinite deterministic hierarchy of moments. Hence, a finite-moment closure suggests a Takens-like finite-dimensional approximation. Second, a discussion based on the Koopman operator theory clarifies that the time evolution of an observable in the Mori-Zwanzig formalism yields a conditional expectation in stochastic systems. Then, the orthogonal "noise" term in the coefficient-space Mori-Zwanzig equation vanishes in the stochastic cases; this result is consistent with the moment-based argument. As an application of this stochastic delay-reconstruction viewpoint, we revisit convergent cross mapping (CCM) for diagnosing certain causal relationships. Although CCM based on the embedding theorem cannot generally be applied to stochastic systems, it is possible to examine certain types of causal relationships by using conditional moments. Using coupled logistic systems with additive and multiplicative coupling mechanisms, we discuss how causal relationships are embedded in stochastic systems.
Figures & tables
Figure 1: An illustrative example of deterministic delay-coordinate reconstruction. A trajectory in the original three-dimensional state space is only partially observed through the scalar time series xt . A delay vector, here illustrated by (xt,xt−τ,xt−2τ) with τ=0.2 , can provide a geometrically equivalent coordinate representation under certain conditions.
Figure 2: Deterministic and stochastic interpretations of delay-coordinate prediction. In a deterministic system, a sufficiently informative delay vector can determine a future value as xt+τ=φ(xt,xt−τ,…) . In a stochastic system, the corresponding deterministic object is a conditional statistic, for example E[Xt+τm∣Xt,Xt−τ,…]=φ(m)(Xt,Xt−τ,…) . For example, the conditional mean describes location, while the variance, obtained from the first and second moments, quantifies stochastic spread.
Figure 3: Sample trajectories of the two stochastic coupled-logistic models used to validate the conditional-moment interpretation. (a) Additive coupling in ( 57 ), for which Xt enters the conditional mean of Yt+1 . (b) Multiplicative coupling in ( 58 ), for which Xt modulates the amplitude of the noise term in Yt+1 . The figure contains 200 samples from trajectories of total retained length 500,000 after 500 relaxation steps.
Figure 4: Numerical results of one-step predictions in (a) additive and (b) multiplicative coupling cases. Upper and lower panels show predictions for X and Y , respectively. Thick pale lines are test data. The three superposed prediction markers compare the three calculations: the Monte Carlo sampling initialized from the full current state (Xt,Yt) , polynomial EDMD based on a single-variable six-dimensional history, and the weighted k -nearest-neighbor estimate based on the same history. Dotted vertical segments show one predicted standard deviation. The small horizontal offsets are only for visualization and do not represent temporal lags.
Figure 5: Definition of the time-shifted conditional reconstruction. The predictor history is HtA=(At,At−1,…) . For lag ℓ , the target is Bt−ℓ . Hence ℓ>0 reconstructs the target’s past, ℓ=0 the simultaneous target, and ℓ<0 the target’s future. The regression approximates E[Bt−ℓ∣HtA] , and the correlation between the conditional-expectation estimate and the realized shifted target defines the reported reconstruction skill.
Figure 6: Library-size dependence of the CCM-like first-moment reconstruction for the additive X→Y model. The tested X→Y direction uses HtY↦Xt−1 ; the reverse uses HtX↦Yt−1 . (a) rx=3.1 , ry=2.8 : the correct-direction skill grows more strongly, but the reverse skill also increases. (b) rx=2.8 , ry=3.1 : both directions increase and, over much of the range, the reverse reconstruction is larger despite the unchanged structural direction X→Y . Hence, convergence with L by itself is not sufficient to identify direction in this stochastic example. The graph plots the mean and standard deviation of the results obtained from generating and analyzing 10 individual time-series datasets.
Figure 7: Sample trajectories for the common-driver systems with structural graph Z→X and Z→Y . (a) Additive driver coupling in ( 86 ). (b) Multiplicative driver coupling in ( 87 ). Although the pair (X,Y) has no direct edge in either model, the two series share information through Z , providing a test of whether a pairwise lag diagnostic can distinguish common-driver structure from direct driver–response structure.
Figure 8: Lag dependence of first-moment conditional reconstruction in the additive common-driver system. ℓ>0 means the reconstruction of the target’s past from the predictor history. (a) Common-driver pair X – Y . The two directions have similar broad profiles because both variables are driven by Z and there is no direct edge between them. (b) Driver–response pair X – Z . The structural direction Z→X is evaluated by HtX↦Zt−ℓ and is markedly stronger, with a maximum near ℓ=5 for the six-coordinate history used here. The peak location is embedding dependent and is not interpreted as a five-step physical coupling delay. The graph plots the mean and standard deviation of the results obtained from generating and analyzing 10 individual time-series datasets.
Figure 9: Lag dependence of first-moment conditional reconstruction in the multiplicative common-driver system. (a) Common-driver pair X – Y : both curves remain near zero because the common driver enters the two response equations through independent zero-mean multiplicative innovations, so the leading one-step information is in conditional variances, not means. (b) Driver–response pair X – Z : the structural Z→X direction is visible indirectly in the first moment and peaks near ℓ=3 or ℓ=4 . The signal arises after nonlinear propagation converts the Z -dependent variance of Xt+1 into a Z2 -dependent mean at later times; ( 90 ) gives the simplest analytic illustration. The graph plots the mean and standard deviation of the results obtained from generating and analyzing 10 individual time-series datasets.
We consider sparse multivariate stochastic systems that evolve in continuous time according to a causal mechanism and present methodology to recover the system's time-infinitesimal transition mechanism from mere cross-sectional data. This observational paradigm is motivated by applications such as gene expression analysis, where destructive experimental techniques may only allow recording data once over a cell's lifetime. Precisely, we assume the system follows a time-homogeneous diffusion process that has reached an equilibrium distribution at observation time. Further, we assume the causal mechanism is fully described by the diffusion drift, is acyclic, and its causal structure graph is known. In this setting, we prove that the full causal mechanism, i.e., the drift function, can be non-parametrically identified under a weak non-explosion criterion. We derive a non-parametric kernel estimator for this challenging inverse problem and prove its consistency. Moreover, we propose a cross-validation scheme for hyperparameter tuning, illustrate the behavior of our estimator in simulations, and we discuss connections with irreversible generative diffusion models and low-frequency sampled data.
Richard Schwank, Mathias Drton
School of Computation, Information and Technology, Technical University of Munich · 2Munich Center for Machine Learning
Pearl's structural causal model (SCM) framework, built on directed acyclic graphs (DAGs) and the do-calculus, is the dominant formal language for causal reasoning. Yet it carries two structural restrictions: every relationship must be pre-specified as a directed causal edge, and feedback cycles are forbidden. This paper examines two classes of phenomena that strain these restrictions. First, symmetric physical and economic constraints, the ideal gas law being the canonical case, carry no intrinsic causal direction. Direction emerges only under intervention, and which variable is solved for must be specified as part of the intervention. We formalize such constraints as causal zeros within an Extended Causal Model by adding an activation operator, subject to local solvability and graph-admissibility conditions. Second, for the class of finite-propagation state-space systems considered here, we treat apparent instantaneous cycles as artifacts of suppressed time and ground both causal zeros and feedback in Causal Differential Equations (CDEs). In these, the transient regime is a time-unrolled acyclic causal process, and causal zeros arise as the defining functions of attracting equilibrium manifolds; periodic and chaotic attractors define further regimes of the same dynamics, treated through attractor-relative intervention. We give the extended do-calculus, identifiability conditions, counterfactual semantics, and open problems.
Sergei V. Kalinin
Department of Materials Science and Engineering, University of Tennessee, Knoxville
We study identifiability in continuous-time linear stationary stochastic differential equations with a known causal structure. Unlike existing approaches, we relax the assumption of a known diffusion matrix, thereby respecting the model's intrinsic scale invariance. Therefore, rather than recovering drift coefficients themselves, we introduce edge-sign identifiability: for a given causal structure, we ask whether the sign of a given drift entry is uniquely determined across all observational covariance matrices induced by parametrisations compatible with that structure. This leads to a trichotomy of edge-sign identifiability: identifiable, non-identifiable, and partially identifiable. This trichotomy introduces the new notion of partial identifiability to the literature, which we show is a genuine category in our setting. Under a notion of faithfulness, we derive criteria to identify membership of each category for general graphs. Applying our criteria to specific causal structures, both analogous to classical causal settings (e.g., instrumental variables) and novel cyclic settings, we determine their edge-sign identifiability and, in some cases, obtain explicit expressions for the sign of a target edge in terms of the observational covariance matrix.