Neural Stochastic Differential Equations (Neural SDEs) provide flexible continuous-time generative models, but generic neural drift and diffusion networks are costly to simulate on long horizons and can give unstable gradients when the training signal is a path functional rather than a pointwise observation. We introduce SLiSDE, a family of Neural SDE models built from structured linear stochastic layers. Parallel-in-time simulation is obtained at the layer level, while expressivity is recovered by gated in-flow stacking: previous-layer paths modulate the next layer's latent flow through learned gates. For functional calibration tasks in which rare paths dominate the loss, we add an optional Girsanov tilt that acts as a learned importance sampler with an exact likelihood-ratio correction. We prove well-posedness, a discretisation error bound, validity of the change of measure, and a universality result: the terminal laws of the gated stack are dense in the space of square-integrable laws. Experiments on functional calibration benchmarks show that the structured model outperforms fully neural SDE baselines while retaining parallel-time simulation and stable importance weights.
Figures & tables
Diagonal
Block-diagonal
Dense
NeuralSDE
Recurrent cost
O(dT)
O(nbbs2T)
O(d2T)
O(d2T)
Scan cost
O(dlogT)
O(nbbs3logT)
O(d3logT)
—
Table 1: Computational cost of composing affine transition maps over T time steps. Here d is the latent dimension, bs is the block size, and nb=d/bs is the number of blocks. Typically nb≪d .
Model
Params
Loss ( 10−4 )
Barrier MAE ( 10−2 )
ms / epoch
Neural SDE – L=2
124 K
1.557±0.206
1.058±0.110
91.7
SLiCE – L=2
142 K
4.690±0.423
2.168±0.102
171.8
SLiSDE gated in-flow – L=2
103 K
0.852±0.168
0.658±0.054
72.9
Neural SDE – L=3
157 K
3.024±0.341
1.637±0.120
110.4
SLiCE – L=3
213 K
5.006±0.914
2.250±0.252
248.2
SLiSDE gated in-flow – L=3
74 K
0.659±0.059
0.672±0.057
68.9
Table 2 : toy results without Girsanov tilt. L is the number of stacked layers; each row is the tuned cell of the search grid for that family and depth, retrained under the common protocol of Appendix D . We report parameter count, calibration loss, threshold-crossing error, and time per epoch.
Model
Params
Loss ( 10−4 )
Far put ( 10−6 )
ms / epoch
Neural SDE – L=2
79 K
1.972±0.117
9.87±2.47
91.4
SLiCE – L=2
142 K
1.605±0.149
8.14±1.69
169.9
SLiSDE gated in-flow – L=2
90 K
1.552±0.070
9.40±2.30
79.7
Neural SDE – L=3
157 K
1.908±0.180
8.20±1.42
114.4
SLiCE – L=3
213 K
1.560±0.086
6.41±1.33
246.9
SLiSDE gated in-flow – L=3
135 K
1.401±0.244
4.07±0.62
129.1
Table 3 : dax results without Girsanov tilt. L is the number of stacked layers; each row is the tuned cell of the search grid for that family and depth, retrained under the common protocol of Appendix D . We report parameter count, held-out loss, held-out far-put (OTM) loss, and time per epoch.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Params
Loss ( 10−4 )
Far put ( 10−6 )
ms / epoch
Neural SDE – L=2
79 K
11.90±0.11
135.4±2.9
94.1
SLiCE – L=2
142 K
6.43±0.12
81.4±4.6
170.1
SLiSDE gated in-flow – L=2
18 K
6.44±0.40
83.7±4.8
44.3
Neural SDE – L=3
157 K
10.29±0.13
136.0±2.8
109.7
SLiCE – L=3
213 K
6.53±0.26
82.9±3.6
248.4
SLiSDE gated in-flow – L=3
34 K
6.29±0.24
70.1±6.6
70.5
Appendix
Table 4 : spx option-surface results without Girsanov tilt, two and three layers. Same masks and protocol as Table 3 ; best cell of the search grid per family and depth, seven seeds, mean ± standard error. The narrower three-layer SLiCE cell diverged on every seed and is excluded.
Variant
Params
Loss
Far put
ms / epoch
Noise type (block size 8 )
multiplicative
0.99×
1.05×
0.97×
0.93×
additive
0.46×
0.95×
1.69×
0.78×
general (both; reference)
1.00×
1.00×
1.00×
1.00×
Block size (general noise)
bs=1 (diagonal)
0.46×
0.95×
1.45×
0.60×
Appendix
Table 5 : Structural ablations on dax , each varying one factor of the two-layer gated in-flow configuration (general noise, block size 8 ). Every entry is the ratio of the seven-seed mean of the variant to that of the reference configuration; the total-loss differences within each block are within one standard error. Top: noise type of the structured layers. Bottom: block size of the block-diagonal structure.
B
T
SLiSDE (seq.)
Neural SDE
SLiCE
speed-up
64
512
9.5
54.0
10.1
5.7 ×
64
2048
35.8
212.0
42.8
5.9 ×
64
8192
117.1
838.9
156.9
7.2 ×
256
2048
79.7
257.1
155.2
3.2 ×
1024
512
69.0
89.4
155.0
1.3 ×
1024
2048
258.9
359.5
OOM
1.4 ×
Appendix
Table 6: Wall-clock time per training step (ms; forward plus backward, post-compilation, median of five repeats) as a function of batch size B and horizon T for SLiSDE ( 136 K parameters, sequential evaluation), the Neural SDE of the same parameter count ( 137 K) and SLiCE ( 142 K). Speed-up is Neural SDE over SLiSDE; OOM: exceeds the 42.6 GiB of the GPU.
B
T
SLiSDE fwd (seq.)
SLiSDE fwd (scan)
NSDE fwd
SLiSDE fwd+bwd
NSDE fwd+bwd
64
512
5.7
4.1
15.3
9.5
54.0
64
2048
14.9
15.8
50.6
35.8
212.0
64
8192
48.4
50.4
202.6
117.1
838.9
256
2048
32.5
44.6
61.0
79.7
257.1
1024
512
27.6
39.9
23.8
69.0
89.4
1024
2048
89.9
OOM
84.2
258.9
359.5
Appendix
Table 7 : Wall-clock time (ms, post-compilation, NVIDIA RTX 6000 Ada) of forward passes and full training steps as a function of batch size and horizon, including the chunked associative scan (best of chunk sizes 16 , 64 , 256 ); the Neural SDE has the same parameter count as SLiSDE ( 137 K against 136 K).
KS (terminal)
W1 (terminal)
KS (running max)
W1 (running max)
Neural SDE ( L=2 )
0.019±0.002
0.020±0.002
0.059±0.002
0.047±0.002
SLiCE ( L=2 )
0.022±0.002
0.031±0.003
0.081±0.003
0.067±0.003
SLiSDE ( L=2 )
0.015±0.001
0.023±0.002
0.036±0.002
0.037±0.002
Appendix
Table 8 : Distances between the learned and the ground-truth laws of the terminal value and of the running maximum on toy ( 2×104 model paths; best two-layer cell of each family; seven seeds, mean ± s.e.).
Arm
paths / step
Loss ( 10−4 )
Far put ( 10−6 )
Far call ( 10−6 )
\ESS/nq put / call
\KL put / call
Vanilla
1024
1.15±0.12
0.93±0.46
3.15±0.33
–
–
Untilted, ρ=0
1024
1.18±0.19
0.87±0.42
3.62±0.78
–
–
Untilted, ρ=0
2048
0.95±0.13
0.81±0.30
2.59±0.49
–
–
Tilt, combined far-tail estimate
1024+2×512
0.98±0.10
0.82±0.27
2.03±0.78
0.05 / 0.09
2.9 / 1.5
Tilt, tilted batch only
1024+2×512
0.88±0.10
0.48±0.08
3.25±1.04
0.05 / 0.12
3.0 / 1.4
Appendix
Table 9 : dax with the Girsanov tilt (seven seeds, mean ± standard error). The tilt study is trained and evaluated on the strikes that lie within the quoted range at each expiry, so its absolute values are not comparable with Table 3 , whose protocol it otherwise follows. Far-tail losses are held-out MSEs on the far strikes; every arm is evaluated by plain Monte Carlo under its reference law with 32768 paths. The tilted arms use ρ=0 , the cross-entropy controller with a 6 -nat KL wall and the initial-state shift; \ESS/nq and \KL(Q∥P) (nats) are end-of-training values of the put / call controllers; time per epoch 130 / 130 / 229 / 303 / 295 ms.
Estimator ( n=1024 ), rel. RMSE
p≈10−2
p≈10−3
p≈10−4
Plain MC
0.320±0.005
1.051±0.019
3.39±0.39
Plain MC, cost-matched ( ×1.18 )
0.278±0.009
0.934±0.047
3.15±0.37
Girsanov tilt, IS
0.156±0.018
0.394±0.104
0.785±0.184
Girsanov tilt, SNIS
0.274±0.021
0.789±0.106
1.88±0.42
Share of plain-MC runs returning exactly 0
0%
34%
90%
Appendix
Table 10: Rare-event estimation on toy : relative RMSE of each estimator at a budget of n=1024 paths per estimate (cost-matched plain Monte Carlo receives 1.18n paths), over 100 replications for each of seven independently calibrated models; mean ± standard error over the models.
rel. RMSE
n=256
n=1024
n=4096
p≈10−2
Plain MC
0.645±0.017
0.320±0.005
0.163±0.005
Plain MC, cost-matched
0.560±0.012
0.278±0.009
0.151±0.003
Tilt, IS
0.411±0.084
0.156±0.018
0.140±0.033
Tilt, IS truncated
0.391±0.070
0.156±0.018
0.119±0.017
Tilt, SNIS
0.645±0.131
0.274±0.021
0.197±0.034
plain MC returning 0
8%
0%
0%
Appendix
Table 11: Rare-event estimation on toy : relative RMSE for three budgets and three nominal probabilities, over 100 replications for each of seven independently calibrated models (mean ± standard error over the models); cost-matched plain Monte Carlo receives 1.18n paths; the last row of each block is the share of plain-Monte-Carlo replications with no path above the threshold.
Controller
Loss ( 10−4 )
Far put ( 10−6 )
Far call ( 10−6 )
\ESS/nq put / call
CE, κ=6 , combined (7 seeds)
0.98±0.10
0.82±0.27
2.03±0.78
0.05 / 0.09
CE, κ=6 , tilted batch only (7 seeds)
0.88±0.10
0.48±0.08
3.25±1.04
0.05 / 0.12
CE, κ=2 , combined (3 seeds)
1.13±0.20
0.59±0.01
2.46±1.03
0.06 / 0.08
SNIS-variance, combined (3 seeds)
1.18±0.35
1.18±0.33
2.71±1.10
1.00 / 0.99
Appendix
Table 12: Controller ablations on dax far-tail calibration ( ρ=0 , protocol of Table 9 ; mean ± standard error over seeds).
Neural stochastic differential equations (SDEs) have emerged as powerful tools for learning noisy or stochastic dynamics directly from data; however, existing approaches largely assume uncoupled and continuous noise, limiting their applicability to realistic stochastic drivers, and often scale poorly in time, requiring expensive autoregressive training. To address these limitations, we propose Neural Kolmogorov Equations (NKEs), a deterministic, infinite-dimensional reformulation of Neural SDEs based on the Kolmogorov Forward equation, transforming the learning problem from modelling individual stochastic trajectories to modelling the evolution of probability densities. NKEs learn general Lévy-type stochastic forcing directly through the operator structure of the KFE, and enable parallel-in-time training via a Lagrangian Galerkin projection and operator splitting. We evaluate NKEs on several stochastic benchmarks, including systems with coupled noise and jump processes, and verify that NKEs provide flexible models that accurately recover deterministic and stochastic dynamics with competitive predictive accuracy and improved training efficiency. Code and pretrained models will be released.
Recovering dynamical systems from noisy observations is a recurring challenge across scientific domains, including neuroscience and physics. Latent stochastic differential equations (SDEs) address this by modeling the system as an unobserved state that evolves according to a learnable SDE and generates the observations. Variational inference (VI) provides a tractable objective for fitting latent SDEs. Traditional VI algorithms evaluate this objective by numerical simulation over a time discretization, trading fidelity for computational cost. A recent class of algorithms, simulation-free VI, sidesteps this tradeoff by parameterizing the posterior through its instantaneous marginals rather than its drift. In this work, we show that the efficiency of existing simulation-free VI algorithms comes at a price: their parameterizations restrict the approximate posterior to a subset of the SDEs available to simulation-based methods, degrading posterior inference and parameter learning. We propose Helmholtz-SDE, a simulation-free VI algorithm that closes this gap by optimizing over path laws compatible with a prescribed collection of marginals. Helmholtz-SDE recovers dynamics more faithfully than prior simulation-free methods, with the largest gains under high posterior uncertainty. It further matches the performance of simulation-based VI at a fraction of the runtime. Code is available at https://github.com/lindermanlab/helmholtz-sde .
Henry D. Smith, Brian L. Trippe, Scott W. Linderman
Stochastic differential equations (SDEs) provide a flexible framework for modeling temporal dynamics in partially observed systems. A central task is to calibrate such models from data, which requires inferring latent trajectories and parameters from sparse, noisy observations. Classical smoothing methods for this problem are often limited by path degeneracy and poor scalability. In this work, we developed a novel method based on characterization of the posterior SDE in terms of conditional backward-in-time score defined as the gradient of a function solving a Kolmogorov backward equation with multiplicative updates at observation times. We learn this conditional score using neural networks trained to satisfy both the governing PDE and the observation-induced jump conditions, thereby integrating continuous-time dynamics with discrete Bayesian updates. The resulting score induces a posterior SDE with the same diffusion coefficient but a modified drift, enabling efficient posterior trajectory sampling. We further derive a likelihood-based objective for learning the SDE parameters, yielding an evidence lower bound (ELBO) for joint state smoothing and parameter estimation. This leads to a variational EM-style procedure, where the neural conditional score is optimized to approximate the smoothing distribution, followed by a maximization step over the SDE parameters using samples from the induced posterior. Experiments on nonlinear systems demonstrate accurate and stable inference with a very few observations demonstrating significant improved scalability compared to classical MCMC methods.
Yu Wang, Arnab Ganguly
Department of Mathematics Louisiana State University