We study the problem of learning transition kernels for time-homogeneous jump-diffusion processes using conditional diffusion models, with the goal of generating new sample paths from training data consisting of N independent trajectories observed on a high-frequency discrete time grid. On the theoretical side, we establish non-asymptotic bounds for the conditional score estimation error and for the KL divergence between the laws of the true and generated discretely observed paths. On the numerical side, we first evaluate our method on synthetic data to assess the theoretical findings and benchmark its performance against the approach of Gao et al. (2025). We then apply our method to real-world data and investigate its performance on a probabilistic forecasting task.
Figures & tables
Figure 1: Comparison of test sample paths with sample paths generated by our method (left) and by the method in Gao et al. (2025) (right).
Figure 2: Empirical convergence rates for d=1 (left) and d=2 (right).
Figure 3: Comparison of the empirical convergence rates for sample paths generated by our method and by Gao et al. (2025) for d=1 (left) and d=2 (right).
Figure 4: One-step forecasts of our estimator on the test segment for BTC and ETH in the Crypto dataset (top) and for the Danish dataset (bottom).
Figure 5: One-step forecasts of the time-homogeneous SBJTS estimator on the test segment for BTC and ETH in the Crypto dataset (top) and for the Danish dataset (bottom). The curves and the band are constructed as in Figure 4 .
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
d
N
Our estimator
Estimator of Gao et al. (2025)
1
100
0.165±0.093
3.265±0.153
1
200
0.140±0.086
3.120±0.186
1
500
0.087±0.037
3.055±0.127
1
1000
0.095±0.032
2.923±0.106
1
2000
0.075±0.033
2.685±0.069
1
5000
0.050±0.019
2.333±0.065
Appendix
Table 1: Joint KL estimates (mean ±95% CI half-width) for our estimator and the estimator of Gao et al. (2025) .
Figure 6: One-step forecasts of our estimator (left) and time-homogeneous SBJTS estimator from De Marco et al. (2026) (right) on the test segment.
We propose a framework for generative modeling of continuous-time processes from irregularly and asynchronously recorded data. It is based on the matching of generators and accommodates discontinuous trajectories. Analytical formulas for diffusion and jump bridges yield a family of reference generators that a neural network is trained to match. The key ingredient is that, for our constructed jump bridge, a parametrization of the jump kernel densities by scaled Gaussians admits closed-form expressions for the Kullback-Leibler divergence, allowing simulation-free training.
O. Pfohl, J. Chemseddine, P. Hagemann +3
Department of Mathematics, Humboldt-Universität zu Berlin, 12489, Berlin, Germany · Department of Mathematics, Technische Universität Berlin, 10623, Berlin, Germany · Bundesanstalt für Materialforschung und -prüfung, 12205, Berlin, Germany +1
Time series driven by unobserved latent states frequently exhibit abrupt jump discontinuities whose timing and magnitude cannot be predicted from observed history alone. Classical jump-diffusion models offer a principled mathematical framework but assume rigid parametric forms, while recent neural jump models operate on fully observed trajectories without inferring the hidden states that govern the dynamics. We propose \textit{Deep ZakaiJ}, a latent-state model for partially observed jump-diffusion systems that embeds the Zakai nonlinear filtering equation into a neural encoder--decoder architecture. The encoder recursively updates a belief over the latent state via Strang splitting into three interpretable substeps: prior propagation, diffusion innovation, and jump innovation, yielding a differentiable, first-order-accurate approximation of the exact filtering evolution. The decoder is a structured jump-diffusion model explicitly conditioned on the filtered belief, preserving the separation between continuous dynamics and discontinuous shocks. On synthetic, financial, and oceanographic datasets, \textit{Deep ZakaiJ} improves distributional forecasts while remaining competitive in point accuracy, achieving calibrated predictive intervals and recovering interpretable latent structure in synthetic and qualitative case studies.
Yan Leng, Thibaut Mastrolia, Hao Wang
University of Texas at Austin · University of California, Berkeley
Sequential Monte Carlo (SMC) methods are a natural tool for post-hoc conditioning of pretrained generative models, but in many applications the mutation kernels used by the particle system are biased approximations of an ideal Feynman--Kac flow. This paper develops a non-asymptotic error analysis for such SMC samplers. Under forward-smoothing forgetting conditions, we decompose the total error into a kernel bias, measuring the effect of replacing the ideal transition kernels by approximate ones, and a finite-particle Monte Carlo error. Our approach relies on extending local Doeblin-type conditions and Lyapunov drift arguments for Markov kernels to conditional distributions, thereby enabling a principled control of the bias. We then instantiate this general framework for conditional sampling with score-based diffusion models, and derive the first non-asymptotic error bound that jointly controls initialization error, time discretization, and score approximation in the reverse diffusion dynamics as well as finite-particle Monte Carlo error.
Stanislas Strasman, Gabriel Victorino Cardoso, Sylvain Le Corff +2
SU, LPSM · Sorbonne Université and Université Paris Cité, CNRS, LPSM, F-75005 Paris, France · LPSM +2