Recent advances in generative modeling have substantially improved time series generation, yet most existing methods either assume a regular temporal grid or focus on feature dynamics under a given sampling structure. This makes them illsuited for generating irregular time series in their native form, where a model must capture not only feature values, but also how many observations occur, when they occur, and which features are observed together. To address this heterogeneous generation problem, we propose ChronoFlow, a unified hierarchical flow matching framework organized by statistical granularity. Following a coarse-to-fine hierarchy, ChronoFlow first generates observation counts and feature-wise frequencies, then jointly generates observation times and feature co-observation patterns, and finally generates values conditioned on the realized pattern. This turns a complex joint generation problem into structurally aligned subproblems while preserving their dependencies. To evaluate complete irregular time series generation, we introduce complementary metrics spanning sample realism, sampling structure, value fidelity, and temporal and cross-feature dependencies, and validate them through controlled corruptions. Across five benchmarks, ChronoFlow achieves strong improvements in generation fidelity over existing baselines, while factorization studies support the proposed hierarchy. Our code is available at https://anonymous.4open.science/r/ChronoFlow.
Figures & tables
Figure 1: Overview of ChronoFlow for irregular time series generation. IntensityFlow generates global statistics, PatternFlow joint times and panels, and ValueFlow conditional values. The feature-by-time display is B⊤ . Figure notation s , zI , and P corresponds to flow time t , state u , and pattern (τ,B) .
Dataset
Model
Set-Discr. ↓
CoOb ↓
Pattern-MMD 2 ↓
Trans-MMD 2 ↓
Value-W1 ↓
Set-Corr. ↓
eICU
Oracle
0.0019 ± 0.0012
0.0306 ± 0.0155
0.0000 ± 0.0000
0.0004 ± 0.0001
0.0312 ± 0.0007
0.0364 ± 0.0015
Poisson + LatentODE
0.4242 ± 0.0317
1.9710 ± 0.0017
0.0047 ± 0.0024
0.0249 ± 0.0038
0.2216 ± 0.0321
0.0643 ± 0.0049
LogNormMix + mTAN
0.4555 ± 0.0021
0.0383 ± 0.0141
0.0011 ± 0.0015
0.0169 ± 0.0025
0.2910 ± 0.0319
0.0813 ± 0.0060
LogNormMix + TFM
0.4518 ± 0.0122
0.0374 ± 0.0042
0.0014 ± 0.0019
0.0247 ± 0.0071
0.3585 ± 0.0321
0.1435 ± 0.0071
FlexTPP
0.3629 ± 0.0288
0.1336 ± 0.1033
0.0032 ± 0.0014
0.0070 ± 0.0030
0.1081 ± 0.0396
0.0532 ± 0.0112
ChronoFlow
0.0995 ± 0.0098
0.0380 ± 0.0047
0.0002 ± 0.0001
0.0045 ± 0.0002
0.0504 ± 0.0047
0.0385 ± 0.0003
Table 1: Generation fidelity across five irregular time series benchmarks. We compare ChronoFlow with four baselines using six complementary metrics covering overall sample realism, sampling structure, temporal dynamics, value fidelity, and cross-feature dependencies. Oracle denotes a real–real reference. Bold indicates the lowest mean among generators for each dataset and metric.
Model
Process
Size
Set-Discr. ↓
CoOb ↓
Pattern-MMD 2 ↓
Trans-MMD 2 ↓
Value-W1 ↓
Set-Corr. ↓
Oracle
–
–
0.0019 ± 0.0012
0.0306 ± 0.0155
0.0000 ± 0.0000
0.0004 ± 0.0001
0.0312 ± 0.0007
0.0364 ± 0.0015
Value → Pattern
M→x→(τ,B)
2.70M
0.4511 ± 0.0217
0.2354 ± 0.0098
0.0051 ± 0.0001
0.0099 ± 0.0003
0.1142 ± 0.0022
0.0461 ± 0.0046
Joint TPV
M→(τ,B,x)
1.12M
0.2639 ± 0.0039
0.0867 ± 0.0033
0.0008 ± 0.0002
0.0275 ± 0.0013
0.3179 ± 0.0044
0.0698 ± 0.0101
Joint All
(M,τ,B,x)
2.21M
0.2445 ± 0.1387
0.0731 ± 0.0055
0.0018 ± 0.0002
0.0074 ± 0.0009
0.0779 ± 0.0011
0.0423 ± 0.0012
w/o OT
(M,r)→(τ,B)→x
2.06M
0.4316 ± 0.0085
0.0854 ± 0.0042
0.0382 ± 0.0031
0.0083 ± 0.0004
0.0731 ± 0.0038
0.0382 ± 0.0037
ChronoFlow
(M,r)→(τ,B)→x
2.06M
0.0995 ± 0.0098
0.0380 ± 0.0047
0.0002 ± 0.0001
0.0045 ± 0.0002
0.0504 ± 0.0047
0.0385 ± 0.0003
Table 2: Ablation study of ChronoFlow on eICU. We compare alternative generation factorizations and evaluate the effect of optimal-transport coupling in PatternFlow. M is the number of observation occasions and r contains feature-wise observation frequencies. τ denotes observation times, B feature panels, and x feature values.
Figure 2: Qualitative analysis on eICU. Top: real and generated samples from LogNormMix + mTAN, FlexTPP, and ChronoFlow. Bottom: ground-truth versus ChronoFlow distributions of observation counts, feature co-observation, and Heart Rate and Temperature values.
Figure 3: Further analysis. (a,b) eICU latency–fidelity and ODE-step sensitivity, with bubble size indicating parameter count in (a). (c) Designated-corruption responses normalized by endpoint change within each dataset, then averaged across five datasets. Shading is between-dataset sample standard deviation. (d) Relative positive responses across corruptions, averaged over available datasets. Definitions and missing-support exceptions are in Section D.7 .
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
τtP←(1−tP)τ0+tPτ1,BtP←(1−tP)B0+tPB1
Appendix
Algorithm 1 Training procedure for ChronoFlow.
(τˉ,Bˉ)←ΦODE(vθP,(τ0,B0);0→1∣M^,r^)
Appendix
Algorithm 2 Sampling procedure for ChronoFlow.
eICU
P12
MIMIC-III
MIMIC-IV
Activity
# Features
11
35
28
28
12
Time Window
24h
48h
48h
48h
2.28s
Min Resolution
1m
1m
1m
1m
10ms
Max Measurements ( Nmax )
672
1,024
1,024
1,024
600
Train (Neg / Pos)
13,817 / 1,346
6,579 / 1,093
10,655 / 1,611
13,806 / 2,046
2,860
Val (Neg / Pos)
3,454 / 336
1,645 / 273
2,662 / 404
3,452 / 512
471
Appendix
Table A3: Dataset statistics for the five irregular time series benchmarks. Clinical split sizes are reported as negative/positive outcome counts, while Activity split sizes indicate the total number of windows across its seven activity classes. Nmax is the measurement budget, not the number of observation occasions.
Model
(M,r)
(τ,B)
x
Set-Discr. ↓
CoOb ↓
Pattern-MMD 2 ↓
Trans-MMD 2 ↓
Value-W1 ↓
Set-Corr. ↓
Oracle
GT
GT
GT
0.0019 ± 0.0012
0.0306 ± 0.0155
0.0000 ± 0.0000
0.0004 ± 0.0001
0.0312 ± 0.0007
0.0364 ± 0.0015
Stage-wise Oracle Intervention
ChronoFlow
sampled
sampled
sampled
0.0995 ± 0.0098
0.0380 ± 0.0047
0.0002 ± 0.0001
0.0045 ± 0.0002
0.0504 ± 0.0047
0.0385 ± 0.0003
+ GT (M,r)
GT
sampled
sampled
0.0153 ± 0.0140
0.0128 ± 0.0011
0.0000 ± 0.0000
0.0040 ± 0.0002
0.0481 ± 0.0037
0.0365 ± 0.0022
+ GT (M,r) , (τ,B)
GT
GT
sampled
0.0928 ± 0.1521
0.0000 ± 0.0000
0.0000 ± 0.0000
0.0003 ± 0.0002
0.0508 ± 0.0100
0.0358 ± 0.0019
Robustness to Upstream Perturbations
Appendix
Table A4: Stage-wise oracle intervention and perturbation analysis of ChronoFlow on eICU. At inference, generated upstream variables are replaced by their ground-truth counterparts to assess sensitivity to upstream conditioning. We further perturb ground-truth upstream variables with relative noise of severity s to evaluate downstream robustness; ground-truth variables are taken from the evaluated test samples, so these rows can fall below the real–real reference. (M,r) denotes global observation statistics, (τ,B) the sampling pattern, and x the feature values.
Model
Optimizer
LR setting
Schedule
Poisson + LatentODE
Adamax
10−2
Exponential decay
LogNormMix + mTAN (Pattern)
Adam
2×10−4
Constant
LogNormMix + mTAN (Value)
Adam
2×10−3
Constant
LogNormMix + TFM
Adam
2×10−4
Constant
FlexTPP
AdamW
5×10−4
One-cycle
ChronoFlow
Adam
4×10−5
Warmup + plateau
Appendix
Table A5: Optimization settings for the evaluated eICU checkpoints. The listed learning rate is the peak rate for FlexTPP and the configured base rate for ChronoFlow. ChronoFlow uses a warmup target of 8×10−4 and a minimum learning rate of 10−5 .
Figure A4: UMAP visualizations of real and generated samples across the five benchmarks: eICU, P12, MIMIC-III, MIMIC-IV, and Activity, shown from top to bottom. For each dataset, the panels compare the real-data representation with samples generated by Poisson + LatentODE, LogNormMix + mTAN, LogNormMix + TFM, FlexTPP, and ChronoFlow. Overlap is interpreted qualitatively, as two-dimensional projection can hide discrepancies.
Figure A5: Qualitative comparison of baseline methods and ChronoFlow on the eICU benchmark. For this dataset, rows correspond to Poisson + LatentODE, LogNormMix + mTAN, LogNormMix + TFM, FlexTPP, and ChronoFlow, respectively. For each model, (a) shows the distribution of the total number of observations per stay. Panel (b) shows feature co-observation rates for all 11 features. The lower triangle shows ground truth, the upper triangle shows generated samples, and the diagonal shows ground-truth marginals. Panels (c, d) show value distributions of the most frequent feature and the median-frequency feature.
Figure A6: Qualitative comparison of baseline methods and ChronoFlow on the P12 benchmark. For this dataset, rows correspond to Poisson + LatentODE, LogNormMix + mTAN, LogNormMix + TFM, FlexTPP, and ChronoFlow, respectively. For each model, (a) shows the distribution of the total number of observations per stay. Panel (b) shows feature co-observation rates for the 12 most frequent features. The lower triangle shows ground truth, the upper triangle shows generated samples, and the diagonal shows ground-truth marginals. Panels (c, d) show value distributions of the most frequent feature and the median-frequency feature.
Figure A7: Qualitative comparison of baseline methods and ChronoFlow on the MIMIC-III benchmark. For this dataset, rows correspond to Poisson + LatentODE, LogNormMix + mTAN, LogNormMix + TFM, FlexTPP, and ChronoFlow, respectively. For each model, (a) shows the distribution of the total number of observations per stay. Panel (b) shows feature co-observation rates for the 12 most frequent features. The lower triangle shows ground truth, the upper triangle shows generated samples, and the diagonal shows ground-truth marginals. Panels (c, d) show value distributions of the most frequent feature and the median-frequency feature.
Figure A8: Qualitative comparison of baseline methods and ChronoFlow on the MIMIC-IV benchmark. For this dataset, rows correspond to Poisson + LatentODE, LogNormMix + mTAN, LogNormMix + TFM, FlexTPP, and ChronoFlow, respectively. For each model, (a) shows the distribution of the total number of observations per stay. Panel (b) shows feature co-observation rates for the 12 most frequent features. The lower triangle shows ground truth, the upper triangle shows generated samples, and the diagonal shows ground-truth marginals. Panels (c, d) show value distributions of the most frequent feature and the median-frequency feature.
Figure A9: Qualitative comparison of baseline methods and ChronoFlow on the Activity benchmark. For this dataset, rows correspond to Poisson + LatentODE, LogNormMix + mTAN, LogNormMix + TFM, FlexTPP, and ChronoFlow, respectively. For each model, (a) shows the distribution of the total number of observations per window. Panel (b) shows feature co-observation rates for the 12 most frequent features. The lower triangle shows ground truth, the upper triangle shows generated samples, and the diagonal shows ground-truth marginals. Panels (c, d) show value distributions of the most frequent feature and the median-frequency feature.
Figure A10: Qualitative comparison of generated samples on the eICU benchmark. Rows correspond to ground truth, Poisson + LatentODE, LogNormMix + mTAN, LogNormMix + TFM, FlexTPP, and ChronoFlow, respectively. Each row shows four samples. Colors indicate standardized feature values, and the total number of observations is shown above each sample.
Figure A11: Qualitative comparison of generated samples on the P12 benchmark. Rows correspond to ground truth, Poisson + LatentODE, LogNormMix + mTAN, LogNormMix + TFM, FlexTPP, and ChronoFlow, respectively. Each row shows four samples. Colors indicate standardized feature values, and the total number of observations is shown above each sample.
Figure A12: Qualitative comparison of generated samples on the MIMIC-III benchmark. Rows correspond to ground truth, Poisson + LatentODE, LogNormMix + mTAN, LogNormMix + TFM, FlexTPP, and ChronoFlow, respectively. Each row shows four samples. Colors indicate standardized feature values, and the total number of observations is shown above each sample.
Figure A13: Qualitative comparison of generated samples on the MIMIC-IV benchmark. Rows correspond to ground truth, Poisson + LatentODE, LogNormMix + mTAN, LogNormMix + TFM, FlexTPP, and ChronoFlow, respectively. Each row shows four samples. Colors indicate standardized feature values, and the total number of observations is shown above each sample.
Figure A14: Qualitative comparison of generated samples on the Activity benchmark. Rows correspond to ground truth, Poisson + LatentODE, LogNormMix + mTAN, LogNormMix + TFM, FlexTPP, and ChronoFlow, respectively. Each row shows four samples. Colors indicate standardized feature values, and the total number of observations is shown above each sample.
Department of Artificial Intelligence, Korea University, Seoul, South Korea · Department of Applied Mathematics and Theoretical Physics, University of Cambridge, Cambridge, United Kingdom.