Organizations: Department of Electrical and Computer Engineering, Johns Hopkins University, Baltimore, MD, USA · Gordon Center for Medical Imaging, Department of Radiology, Massachusetts General Hospital and Harvard Medical School, Boston, MA, USA · Center for Advanced Medical Computing and Analysis, Department of Radiology, Massachusetts General Hospital and Harvard Medical School, Boston, MA, USA
Cine cardiovascular magnetic resonance (CMR) captures the cardiac cycle as a four-dimensional (4D) sequence, but standard acquisition requires electrocardiogram (ECG) gating and repeated breath holds. Visual realism alone does not establish accurate patient-specific ejection fraction (EF) or ventricular volumes. We present PhaseFlow3D, a generative framework that synthesizes a complete 4D cine sequence from a single end-diastolic (ED) three-dimensional (3D) volume without ECG. To capture asymmetric systolic and diastolic dynamics, it represents the cardiac cycle as a piecewise linear phase anchored at ED and end-systolic (ES) time points. At inference, a population-level canonical template supplies this phase without patient-specific temporal information. A phase-conditioned rectified flow model generates a cardiac motion trajectory in latent space. Radial Contraction Decomposition converts each latent state into a 3D displacement field, combining a physics-informed radial component for centripetal myocardial contraction with an image-conditioned residual for rotation and out-of-plane motion. Each frame is generated by directly warping the ED volume, bypassing variational autoencoder decoding. On the combined ACDC and M&Ms benchmark, PhaseFlow3D achieves the lowest EF mean absolute error, the only positive left-ventricular volume-curve R2, and the best distributional quality among compared methods. Ablations confirm each component's contribution. Downstream evaluations demonstrate the utility of the synthesized sequences and displacement fields for segmentation, pathology classification, label propagation, and myocardial strain analysis.
Figures & tables
Figure 1: Uniform linear phase (top) scatters ES across ϕ values; the piecewise linear phase (bottom) aligns ES at ϕ=π by construction across pathological types.
Figure 2: PhaseFlow3D encodes the ED 3D volume into a latent anchor z0 , predicts a phase-conditioned latent velocity via rectified flow matching, and decodes a 3D displacement field through a Radial Contraction Decomposition. The displacement field directly warps the ED volume to synthesize the full 4D cine sequence.
Image & Video Quality
Physiological Fidelity
Method
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
FVD ↓
EF MAE ↓
Vol Corr ↑
Vol R2↑
Vol MAE ↓
ConvLSTM [ 45 ]
19.16
0.426
0.474
292.0
23.31
—
—
− 37.80
0.734
GAN-cMRI [ 14 ]
22.64
0.542
0.409
185.8
19.70
39.35%
0.308
− 32.91
0.673
LFDM [ 26 ]
28.50
0.800
0.200
82.53
11.24
29.83%
0.809
− 5.547
0.221
CVAE
33.01
0.899
0.066
27.95
4.141
32.39%
0.781
− 0.582
0.179
DragNet [ 7 ]
35.61
0.928
0.030
5.450
1.690
37.77%
0.757
− 0.484
0.175
Table 1: Main results on the combined ACDC and M&Ms test set. Best in bold , second-best underlined . Baselines are evaluated with image-based EF (SegUNet on generated frames), whereas PhaseFlow3D uses warped ED masks; EF under unified measurement protocols is cross-validated in Appendix D (Table 4 ); “—” denotes inapplicable metrics.
Wilcoxon p (PhaseFlow3D vs.)
Metric
ED Rep.
ConvLSTM
GAN-cMRI
LFDM
CVAE
DragNet
SimVP
Dir. Reg.
EF MAE
4.5×10−28
—
9.5×10−20
2.5×10−18
1.1×10−19
9.8×10−23
1.1×10−14
1.6×10−17
Vol Corr
—
—
2.2×10−22
0.84
5.6×10−3
1.0×10−2
1.1×10−10
3.9×10−9
Vol R2
1.3×10−21
5.4×10−29
8.4×10−29
4.7×10−13
6.7×10−11
6.1×10−10
2.5×10−12
1.3×10−13
Vol MAE
2.2×10−27
5.4×10−29
7.9×10−29
3.7×10−18
1.1×10−17
2.8×10−16
2.5×10−19
5.0×10−19
Table 2: Paired Wilcoxon signed-rank test ( n=166 patients) against all baselines. Every cell is a p -value; smaller means stronger evidence that PhaseFlow3D is better, and in every significant cell the per-patient means favour PhaseFlow3D. Bold = significant at α=0.05 . The only non-significant cell is Vol Corr vs. LFDM; “—” marks undefined comparisons (constant or empty predicted volume curves).
Figure 3: Per-patient distribution of image quality (PSNR, SSIM, LPIPS) and physiological fidelity (EF MAE, Vol Corr, Vol R2 ) on the combined ACDC and M&Ms test set ( n=166 ). Each box spans the interquartile range with the median line and 5–95th percentile whiskers. PhaseFlow3D (orange) is compared against the eight baselines (blue); empty slots mark metrics that are undefined for a given method, matching the dashes in Table 1 . Vol R2 uses a symmetric-log axis to resolve both the high-performing methods near one and the collapsed baselines.
Image & Video Quality
Physiological Fidelity
Variant
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
FVD ↓
EF MAE ↓
Vol Corr ↑
Vol R2↑
Vol MAE ↓
A1: w/o deform decoder
35.37
0.929
0.012
1.949
0.724
55.33%
—
− 1.766
0.250
A2: w/o radial head
34.48
0.924
0.017
2.155
0.765
15.00%
0.851
− 0.244
0.114
A3: radial only (w/o residual)
33.14
0.922
0.020
3.098
1.058
18.48%
0.823
− 0.415
0.132
A4: w/o vol curve loss
34.33
0.924
0.020
3.063
0.998
14.66%
0.785
− 0.643
0.143
A5: linear phase
34.91
0.927
0.017
1.853
0.697
12.60%
0.741
− 0.198
0.126
Table 3: Ablation study on the combined ACDC and M&Ms test set. Best in bold , second-best underlined .
Figure 4: Group-mean piecewise phase templates (coloured; ACDC pathology groups and pooled M&Ms) vs. the canonical template used at inference (black) and the uniform linear baseline (grey dashed). Each template reaches ϕ=π (ES) at the group-mean ES fraction, encoding pathology-specific systole–diastole timing.
Figure 5: Per-frame phase RMSE (% of cycle) of the canonical template (solid blue), the group-mean oracle template (dotted green), and the uniform linear baseline (grey dashed), per ACDC pathology group and for the pooled M&Ms cohort. Shaded region: error saved by the canonical template over the linear baseline.
Figure 6: det(Jφ) distribution of the composed 3D displacement field (radial and scaling-and-squaring residual, K=7 ) on the ACDC + M&Ms test set. Left : global histogram (log scale); red dashed: fold boundary ( det=0 ). Right : per-group mean ± std. Folding voxels are 0.0005% on average ( <0.002% per patient), indicating a low measured folding rate.
Figure 7: Spatial view of deformation regularity for a representative HCM subject (patient036, mid-ventricular slice) at six frames across the cardiac cycle. Rows: ground-truth cine, PhaseFlow3D generation, displacement magnitude ∣φ∣ , and Jacobian determinant det(Jφ) . The deformation is localized to the myocardium; det(Jφ)<1 (blue) marks local contraction and >1 (red) local expansion, and it stays strictly positive everywhere, confirming a fold-free diffeomorphic mapping.
Figure 8: Qualitative comparison (MINF and NOR, mid-ventricular slice). Rows 1–3: ground truth, the VAE-decoded variant (A1), and PhaseFlow3D at ED, mid-systole, ES, and mid-diastole. Row 4: predicted displacement magnitude ∣φ∣ . PhaseFlow3D produces sharp, anatomically coherent frames; the VAE-decoded variant exhibits characteristic reconstruction blur.
Figure 9: Per-pixel absolute reconstruction error ∣Gen−GT∣ for three representative subjects (rows) across six frames spanning the cardiac cycle (columns), on the mid-ventricular slice. The error stays near zero over the static background and concentrates in a thin band at the myocardial boundary, peaking around ES where motion is largest.
Figure 10: LV volume curves (% of ED volume). Ground-truth (solid) vs. PhaseFlow3D-predicted (dashed) trajectories for one representative patient per ACDC pathology group.
Figure 11: Segmentation substitution. Test Dice (LV and myocardium mean, 3-seed mean) vs. the fraction of real supervision replaced by synthetic supervision, for phase-aligned substitution (pipeline A, blue) and full-cycle substitution (pipeline B, red). Dashed lines: each pipeline’s all-real baseline. Right edge: retention, the ratio of fully synthetic to all-real Dice.
Figure 12: Classification substitution. Balanced accuracy and macro-F1 (3-seed mean ± std) vs. the fraction of real geometry samples replaced by synthetic ones. Dashed line: all-real balanced accuracy; dotted line: 5-class chance. Right edge: balanced-accuracy retention.
Figure 13: Label propagation. LV Dice of the ED annotation warped to ES, per ACDC pathology group and pooled (All), for the no-motion floor, our blind single-frame propagation, and the Demons two-frame oracle that sees the real ES image. Bars: patient mean; error bars: patient std.
Figure 14: Peak-systolic GCS from the synthesized motion, by ACDC pathology group (blue) and for the pooled M&Ms cohort (grey). Bars: group mean; error bars: patient standard deviation. The Kruskal–Wallis test is computed across the five ACDC groups.
Figure 15: Perceptual realism. Left: pooled out-of-fold ROC of the real-vs-synthetic discriminator (3-seed mean, shaded band: seed std); the dotted diagonal is chance. Right: AUC by cycle third; mid-systole carries most of the residual separability.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Image-based
Mask propagation
GAN-cMRI [ 14 ]
38.68±18.07
—
LFDM [ 26 ]
29.67±15.76
—
CVAE
31.70±18.55
—
DragNet [ 7 ]
36.49±16.43
37.73±16.80
SimVP [ 25 ]
26.22±18.08
—
Direct Reg
29.71±17.66
—
Appendix
Table 4: EF MAE (%, patient mean ± std, n=166 ) under the two unified measurement protocols, both referenced against the expert annotations. “—”: the method exposes no pixel-space displacement field, so mask propagation is undefined. ConvLSTM is omitted because its EF is undefined under either protocol (cf. Table 1 ). For GAN-cMRI, 38 cases whose generated frames yield an empty LV cavity are excluded ( n=128 ).
Metric
Mean ± Std
Median [IQR]
Image quality
PSNR (dB)
34.97±3.24
34.67[32.66,37.32]
SSIM
0.927±0.012
0.926[0.917,0.936]
LPIPS
0.017±0.006
0.015[0.013,0.021]
Physiological fidelity
EF MAE (%)
12.85±11.03
9.49[4.58,18.41]
Appendix
Table 5: Per-patient dispersion of PhaseFlow3D on the joint ACDC + M&Ms test set ( n=166 patients), reported as mean ± standard deviation and as median with interquartile range.
Quantity
Train ( n=255 )
Val ( n=49 )
σ of z1
0.800
0.794
σ of anchor Evae(xED)
0.798
0.793
σ of z1−repeatT(⋅)
0.131
0.122
α/σ(z1)
12.5%
12.6%
α/σ(z1−repeatT(⋅))
76.6%
82.1%
Appendix
Table 6: Empirical latent statistics ( σ = per-element standard deviation, averaged over patients). The two splits agree closely, indicating the latent scale is a stable property of the frozen encoder rather than a split artifact.
Figure 16: Per-pathology qualitative results (one representative patient per group). Top row : ground truth; bottom row : PhaseFlow3D synthesis. ED frame ( t=0 , leftmost) is the sole model input.
Synthetic fraction
Pipeline A
Pipeline B
0%
0.894±0.005
0.891±0.008
25%
0.898±0.003
0.898±0.005
50%
0.900±0.001
0.885±0.003
75%
0.896±0.000
0.890±0.006
90%
0.890±0.003
0.868±0.011
100%
0.887±0.002
0.863±0.009
Appendix
Table 7: Segmentation substitution: test Dice (LV and myocardium mean, 3-seed mean ± std) at each synthetic supervision fraction.
Synthetic fraction
Balanced accuracy
Macro-F1
0%
0.733±0.048
0.731±0.052
25%
0.711±0.048
0.705±0.052
50%
0.600±0.065
0.593±0.067
75%
0.570±0.055
0.551±0.064
90%
0.526±0.010
0.489±0.009
100%
0.533±0.018
0.492±0.021
Appendix
Table 8: Classification substitution: five-class balanced accuracy and macro-F1 (3-seed mean ± std) at each synthetic substitution fraction.
Structure
No-motion floor
Ours (single-frame)
Demons (oracle)
LV
0.622±0.177
0.712±0.129
0.738±0.138
Myocardium
0.486±0.164
0.572±0.121
0.700±0.093
Appendix
Table 9: Label propagation on the ACDC test set ( n=30 ): Dice of the ED annotation warped to ES (patient mean ± std) per structure and arm.
Cine cardiovascular magnetic resonance (CMR) analysis relies on multi-frame sequences capturing the full cardiac cycle. However, standard multi-frame acquisition depends heavily on electrocardiogram (ECG) gating and repeated breath-holds, posing challenges in uncooperative populations, resource-limited settings, and temporally corrupted datasets. Existing methods that synthesize full cardiac sequences either rely on explicit ECG signals to parameterize myocardium function, or employ deformable registration without physiological constraints, failing to faithfully reproduce clinically relevant dynamic metrics such as ejection fraction (EF) and ventricular contraction magnitude. We present PhaseFlow, a unified generative framework that overcomes both limitations. PhaseFlow estimates a non-linear cardiac phase signal directly from the input sequence via a segmentation-derived left-ventricular (LV) area curve, capturing the asymmetric dynamics of systole and diastole without any ECG dependency. At inference, this phase signal is provided by a pathology-specific template, informing phase-specific frame generation. A rectified flow model conditioned on the phase and slice position synthesizes the full cardiac motion trajectory in the latent space, decoded into a diffeomorphic displacement field that warps end-diastole pixel intensities directly, eliminating the reconstruction blur often accompanying the variational autoencoder. On the ACDC benchmark, PhaseFlow achieves superior physiological fidelity and image realism, with best LV volume curve R2, structural similarity (SSIM) and generative quality (FID) among all baselines. Ablation studies confirm that each proposed component contributes measurably to the overall performance.
Shiyi Wang, Ruochen Sun, Peirong Liu +2
Department of Electrical and Computer Engineering, Johns Hopkins University, Baltimore, MD, USA · Gordon Center for Medical Imaging, Department of Radiology, Massachusetts General Hospital and Harvard Medical School, Boston, MA, USA · Center for Advanced Medical Computing and Analysis, Department of Radiology, Massachusetts General Hospital and Harvard Medical School, Boston, MA, USA
Cardiac Magnetic Resonance (CMR) imaging provides a comprehensive assessment of cardiac structure and function but remains constrained by high acquisition costs and reliance on expert annotations, limiting the availability of large-scale labeled datasets. In contrast, electrocardiograms (ECGs) are inexpensive, widely accessible, and offer a promising modality for conditioning the generative synthesis of cine CMR. To this end, we propose ECGFlowCMR, a novel ECG-to-CMR generative framework that integrates a Phase-Aware Masked Autoencoder (PA-MAE) and an Anatomy-Motion Disentangled Flow (AMDF) to address two fundamental challenges: (1) the cross-modal temporal mismatch between multi-beat ECG recordings and single-cycle CMR sequences, and (2) the anatomical observability gap due to the limited structural information inherent in ECGs. Extensive experiments on the UK Biobank and a proprietary clinical dataset demonstrate that ECGFlowCMR can generate realistic cine CMR sequences from ECG inputs, enabling scalable pretraining and improving performance on downstream cardiac disease classification and phenotype prediction tasks.
Xiaocheng Fang, Zhengyao Ding, Guangkun Nie +9
State Key Laboratory of General Artificial Intelligence, Peking University · School of Intelligence Science and Technology, Peking University, Beijing, China · Polytechnic Institute of Zhejiang University, Zhejiang University, Hangzhou, China +4
Accurate 4D whole-heart mesh reconstruction from sparse cine MRI is critical for creating cardiac digital twins, but remains challenging due to limited 2D slice coverage and the complex coupling between cardiac shape and motion. Existing methods often rely on intermediate contour fitting and typically reconstruct static, single-phase, or partial cardiac geometries, limiting their ability to capture full-chamber dynamics. We propose a novel end-to-end framework for reconstructing temporally resolved whole-heart meshes from multi-view 2D cine MRI sequences by learning an image-to-mesh mapping. The framework incorporates a differentiable contour renderer inspired by the Beer-Lambert attenuation principle, enabling anatomy-aware supervision of 3D+t mesh deformation through contour-based projection losses. To improve temporal consistency across the cardiac cycle, we further introduce a multi-scale temporal modeling module that integrates global cycle-level dynamics with local inter-frame coherence to generate smooth and physiologically plausible mesh trajectories. The proposed method achieved a whole-heart mean absolute error of 1.68 ± 0.31 mm and a motion jitter of 0.77 ± 0.17 mm/frame3, outperforming existing methods with lower reconstruction error and substantially improved motion smoothness. It also improved 2D contour alignment across multiple cine MRI views and supported downstream proof-of-concept electrophysiological simulation. The code will be released publicly upon acceptance of the manuscript for publication.
Xiaoyue Liu, Dongcheng Cang, Xiaohan Yuan +3
Department of Biomedical Engineering, National University of Singapore, Singapore · School of Automation, Southeast University, Nanjing, China · Department of Medicine, National University of Singapore, Singapore +1