Acquiring a complete set of magnetic resonance imaging (MRI) contrasts is time-intensive and uncomfortable for patients, despite the diagnostic value of multi-contrast imaging. This motivates synthesizing missing contrasts from those already acquired, which is an inherently 3D problem requiring anatomical coherence across axial, sagittal, and coronal planes. However, fully 3D generative models are often impracti- cal under computational resources that scale cubically with volume size. We propose a unified Multi-Plane Autoregressive Diffusion (MPAD), a latent diffusion framework that achieves full-volume 3D synthesis using efficient plane-wise 2D operations while preserving volumetric coherence. A 3D autoencoder first compresses MRI scans into an isotropic 3D la- tent representation. A 2D diffusion model is then trained to reconstruct masked latent slices of the target contrast, conditioned on both source- contrast slices and unmasked target-contrast slices. During inference, we introduce plane-wise autoregressive synthesis with inter-plane priors. Slices are generated autoregressively in random order within one plane orientation to maintain intra-plane continuity, then propagated as con- ditioning priors to orthogonal plane orientations to enforce inter-plane consistency. Compared to 3D latent diffusion baselines, MPAD reduces training and inference FLOPs by 7x and 3x, respectively, while also lowering inference time and peak memory consumption. Experiments on multiple datasets demonstrate that MPAD achieves superior perfor- mance, generating high-fidelity 3D volumes and supporting one-to-many translation within a single unified model.
Figures & tables
Figure 1 : Overview of the proposed 3D multi-contrast MR image translation framework. Given a source-contrast (e.g., T1), the model synthesizes a target contrast (e.g., T2 or PD) in latent space using a Multi-modal Conditioning Encoder (MCE) and a 2D diffusion model with plane-wise autoregressive synthesis. Bottom: Training FLOPs per iteration (GFLOPs), inference FLOPs per volume (TFLOPs), inference time, peak memory, and PSNR averaged across all translation tasks on the ADNI dataset.
Figure 2 : Overview of the proposed training framework for 3D multi-contrast MR image synthesis. Stage 1: A 3D autoencoder encodes MR volumes into isotropic latent representations. Stage 2: A 2D diffusion model generates target-contrast latent slices conditioned on source-contrast slices, unmasked target-contrast slices, and text embeddings specifying the target modality. The Multi-modal Conditioning Encoder (MCE) fuses source and target features, which are then injected into the 2D diffusion model via SPADE modulation layers.
Figure 3 : (a): Random-order slice prediction within each plane, where slices are generated sequentially conditioned on previously predicted slices (purple) while unknown slices (gray) remain to be synthesized. (b): Four-pass generation process across planes. The first plane π1 is generated from pure noise at timestep T . Subsequent planes π2 and π3 start from a smaller timestep τ using previously completed planes as priors. A refinement pass on π1 incorporates information from all planes. The four candidates are aggregated via voxel-wise averaging to produce the final output.
T1 → T2
T1 → PD
T2 → T1
T2 → PD
PD → T1
PD → T2
Method
PSNR ↑
SSIM ↑
NMSE ↓
PSNR ↑
SSIM ↑
NMSE ↓
PSNR ↑
SSIM ↑
NMSE ↓
PSNR ↑
SSIM ↑
NMSE ↓
PSNR ↑
SSIM ↑
NMSE ↓
PSNR ↑
SSIM ↑
NMSE ↓
CycleGAN-3D [ 40 ]
24.88
0.810
0.130
20.09
0.774
0.179
21.49
0.807
0.197
19.63
0.761
0.204
19.08
0.751
0.248
19.23
0.751
0.450
EaGAN [ 36 ]
21.72
0.826
0.273
19.80
0.815
0.203
21.14
0.829
0.199
20.32
0.834
0.184
19.54
0.779
0.269
20.42
0.784
0.362
LDM-3D [ 26 ]
22.89
0.803
0.214
23.59
0.816
0.085
21.63
0.817
0.161
25.29
0.843
0.068
19.82
0.772
0.261
22.07
0.785
0.257
cWDM [ 8 ]
22.79
0.814
0.221
23.58
0.763
0.072
21.26
0.825
0.158
23.44
0.779
0.081
18.03
0.747
0.399
14.69
0.651
1.337
ALDM [ 15 ]
20.77
0.770
0.352
22.96
0.797
0.099
20.15
0.789
0.236
24.00
0.816
0.085
19.15
0.754
0.307
20.15
0.751
0.412
Table 1 : Quantitative comparison on the ADNI dataset. All metrics are evaluated on 3D volumes. Best scores are bolded.
T1 → T2
T1 → PD
T2 → T1
T2 → PD
PD → T1
PD → T2
Method
PSNR ↑
SSIM ↑
NMSE ↓
PSNR ↑
SSIM ↑
NMSE ↓
PSNR ↑
SSIM ↑
NMSE ↓
PSNR ↑
SSIM ↑
NMSE ↓
PSNR ↑
SSIM ↑
NMSE ↓
PSNR ↑
SSIM ↑
NMSE ↓
CycleGAN-3D [ 40 ]
27.67
0.845
0.126
24.84
0.829
0.089
22.67
0.803
0.314
25.81
0.853
0.067
23.71
0.815
0.169
24.06
0.834
0.289
EaGAN [ 36 ]
26.17
0.874
0.202
22.75
0.874
0.149
23.66
0.878
0.345
23.70
0.888
0.130
22.90
0.870
0.360
24.07
0.873
0.329
LDM-3D [ 26 ]
28.79
0.878
0.107
27.74
0.880
0.044
27.66
0.884
0.112
29.09
0.889
0.034
28.33
0.881
0.070
30.31
0.881
0.070
cWDM [ 8 ]
24.62
0.853
0.319
27.28
0.870
0.057
20.86
0.836
0.526
26.26
0.862
0.086
24.53
0.850
0.235
21.98
0.836
0.585
ALDM [ 15 ]
29.32
0.874
0.087
27.28
0.875
0.049
27.98
0.883
0.127
28.31
0.885
0.040
27.15
0.876
0.111
29.48
0.877
0.084
Table 2 : Quantitative comparison on the IXI dataset. All metrics are evaluated on 3D volumes. Best scores are bolded.
Figure 4 : Qualitative results on the ADNI dataset for T2→T1 synthesis. Each row shows axial, coronal, and sagittal views. Yellow arrows highlight anatomical regions where baseline methods show visible discrepancies compared to the ground truth (GT).
Figure 5 : Qualitative results on the IXI dataset for T1→T2 synthesis. Each row shows axial, coronal, and sagittal views. Yellow arrows highlight anatomical regions where baseline methods show visible discrepancies compared to the ground truth (GT).
ADNI
IXI
AR
g
PSNR ↑
NMSE ↓
PSNR ↑
NMSE ↓
✗
-
22.79
0.089
28.89
0.064
✓
16
22.83
0.089
29.06
0.063
✓
4
22.85
0.089
29.08
0.063
✓
1
22.87
0.088
29.14
0.061
Table 3: Ablation study on the autoregressive (AR) framework. PSNR and NMSE are averaged across all six translation tasks (T1 ↔ T2, T1 ↔ PD, T2 ↔ PD) on ADNI and IXI datasets.
ADNI
IXI
PSNR ↑
SSIM ↑
NMSE ↓
PSNR ↑
SSIM ↑
NMSE ↓
MCE
2D
22.82
0.813
0.091
28.18
0.879
0.090
3D
+0.72
+0.019
-0.013
+1.19
+0.007
-0.029
Masking
Learnable
23.22
0.803
0.133
27.37
0.869
0.102
Gaussian Noise
+0.32
+0.029
-0.055
+2.00
+0.017
-0.041
Prior
✗
22.42
0.795
0.087
27.49
0.873
0.098
Table 4: Ablation study on multi-plane conditional encoding design. PSNR, SSIM, and NMSE are averaged across all six translation tasks (T1 ↔ T2, T1 ↔ PD, T2 ↔ PD) on ADNI and IXI datasets. Each design choice is evaluated independently: MCE architecture (2D vs 3D convolution), masking type (learnable vs Gaussian noise), and prior information (off vs on). The first row in each block shows baseline performance, while the second row shows the performance change ( Δ ) when switching the design. Bold values indicate the magnitude of improvement or degradation.
Planes
ADNI
IXI
π1
π2
π3
π1′
PSNR ↑
SSIM ↑
NMSE ↓
PSNR ↑
SSIM ↑
NMSE ↓
✓
22.82
0.820
0.089
29.17
0.883
0.060
✓
✓
23.32
0.828
0.081
29.00
0.884
0.067
✓
✓
23.27
0.827
0.081
29.25
0.884
0.065
✓
✓
23.33
0.827
0.079
29.15
0.885
0.064
✓
✓
✓
23.49
0.830
0.078
29.29
0.886
0.062
Table 5: Ablation study on multi-plane generation. PSNR, SSIM, and NMSE are averaged across all six translation tasks (T1 ↔ T2, T1 ↔ PD, T2 ↔ PD) on ADNI and IXI datasets. Each row represents a different combination of planes used for generation. Checkmarks indicate which planes ( π1 : coronal, π2 : sagittal, π3 : axial, π1′ : refined coronal) are generated and averaged. Best scores are bolded.
Planes Ordering
PSNR ↑
SSIM ↑
π3→π1→π2→π3′
23.48
0.830
π2→π3→π1→π2′
23.49
0.830
π1→π2→π3→π1′
23.54
0.832
Table 6 : Ablation study on the plane ordering for translation tasks (T1 ↔ T2, T2 ↔ PD, and PD ↔ T1) on the ADNI dataset.
State Key Laboratory of Virtual Reality Technology and Systems, Beihang University · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences · Faculty of Computer Science and Control Engineering, Shenzhen University of Advanced Technology
Department of Computer Science, Stanford University, Stanford, CA, USA · Machine and Hybrid Intelligence Lab, Feinberg School of Medicine, Northwestern University, Chicago, IL, USA