This paper addresses cross-modal medical image segmentation, focusing on MRI-CT transfer in a source-only domain generalization setting. During training, only source-modality samples are available, while unlabeled target-modality images are used for testing. We propose LowBridge, which builds on the observation that cross-modal images share similar low-level features (e.g. edges) as they depict the same types of anatomical structures. Specifically, we first train a generative model to recover the source images from their edge features, followed by training a segmentation model on the generated source images, separately. At test time, edge features from the target images are input to the pretrained generative model to generate source-style target domain images, which are then segmented using the pretrained segmentation network. Experiments on various public datasets demonstrate that LowBridge achieves state-of-the-art performance, outperforming ten existing approaches. Ablation studies further show that LowBridge is compatible with different types of generative and segmentation models, suggesting its generalizability and potential to benefit from future advances in these models. The code will be available at https://github.com/JoshuaLPF/LowBridge.
Figures & tables
Figure 1: Framework of our LowBridge. Low-level features ( e.g. edge features) are treated as domain-invariant representations to train a generative model, G , to generate source-style images. Edge features extracted from unlabeled target images are then input to G , followed by a segmentation model, S , pretrained on the source data, to output the predictions.
Liver MR → CT
Liver CT → MR
Methods
Dice ↑
ASD ↓
Dice ↑
ASD ↓
Supervised training
94.3
1.0
91.9
0.7
W/o adaptation
50.5
25.4
42.2
17.1
SIFAv1 19 [ 12 ]
78.7
9.6
66.5
12.3
SIFAv2 20 [ 13 ]
79.2
9.4
67.8
11.6
DDFseg 21 [ 7 ]
78.4
10.2
–
–
Table 1: Quantitative comparison between our LowBridge and existing SOTA models on the CHAOS dataset for liver segmentation. Best results are in bold . “-” Indicates the data is not available.
Figure 2: Visual results on the CHAOS and MMWHS datasets.
Cardiac MR → CT
Methods
Dice ↑
ASD ↓
AA
LAC
LVC
MYO
Average
Average
Supervised training
95.1
93.0
90.1
89.4
91.9
1.6
W/o adaptation
49.3
57.3
59.3
63.7
57.4
12.8
SIFAv1 19 [ 12 ]
81.1
76.4
75.70
58.7
73.0
8.1
SIFAv2 20 [ 13 ]
81.3
79.5
73.8
61.6
74.1
7.0
Table 2: Quantitative comparison between our LowBridge and existing SOTA models on the MMWHS dataset for cardiac substructure segmentation. Best results are in bold .
Liver MR → CT
Liver CT → MR
No.
Gen.
Seg.
PSNR ↑
Dice ↑
ASD ↓
PSNR
Dice ↑
ASD ↓
1
FreUNet
SwinUNet
21.69
84.6
4.5
27.65
78.3
4.8
2
FreUNet
FreUNet
21.69
86.1
5.0
27.65
76.1
8.6
3
FreUNet
3D UNet
21.69
82.8
5.6
27.65
77.8
4.3
4
FreUNet
UNet
21.69
86.2
5.7
27.65
76.6
9.4
5
GAN
UNet
20.15
76.8
10.7
17.02
59.4
7.3
Table 3: Ablation Studies of LowBridge on the CHAOS Dataset.
Figure 3: Visual results of the ablation studies on the CHAOS dataset.
Deep learning-based medical image segmentation is increasingly used to support clinical diagnosis and develop new treatment strategies. However, model performance remains limited by the scarcity of high-quality annotated data and insufficient generalization across imaging protocols. This limitation is particularly evident in MRI and CT, where models are typically trained on a single acquisition sequence and exhibit reduced robustness when applied to unseen sequences or contrasts. Although data augmentation is widely used to improve general robustness on medical images, its impact on cross-modality generalization has not been quantitatively explored. In this work, we study a targeted set of data augmentation techniques designed to improve cross-modality transfer. We train three spine segmentation models, each on a single-modality/sequence dataset, and evaluate them across seven out-of-distribution datasets (spanning CT and MRI), reflecting a realistic single-sequence training and multi-sequence/contrast/modality deployment scenario. Our results demonstrate substantial performance gains on unseen domains (average Dice gain of 155 %) while preserving in-domain accuracy (average Dice decrease of 0.008 %), including effective transfer between CT and MRI. To mitigate the computational cost typically associated with strong data augmentation, we implement GPU-optimized augmentations that maintain, and even improve, training efficiency by approximately 10 %. We release our approach as an open-source toolbox, enabling seamless integration into commonly used frameworks such as nnUNet and MONAI. These augmentations significantly enhance robustness to heterogeneous clinical imaging scenarios without compromising training speed.
Nathan Molinier, Hendrik Möller, Thomas Dagonneau +6
NeuroPoly Lab, Institute of Biomedical Engineering, Polytechnique Montreal · Mila – Quebec AI Institute · Department for Interventional and Diagnostic Neuroradiology, TUM UniversityMay Hospital +3
Robust medical image segmentation across imaging modalities is challenging because of large differences in appearance and intensity distributions. Models trained on a single modality often show substantial performance drops when applied to unseen domains. In this work, we develop a unified 3D pancreas segmentation framework that applies domain-adversarial learning to 4,604 heterogeneous CT and MRI scans to learn anatomical representations. A shared nnU-Net encoder-decoder is trained for whole-pancreas segmentation, with a latent domain discriminator encouraging CT-MRI feature alignment. The learned encoder is subsequently transferred to pancreatic head-body-tail segmentation using limited MRI-only subregion annotations. An average Dice score of 87.31% on the in-distribution test set and Dice scores ranging from 84.20% to 88.09% across external OOD datasets were achieved in whole pancreas segmentation. Dice scores of 80.53% on MRI and 83.05% on CT were achieved for downstream subregion segmentation, without using CT subregion annotations. These results demonstrate that a unified anatomical representation can support both cross-modality pancreas segmentation and label-efficient downstream transfer.
Ziliang Hong, Hongyi Pan, Halil Ertugrul Aktas +7
Department of Radiology, Northwestern University, Chicago, IL, USA · Division of Gastroenterology and Hepatology, Mayo Clinic Florida, Jacksonville, FL, USA · Department of Gastroenterology and Hepatology, Northwestern University, Chicago, IL, USA
Magnetic resonance imaging comes in various modality contrasts that provide complementary anatomical and pathological information. Complete multimodal acquisitions are often unavailable due to time and protocol constraints. This leads to real-world datasets with missing modalities, where conventional medical image translation methods are typically limited to fixed source-target settings or require retraining for each observed source-target pair. We propose a unified framework that formulates missing-modality generation as a linear inverse problem under a joint distribution and solves it via posterior sampling with a flow matching model. By learning a joint prior over the complete modality set, our method can reconstruct arbitrary missing modalities at inference time by guiding the sampling trajectory to enforce measurement consistency with observed modalities. We further mitigate inter-modality error propagation in multi-target generation by adopting a many-to-one sampling strategy. Experiments on BraTS and IXI datasets show that our method achieves the best performance over baselines across most missing-modality scenarios. In downstream tumor segmentation, synthesized images from our method result in higher segmentation performance, indicating better preservation of clinically relevant structures.
Jonghun Kim
Department of Electrical and Computer Engineering, Sungkyunkwan University, Suwon, Korea