Organizations: School of Computer Science and Engineering, Central South University, Changsha, China · School of Software Engineering, Xi’an Jiaotong University, Xi’an, Shaanxi, China · College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China
Domain Generalization (DG) for medical image segmentation is both highly challenging and critically important. However, existing medical DG methods largely overlook the issue of Catastrophic Forgetting (CF): \textbf{Models often sacrifice their ability to retain source-domain knowledge while pursuing cross-domain robustness.} This can directly threaten diagnostic safety in already-deployed clinical scenarios. To address this, we investigate data augmentation strategies and catastrophic forgetting for medical image DG segmentation. First, we propose a structure-guided style diffusion augmentation method. Constrained by anatomical structure consistency in the frequency domain, this method performs cross-domain diffusion on the amplitude spectrum, generating samples with more diverse and broader style coverage to better support domain generalization. Then, we design a collaborative learning network with a dual-branch interactive architecture (CoDG-Net), together with a novel learning bias-guided strategy that adaptively regulates knowledge transfer at both the layer level and the task level, thereby effectively mitigating catastrophic forgetting on the source domain. Experiments and ablation studies on single-source and multi-source medical DG benchmark datasets demonstrate that CoDG-Net not only outperforms existing state-of-the-art methods in target-domain segmentation performance, but also achieves a lower forgetting rate on the source-domain data. The code is available at: https://github.com/wangprocess/CoDG-Net.
Figures & tables
Figure 1: Conceptual overview of our research motivation. (a) Original amplitude and phase spectra of fundus images acquired from different devices. (b) Coverage of extended domains after applying traditional data augmentation and diffusion-based generation methods on source domain 1. (c) Comparison of distribution-fitting capability on the source domain.
Figure 2: Overall workflow of the proposed method. First, pseudo-domain style exploration is applied to obtain DP . Second, a diffusion-based style diversification method is used to generate a style-diverse DG . Finally, the generated data are fed into the proposed CoDG-Net for training.
Figure 3: Overall architecture of the proposed CoDG-Net and the bias-guided learning strategies (for clarity, skip connections within each U-Net branch are omitted).
Table 1: Comparative results with other related methods on the BraTS dataset (including using T2 as the source domain and using T1CE as the source domain. Red indicates optimal, and underlining indicates suboptimal.)
Method
Target Domain: Domain1
Target Domain: Domain2
Target Domain: Domain3
Target Domain: Domain4
Disc
Cup
Disc
Cup
Disc
Cup
Disc
Cup
Dice
ASD
Dice
ASD
Dice
ASD
Dice
ASD
Dice
ASD
Dice
ASD
Dice
ASD
Dice
ASD
No Generalization
82.80
19.03
64.55
26.20
85.99
19.98
76.03
20.31
88.78
16.63
84.14
10.35
85.43
9.93
68.85
25.11
RAM-DSIR Zhou et al. (2022a)
95.75
7.12
85.48
16.05
89.43
13.86
78.82
14.01
94.67
7.11
87.44
9.02
94.10
7.06
85.84
8.29
DCAC Hu et al. (2022)
95.54
6.35
81.43
19.20
87.85
18.28
77.72
17.15
94.28
8.11
86.80
9.14
95.40
5.20
87.68
7.12
DFQ Bi et al. (2024)
96.50
6.01
87.30
15.72
92.52
12.09
81.92
13.05
95.04
7.05
88.95
7.70
94.85
5.84
87.47
6.55
Table 2: Comparison of domain generalization results on Optic Disc/Cup segmentation from fundus images (We follow the practice in domain generalization literature to adopt the leave-one-domain-out strategy. Red indicates optimal, and underlining indicates suboptimal.)
Method
Source Domain: T2
Source Domain: T1CE
Dice ↑
SFM (Dice) ↓
Dice ↑
SFM (Dice) ↓
No Generalization
84.52
–
78.71
–
Fed-DG
68.68
15.84
59.32
19.39
SADN
73.71
10.81
55.15
23.56
EGSDG
75.93
8.59
63.67
15.04
SLAug
73.53
10.99
60.61
18.1
Table 3: Comparison of source-domain performance and forgetting rate on the BraTS dataset.
Figure 4: Visual comparison of model predictions between the T2 source domain and the T1ce target domain.
Method
SE
SD
DBS
LW
TW
T2(Source)
T2-other(Avg)
T1CE(Source)
T1CE-other(Avg)
Dice ↑
SFM ↓
Dice ↑
Dice ↑
SFM ↓
Dice ↑
VA1
×
×
×
×
×
84.52
–
29.69
78.71
–
40.85
VA2
✓
×
×
×
×
80.56
3.96
51.82
74.98
3.73
44.30
VA3
×
✓
×
×
×
83.16
1.36
33.35
76.42
2.29
44.13
VA4
✓
✓
×
×
×
79.33
5.19
67.59
72.83
5.88
61.71
VA5
✓
✓
✓
×
×
81.09
3.43
52.12
75.25
3.46
56.82
Table 4: Ablation study results on the BraTS dataset.
Figure 5: Activation frequency of collaborative interactions at each intermediate layer in CoDG-Net.
Figure 6: Hyper-parameter sensitivity analysis of the random perturbation probability δ .
Reliable clinical deployment of deep medical image models is hindered by distribution shifts across scanners, sites, and acquisition protocols. Existing domain generalization (DG) methods often focus on style or intensity diversification, but they can still leave networks dependent on domain-specific texture correlations. Inspired by evidence that Fourier phase encodes semantic structure, we introduce PhaseAT, a phase-aware adversarial training framework for medical DG. PhaseAT forms phase-perturbed training views in the Fourier domain by iteratively updating a bounded phase perturbation while keeping the amplitude spectrum unchanged, thereby stressing spatial organization under matched appearance statistics. Perturbations are applied only to the luminance channel in YCbCr color space to avoid chromatic artifacts. Additionally, a simple phase-saliency mask concentrates updates on the most influential frequencies. The model is trained with a weighted combination of losses on clean and phase-perturbed samples, supporting both single-source and multi-source DG. We validate our method on two challenging medical datasets and demonstrate that PhaseAT achieves over 20% improvement in single-source domain generalization, outperforming several state-of-the-art DG methods. The code implementation is available at: https://github.com/ahmed-sharshar/PhaseAT.
Ahmed Sharshar, Asif Hanif, Naveen Kumar Kummari +2
Division of Computing and Mathematical Sciences Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE
Medical image segmentation is a critical task in computer-aided diagnosis and treatment planning. However, deep learning models often struggle to generalize across datasets due to domain shifts arising from variations in imaging protocols, scanner types, and patient populations. Traditional domain generalization (DG) methods utilize causal feature learning, adversarial consistency, and style augmentation to improve segmentation robustness. While effective, these approaches rely on explicit feature alignment, adversarial objectives, or handcrafted augmentations, which may not fully exploit the capabilities of foundation models. Recently, the Segment Anything Model (SAM) has demonstrated strong generalization capabilities in segmentation tasks. SAM-based DG methods attempt to improve medical image segmentation. However, these approaches primarily operate in the spatial domain and overlook frequency-based discrepancies that significantly affect model robustness. In this work, we propose Frequency-based Domain Generalization with SAM (FSAM), a novel framework that integrates Low-Rank Adaptation (LoRA) for efficient fine-tuning and a frequency adapter to incorporate frequency-domain representations for single-source domain generalization. FSAM enhances SAM's segmentation robustness by extracting domain-invariant high-frequency features, mitigating frequency-related domain shifts. Experimental results on fundus and prostate datasets demonstrate that FSAM outperforms existing traditional DG and SAM-based DG approaches in domain generalization. Codes and pre-trained models will be made available on GitHub.
Hardware shifts, color variations, and changing patient characteristics between development and deployment routinely break trained medical image classifiers. Existing remedies fall short: standard color jittering provides insufficient diversity, while deep generative style transfer algorithms hallucinate features, destroy clinically relevant structures, and waste massive compute resources. To address this, we revisit classical statistical color matching and repurpose it as Colorist, a highly efficient data augmentation strategy that applies global mean-standard deviation matching directly in the RGB color space. We demonstrate that this training-free, fully interpretable approach safely generates structurally intact domain variations, outperforming deep generative models in structural fidelity and color alignment. Across out-of-distribution histopathology, peripheral blood, dermatology, and retinal datasets, it improves balanced accuracy by up to +9% over state-of-the-art domain generalization regularizers and by +13% over an unaugmented baseline. Moreover, by avoiding neural networks in the augmentation loop, Colorist preserves anatomical structure, minimizes carbon footprint, and integrates seamlessly into standard dataloaders. Together, these findings establish statistical matching as a safe, interpretable, yet overlooked alternative to deep architectures for clinical robustness. Source code is available at https://github.com/sdoerrich97/colorist.
Sebastian Doerrich, Francesco Di Salvo, Shyam Nandan Rai +2
xAILab Bamberg, University of Bamberg, Bamberg, Germany