Out-of-domain (OOD) robustness is challenging to achieve in real-world computer vision, especially in unsupervised domain adaptation scenarios, where shifts in image background, style, and acquisition instruments often degrade model performance. Generic augmentations show inconsistent gains under such shifts, whereas dataset-specific augmentations require expert knowledge and prior analysis. Moreover, prior studies show that neural networks adapt poorly to domain shifts because they exhibit a learning bias to domain-specific frequency components. Perturbing frequency values can mitigate such bias but overlooks pixel-level details, leading to suboptimal performance. To address these limitations, we propose D-GAP, a Dataset-agnostic and Gradient-guided augmentation method for the Amplitude spectrum (in frequency space) and the Pixel values. Unlike conventional handcrafted augmentations, D-GAP computes sensitivity maps in the frequency space from task gradients, which reflect how strongly the deep models respond to different frequency components, and uses the maps to adaptively interpolate amplitudes between source and target samples. We further propose a dual-space augmentation that jointly controls spectral bias and spatial fidelity by introducing a complementary pixel-space blending branch. This way, D-GAP turns augmentation from fixed, random, or manually designed perturbation into a model-response-adaptive intervention. Extensive experimental results show that the proposed method consistently outperforms both generic and dataset-specific domain adaptation methods, improving average OOD performance by +5.3% on four real-world datasets and +1.9% on three benchmark datasets. Code is available at https://github.com/RapidsAtHKUST/D-GAP.
Figures & tables
Figure 1 : Average amplitude sensitivity maps of ResNet50 trained across domains.
Figure 2 : Feature decomposition and augmentation examples. Top: For each dataset (iWildCam, Camelyon17, BirdCalls, Galaxy10), we show a source image, its corresponding augmented image generated by our method. Bottom: Based on the decomposition framework [ 49 ] , we annotate representative features in the datasets across xobj , xd:robust , xd:spu , xnoise . Our method effectively randomizes xd:spu , varies xd:robust while preserving or xobj .
Figure 3 : Amplitude-only and phase-only reconstructions of two images. The amplitudes of Image A and B are first mixed, and the mixed amplitude is combined with the phase of Image A to generate an augmentation.
Aug
α
β
γ
α/γ
β/γ
w/o
0.007
0.008
0.021
0.33
0.38
Ours
0.212
0.030
0.051
4.16
0.59
Table 1 : Connectivity values on iWildCam.
Figure 4 : Overview of the Gradient-guided Amplitude Mix procedure. Given a source image x1 and a target-domain image x2 , we compute the Sensitivity Map G(u,v) and the Mixing Map D(u,v) . The mixed amplitude is then combined with the phase of x1 .
Figure 5 : We plot the in-domain (ID) versus OOD performance across datasets (mean and std). ‘CL’ refers to ‘Connect Later’. ‘CP’ refers to ‘Copy Paste’. ‘SCJ’ refers to ‘Stain Color Jitter’.
Methods
PACS
OfficeHome
Art
Cartoon
Photo
Sketch
Avg.
Art
Clipart
Product
Real
Avg.
DeepAll
84.94
76.98
97.64
76.75
84.08
57.88
52.72
73.50
74.80
64.72
FACT (2023) [ 66 ]
89.68
81.06
97.02
83.76
87.88
60.34
54.85
74.48
76.55
66.56
SAM (2023) [ 65 ]
85.25
82.02
96.75
85.04
87.27
60.79
55.47
74.37
76.37
66.75
Su et al. (2024) [ 52 ]
79.88
78.03
95.93
73.15
81.75
64.13
54.83
75.60
75.36
67.71
AGST (2025) [ 62 ]
83.50
80.83
96.49
79.57
85.35
62.50
53.87
74.57
75.70
66.66
Table 2 : Model accuracy of leave-one-domain-out evaluation on PACS and OfficeHome.
Methods
MNIST
MNIST-M
SVHN
SYN
Avg.
DeepAll
95.8
58.8
61.7
78.6
73.7
FACT (2023) [ 66 ]
97.9
65.6
72.4
90.3
81.5
SAM (2023) [ 65 ]
97.8
66.8
73.2
90.7
82.1
Su et al. (2024) [ 52 ]
97.3
65.9
74.5
92.5
82.6
AGST (2025) [ 62 ]
97.3
63.1
70.2
93.2
81.0
Ours
98.3
68.5
77.2
93.8
84.5
Table 3 : Model accuracy of leave-one-domain-out evaluation on Digits-DG.
iWildCam (F1)
Camelyon17 (Acc)
BirdCall (F1)
Galaxy10 (Acc)
ID
OOD
ID
OOD
ID
OOD
ID
OOD
Ours w/o Augmentations
46.4
31.2
92.3
91.4
73.1
30.1
93.6
54.3
Pixel-only
45.4
29.1
87.5
69.8
75.2
32.4
95.5
68.2
Frequency-only
52.0
35.6
98.3
95.7
73.4
37.6
95.1
77.9
Mask-low
51.2
36.0
98.3
96.1
75.5
38.4
95.1
80.5
Mask-high
51.5
35.4
97.9
95.4
75.1
36.4
95.0
79.5
Table 4 : Ablation results across four real-world datasets.
Figure 6 : Performance of ConvNeXt and ViT on Galaxy10.
Method
PACS
Digits
OfficeHome
ConvNeXt
ViT
ConvNeXt
ViT
ConvNeXt
ViT
SAM (2023) [ 65 ]
85.3
85.4
88.0
85.1
73.7
63.8
Su et al. (2024) [ 52 ]
87.9
75.3
88.9
85.6
73.9
64.6
AGST (2025) [ 62 ]
87.4
79.0
88.5
85.8
73.7
64.4
Ours
88.7
86.4
89.8
86.8
75.0
66.0
Table 5 : Average OOD performance on PACS, Digits-DG, and OfficeHome with ConvNeXt and ViT.
Galaxy10
OfficeHome
High-gradient region
SC-CD
DC-SD
SC-CD
DC-SD
Top 10%
0.5449
0.5551
0.2774
0.3035
Top 20%
0.5485
0.5734
0.3332
0.3554
Top 30%
0.5497
0.5847
0.4372
0.4558
Table 6 : Jaccard similarity of top-ranked high-gradient frequency regions. SC-CD denotes same-class cross-domain pairs, while DC-SD denotes different-class same-domain pairs.
Camelyon17
iWildCam
Augmentations
α
β
γ
α/γ
β/γ
Augmentations
α
β
γ
α/γ
β/γ
w/o Augmentations
0.015
0.189
0.002
7.50
94.5
w/o Augmentations
0.007
0.008
0.021
0.33
0.38
RandAugment
0.181
0.186
0.114
1.59
1.63
RandAugment
0.224
0.125
0.130
1.72
0.96
Stain Color Jitter
0.029
0.116
0.004
7.25
29.0
Copy-Paste
0.254
0.031
0.063
4.03
0.49
Ours
0.120
0.073
0.003
40.0
24.3
Ours
0.212
0.030
0.051
4.16
0.59
Table 7 : Connectivity values across class, domain, and both on Camelyon17 and iWildCam datasets.
Method
Time / Iter. (ms)
Throughput (img/s)
Peak GPU Mem. (GB)
ERM
24.46±0.17
1308.3±9.1
0.86
MixUp
24.71±0.16
1294.9±8.1
0.88
DANN
48.62±0.24
658.2±3.2
1.55
LISA
66.28±0.19
482.8±1.3
0.88
SAM
48.00±0.30
666.6±4.2
0.91
FACT
63.00±0.31
507.9±2.5
1.94
Table 8 : Computational cost comparison under the same setting.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Methods
MNIST
MNIST-M
SVHN
SYN
Avg.
DeepAll
95.8
58.8
61.7
78.6
73.7
FACT (2023) [ 66 ]
97.9
65.6
72.4
90.3
81.5
SAM (2023) [ 65 ]
97.8
66.8
73.2
90.7
82.1
Su et al. (2024) [ 52 ]
97.3 ± 0.33
65.9 ± 0.51
74.5 ± 0.48
92.5 ± 0.31
82.6
AGST (2025) [ 62 ]
97.3 ± 0.14
63.1 ± 0.45
70.2 ± 0.78
93.2 ± 0.37
81.0
Ours
98.3 ± 0.12
68.5 ± 0.26
77.2 ± 0.53
93.8 ± 0.24
84.5
Appendix
Table 9 : Model accuracy of leave-one-domain-out evaluation on Digits-DG (mean and standard deviation).
Methods
Art
Clipart
Product
Real
Avg.
DeepAll
57.88
52.72
73.50
74.80
64.72
FACT (2023) [ 66 ]
60.34
54.85
74.48
76.55
66.56
SAM (2023) [ 65 ]
60.79
55.47
74.37
76.37
66.75
Su et al. (2024) [ 52 ]
64.13 ± 0.50
54.83 ± 0.49
75.60 ± 0.61
75.36 ± 1.04
67.71
AGST (2025) [ 62 ]
62.50 ± 0.61
53.87 ± 0.29
74.57 ± 0.48
75.70 ± 0.78
66.66
Ours
64.82 ± 0.55
55.21 ± 0.37
77.76 ± 0.43
83.08 ± 0.25
70.22
Appendix
Table 10 : Model accuracy of leave-one-domain-out evaluation on OfficeHome (mean and std).
Methods
Art
Cartoon
Photo
Sketch
Avg.
DeepAll
84.94
76.98
97.64
76.75
84.08
FACT (2023) [ 66 ]
89.68
81.06
97.02
83.76
87.88
SAM (2023) [ 65 ]
85.25
82.02
96.75
85.04
87.27
Su et al. (2024) [ 52 ]
79.88 ± 0.53
78.03 ± 0.57
95.93 ± 0.21
73.15 ± 0.39
81.75
AGST (2025) [ 62 ]
83.50 ± 0.76
80.83 ± 0.48
96.49 ± 0.23
79.57 ± 0.62
85.35
Ours
89.34 ± 0.34
82.81 ± 0.63
97.99 ± 0.14
85.98 ± 0.19
89.03
Appendix
Table 11 : Model accuracy of leave-one-domain-out evaluation on PACS (mean and std).
Figure 7 : Effect of the frequency-window side-length ratio r on Galaxy10 and BirdCalls datasets.
Figure 8 : Effect of the mixing ratios λ1 and λ2 on Galaxy10 and BirdCalls datasets.
Method
Hyperparameters
ERM
Learning rate =10−4 Weight decay =0
LISA
Learning rate =10−4 Weight decay =0 Transform probability =0.9
MixUp
Learning rate =10−4 Weight decay =0 Transform probability =0.9α=0.4
CutMix
Learning rate =10−4 Weight decay =0 Transform probability =0.9α=0.4
Cutout
Learning rate =10−4 Weight decay =0 Transform probability =0.9α=0.4
Randaugment
Learning rate =10−4 Transform probability =0.9 Version = Original
Appendix
Table 12: Hyperparameter settings for baseline methods on Galaxy10.