Out-of-domain (OOD) robustness is challenging to achieve in real-world computer vision, especially in unsupervised domain adaptation scenarios, where shifts in image background, style, and acquisition instruments often degrade model performance. Generic augmentations show inconsistent gains under such shifts, whereas dataset-specific augmentations require expert knowledge and prior analysis. Moreover, prior studies show that neural networks adapt poorly to domain shifts because they exhibit a learning bias to domain-specific frequency components. Perturbing frequency values can mitigate such bias but overlooks pixel-level details, leading to suboptimal performance. To address these limitations, we propose D-GAP, a Dataset-agnostic and Gradient-guided augmentation method for the Amplitude spectrum (in frequency space) and the Pixel values. Unlike conventional handcrafted augmentations, D-GAP computes sensitivity maps in the frequency space from task gradients, which reflect how strongly the deep models respond to different frequency components, and uses the maps to adaptively interpolate amplitudes between source and target samples. We further propose a dual-space augmentation that jointly controls spectral bias and spatial fidelity by introducing a complementary pixel-space blending branch. This way, D-GAP turns augmentation from fixed, random, or manually designed perturbation into a model-response-adaptive intervention. Extensive experimental results show that the proposed method consistently outperforms both generic and dataset-specific domain adaptation methods, improving average OOD performance by +5.3% on four real-world datasets and +1.9% on three benchmark datasets. Code is available at https://github.com/RapidsAtHKUST/D-GAP.
Figures & tables
Figure 1 : Average amplitude sensitivity maps of ResNet50 trained across domains.
Figure 2 : Feature decomposition and augmentation examples. Top: For each dataset (iWildCam, Camelyon17, BirdCalls, Galaxy10), we show a source image, its corresponding augmented image generated by our method. Bottom: Based on the decomposition framework [ 49 ] , we annotate representative features in the datasets across xobj , xd:robust , xd:spu , xnoise . Our method effectively randomizes xd:spu , varies xd:robust while preserving or xobj .
Figure 3 : Amplitude-only and phase-only reconstructions of two images. The amplitudes of Image A and B are first mixed, and the mixed amplitude is combined with the phase of Image A to generate an augmentation.
Aug
α
β
γ
α/γ
β/γ
w/o
0.007
0.008
0.021
0.33
0.38
Ours
0.212
0.030
0.051
4.16
0.59
Table 1 : Connectivity values on iWildCam.
Figure 4 : Overview of the Gradient-guided Amplitude Mix procedure. Given a source image x1 and a target-domain image x2 , we compute the Sensitivity Map G(u,v) and the Mixing Map D(u,v) . The mixed amplitude is then combined with the phase of x1 .
Figure 5 : We plot the in-domain (ID) versus OOD performance across datasets (mean and std). ‘CL’ refers to ‘Connect Later’. ‘CP’ refers to ‘Copy Paste’. ‘SCJ’ refers to ‘Stain Color Jitter’.
Methods
PACS
OfficeHome
Art
Cartoon
Photo
Sketch
Avg.
Art
Clipart
Product
Real
Avg.
DeepAll
84.94
76.98
97.64
76.75
84.08
57.88
52.72
73.50
74.80
64.72
FACT (2023) [ 66 ]
89.68
81.06
97.02
83.76
87.88
60.34
54.85
74.48
76.55
66.56
SAM (2023) [ 65 ]
85.25
82.02
96.75
85.04
87.27
60.79
55.47
74.37
76.37
66.75
Su et al. (2024) [ 52 ]
79.88
78.03
95.93
73.15
81.75
64.13
54.83
75.60
75.36
67.71
AGST (2025) [ 62 ]
83.50
80.83
96.49
79.57
85.35
62.50
53.87
74.57
75.70
66.66
Table 2 : Model accuracy of leave-one-domain-out evaluation on PACS and OfficeHome.
Methods
MNIST
MNIST-M
SVHN
SYN
Avg.
DeepAll
95.8
58.8
61.7
78.6
73.7
FACT (2023) [ 66 ]
97.9
65.6
72.4
90.3
81.5
SAM (2023) [ 65 ]
97.8
66.8
73.2
90.7
82.1
Su et al. (2024) [ 52 ]
97.3
65.9
74.5
92.5
82.6
AGST (2025) [ 62 ]
97.3
63.1
70.2
93.2
81.0
Ours
98.3
68.5
77.2
93.8
84.5
Table 3 : Model accuracy of leave-one-domain-out evaluation on Digits-DG.
iWildCam (F1)
Camelyon17 (Acc)
BirdCall (F1)
Galaxy10 (Acc)
ID
OOD
ID
OOD
ID
OOD
ID
OOD
Ours w/o Augmentations
46.4
31.2
92.3
91.4
73.1
30.1
93.6
54.3
Pixel-only
45.4
29.1
87.5
69.8
75.2
32.4
95.5
68.2
Frequency-only
52.0
35.6
98.3
95.7
73.4
37.6
95.1
77.9
Mask-low
51.2
36.0
98.3
96.1
75.5
38.4
95.1
80.5
Mask-high
51.5
35.4
97.9
95.4
75.1
36.4
95.0
79.5
Table 4 : Ablation results across four real-world datasets.
Figure 6 : Performance of ConvNeXt and ViT on Galaxy10.
Method
PACS
Digits
OfficeHome
ConvNeXt
ViT
ConvNeXt
ViT
ConvNeXt
ViT
SAM (2023) [ 65 ]
85.3
85.4
88.0
85.1
73.7
63.8
Su et al. (2024) [ 52 ]
87.9
75.3
88.9
85.6
73.9
64.6
AGST (2025) [ 62 ]
87.4
79.0
88.5
85.8
73.7
64.4
Ours
88.7
86.4
89.8
86.8
75.0
66.0
Table 5 : Average OOD performance on PACS, Digits-DG, and OfficeHome with ConvNeXt and ViT.
Galaxy10
OfficeHome
High-gradient region
SC-CD
DC-SD
SC-CD
DC-SD
Top 10%
0.5449
0.5551
0.2774
0.3035
Top 20%
0.5485
0.5734
0.3332
0.3554
Top 30%
0.5497
0.5847
0.4372
0.4558
Table 6 : Jaccard similarity of top-ranked high-gradient frequency regions. SC-CD denotes same-class cross-domain pairs, while DC-SD denotes different-class same-domain pairs.
Camelyon17
iWildCam
Augmentations
α
β
γ
α/γ
β/γ
Augmentations
α
β
γ
α/γ
β/γ
w/o Augmentations
0.015
0.189
0.002
7.50
94.5
w/o Augmentations
0.007
0.008
0.021
0.33
0.38
RandAugment
0.181
0.186
0.114
1.59
1.63
RandAugment
0.224
0.125
0.130
1.72
0.96
Stain Color Jitter
0.029
0.116
0.004
7.25
29.0
Copy-Paste
0.254
0.031
0.063
4.03
0.49
Ours
0.120
0.073
0.003
40.0
24.3
Ours
0.212
0.030
0.051
4.16
0.59
Table 7 : Connectivity values across class, domain, and both on Camelyon17 and iWildCam datasets.
Method
Time / Iter. (ms)
Throughput (img/s)
Peak GPU Mem. (GB)
ERM
24.46±0.17
1308.3±9.1
0.86
MixUp
24.71±0.16
1294.9±8.1
0.88
DANN
48.62±0.24
658.2±3.2
1.55
LISA
66.28±0.19
482.8±1.3
0.88
SAM
48.00±0.30
666.6±4.2
0.91
FACT
63.00±0.31
507.9±2.5
1.94
Table 8 : Computational cost comparison under the same setting.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Methods
MNIST
MNIST-M
SVHN
SYN
Avg.
DeepAll
95.8
58.8
61.7
78.6
73.7
FACT (2023) [ 66 ]
97.9
65.6
72.4
90.3
81.5
SAM (2023) [ 65 ]
97.8
66.8
73.2
90.7
82.1
Su et al. (2024) [ 52 ]
97.3 ± 0.33
65.9 ± 0.51
74.5 ± 0.48
92.5 ± 0.31
82.6
AGST (2025) [ 62 ]
97.3 ± 0.14
63.1 ± 0.45
70.2 ± 0.78
93.2 ± 0.37
81.0
Ours
98.3 ± 0.12
68.5 ± 0.26
77.2 ± 0.53
93.8 ± 0.24
84.5
Appendix
Table 9 : Model accuracy of leave-one-domain-out evaluation on Digits-DG (mean and standard deviation).
Methods
Art
Clipart
Product
Real
Avg.
DeepAll
57.88
52.72
73.50
74.80
64.72
FACT (2023) [ 66 ]
60.34
54.85
74.48
76.55
66.56
SAM (2023) [ 65 ]
60.79
55.47
74.37
76.37
66.75
Su et al. (2024) [ 52 ]
64.13 ± 0.50
54.83 ± 0.49
75.60 ± 0.61
75.36 ± 1.04
67.71
AGST (2025) [ 62 ]
62.50 ± 0.61
53.87 ± 0.29
74.57 ± 0.48
75.70 ± 0.78
66.66
Ours
64.82 ± 0.55
55.21 ± 0.37
77.76 ± 0.43
83.08 ± 0.25
70.22
Appendix
Table 10 : Model accuracy of leave-one-domain-out evaluation on OfficeHome (mean and std).
Methods
Art
Cartoon
Photo
Sketch
Avg.
DeepAll
84.94
76.98
97.64
76.75
84.08
FACT (2023) [ 66 ]
89.68
81.06
97.02
83.76
87.88
SAM (2023) [ 65 ]
85.25
82.02
96.75
85.04
87.27
Su et al. (2024) [ 52 ]
79.88 ± 0.53
78.03 ± 0.57
95.93 ± 0.21
73.15 ± 0.39
81.75
AGST (2025) [ 62 ]
83.50 ± 0.76
80.83 ± 0.48
96.49 ± 0.23
79.57 ± 0.62
85.35
Ours
89.34 ± 0.34
82.81 ± 0.63
97.99 ± 0.14
85.98 ± 0.19
89.03
Appendix
Table 11 : Model accuracy of leave-one-domain-out evaluation on PACS (mean and std).
Figure 7 : Effect of the frequency-window side-length ratio r on Galaxy10 and BirdCalls datasets.
Figure 8 : Effect of the mixing ratios λ1 and λ2 on Galaxy10 and BirdCalls datasets.
Method
Hyperparameters
ERM
Learning rate =10−4 Weight decay =0
LISA
Learning rate =10−4 Weight decay =0 Transform probability =0.9
MixUp
Learning rate =10−4 Weight decay =0 Transform probability =0.9α=0.4
CutMix
Learning rate =10−4 Weight decay =0 Transform probability =0.9α=0.4
Cutout
Learning rate =10−4 Weight decay =0 Transform probability =0.9α=0.4
Randaugment
Learning rate =10−4 Transform probability =0.9 Version = Original
Appendix
Table 12: Hyperparameter settings for baseline methods on Galaxy10.
Domain generalization (DG) and neural network pruning are conventionally treated as distinct objectives, targeting out-of-distribution (OOD) robustness and model efficiency, respectively. In this work, we bridge this gap by introducing Domain-Aware Pruning (DAP), a framework that leverages network sparsity as a mechanism to implicitly enhance generalization to unseen domains. Diverging from standard binary mask optimization, DAP learns a continuous parameter retention probability p∈[0,1], framing network compression as a continuous probabilistic masking problem. By introducing a regularization objective that actively penalizes the retention of domain-sensitive weights during the mask training, DAP identifies a domain-invariant subnetwork. Empirical results across five DG benchmark datasets demonstrate that DAP achieves significant sparsity while consistently matching or exceeding the OOD performance of its dense counterparts. Crucially, DAP is an algorithm-agnostic framework that integrates seamlessly with existing DG pipelines without necessitating post-hoc fine-tuning. Beyond efficiency and generalization, we show that DAP natively provides increased robustness to adversarial perturbations and yields highly interpretable models, where the retained weights reliably encapsulate the most domain-invariant and task-critical representations.
Domain Generalization (DG) aims to learn representations that remain robust under out-of-distribution (OOD) shifts and generalize effectively to unseen target domains. While recent invariant learning strategies and architectural advances have achieved strong performance, explicitly discovering a structured domain-invariant subspace through second-order statistics remains underexplored. In this work, we propose CPCANet, a novel framework grounded in Common Principal Component Analysis (CPCA), which unrolls the iterative Flury-Gautschi (FG) algorithm into fully differentiable neural layers. This approach integrates the statistical properties of CPCA into an end-to-end trainable framework, enforcing the discovery of a shared subspace across diverse domains while preserving interpretability. Experiments on four standard DG benchmarks demonstrate that CPCANet achieves state-of-the-art (SOTA) performance in zero-shot transfer. Moreover, CPCANet is architecture-agnostic and requires no dataset-specific tuning, providing a simple and efficient approach to learning robust representations under distribution shift. Code is available at https://github.com/wish44165/CPCANet.
Dataset Distillation (DD) synthesizes a compact synthetic dataset that preserves the training utility of a full dataset. However, its standard formulation assumes that test data follow the same distribution as training data, an assumption that rarely holds in practice. A straightforward extension-applying post-hoc Domain Generalization (DG) techniques to distilled data-is ill-suited because existing DG methods rely on the natural diversity of real datasets, which compact synthetic sets inherently lack, while also incurring substantial augmentation overhead that conflicts with the efficiency objective of dataset distillation. To address this limitation, we introduce Domain Generalizable Dataset Distillation (DGDD), a new problem setting that explicitly targets out-of-distribution (OOD) generalization of distilled datasets. We study this problem through a widely adopted DD baseline of Distribution Matching (DM). We attribute the OOD vulnerability of DM to the entanglement of class-discriminative and domain-specific information within the compressed synthetic set, and propose Spectral Gradient Surgery (SGS) to disentangle the two. The key insight of SGS is that cross-domain agreement among domain-wise gradients in the spectral domain reveals which gradient components are shared across source domains-and are therefore class-discriminative-and which are domain-specific. Based on this observation, SGS augments the standard DM update with two complementary gradients: one that reinforces cross-domain shared components and another that explicitly promotes diversity within the distilled dataset. Extensive experiments on diverse-scale benchmarks demonstrate that SGS substantially improves OOD generalization while remaining plug-and-play compatible with existing DM methods.
Minyoung Oh, Najeong Chae, Jae-Young Sim
Graduate School of Artificial Intelligence Ulsan National Institute of Science and Technology (UNIST) Ulsan, Republic of Korea