Improving Mixup Calibration with Wasserstein Distributionally Robust Optimization
Authors: Jiaming Hu, Yeping Jin, Debarghya Mukherjee, Ioannis Ch. Paschalidis
Organizations: Department of Mathematics & Statistics Boston University · Department of Electrical and Computer Engineering, Division of Systems Engineering Department of Biomedical Engineering Faculty of Computing & Data Sciences Boston University
In many real-world applications, ensuring the robustness and stability of deep neural networks (DNNs) is crucial, particularly for image classification tasks that encounter various input perturbations. While Mixup-based data augmentation techniques have been widely adopted to enhance the resilience of trained models against such perturbations, our experiments reveal an important corruption robustness-calibration trade-off: stronger Mixup-based augmentation can improve robustness against corrupted data while substantially increasing expected calibration error (ECE). To address this challenge, we introduce DRO-Augment, a framework that integrates Wasserstein Distributionally Robust Optimization (W-DRO) with various Mixup-based data augmentation strategies to mitigate this trade-off. Our method substantially reduces ECE under strong Mixup-based augmentation while largely preserving corruption accuracy across CIFAR-10, CIFAR-100, CIFAR-10-C, and CIFAR-100-C. On the theoretical side, we establish novel generalization error bounds for neural networks trained using a variation-regularized loss function with augmented data, closely related to the W-DRO problem. Furthermore, we introduce a refined CIFAR-C benchmark that corrects inconsistencies in corruption intensities, providing a more reliable evaluation for future robustness research.
Figures & tables
Method
CIFAR-10
CIFAR-100
Clean
-C
Clean
-C
Acc. ↑
ECE ↓
Acc. ↑
ECE ↓
Acc. ↑
ECE ↓
Acc. ↑
ECE ↓
Baseline
94.74
3.34
75.45
17.56
77.99
10.07
49.40
27.01
Mixup
95.60
14.03
79.25
12.44
79.12
14.92
53.29
11.64
Mixup + DRO
96.04
7.54
80.75
7.25
80.25
8.95
54.24
9.80
Manifold Mixup
95.83
17.44
78.71
15.11
80.19
12.83
53.31
12.73
Table 1: Average accuracy (%) and ECE (%) of PreActResNet-18 on CIFAR and CIFAR-C benchmarks. Higher accuracy is better and lower ECE is better.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Method
CIFAR-10
CIFAR-100
Clean
-C
Clean
-C
Acc. ↑
ECE ↓
Acc. ↑
ECE ↓
Acc. ↑
ECE ↓
Acc. ↑
ECE ↓
Baseline
94.86
3.24
74.37
17.95
75.95
11.67
44.51
33.58
Mixup
94.72
20.95
77.96
16.69
75.17
16.58
48.72
13.05
Mixup + DRO
95.77
12.12
78.72
9.67
77.36
5.78
49.71
12.62
Manifold Mixup
95.81
12.12
75.38
11.41
77.79
8.36
48.84
13.12
Appendix
Table 2: Average accuracy (%) and ECE (%) of WideResNet-28 × 2 on CIFAR and CIFAR-C benchmarks. Higher accuracy is better and lower ECE is better.
Method
ρ
CIFAR - 10
CIFAR - 100
Clean
- C
Clean
- C
Acc. ↑
ECE ↓
Acc. ↑
ECE ↓
Acc. ↑
ECE ↓
Acc. ↑
ECE ↓
Mixup + DRO
1.0×10−4
96.06
6.30
80.08
7.20
80.15
5.77
54.41
9.24
2.0×10−4
96.09
6.03
79.52
6.84
80.17
7.57
54.43
10.54
3.0×10−4
95.86
6.46
79.74
7.92
80.38
9.81
54.44
9.73
5.0×10−4
96.13
6.18
80.04
6.79
80.07
8.08
54.31
9.13
Appendix
Table 3: Sensitivity study of the DRO radius ρ for Mixup + DRO, NoisyMix + DRO, and Manifold Mixup + DRO on CIFAR and CIFAR-C benchmarks using PreActResNet-18. Metrics follow Table 1 : higher accuracy is better and lower ECE is better.
Method
αmix
CIFAR - 10
CIFAR - 100
Clean
- C
Clean
- C
Acc. ↑
ECE ↓
Acc. ↑
ECE ↓
Acc. ↑
ECE ↓
Acc. ↑
ECE ↓
Mixup [+DRO]
0.2
95.69 [95.75]
2.52 [1.12]
76.97 [77.99]
10.14 [10.98]
79.02 [79.33]
3.53 [1.40]
52.17 [51.96]
8.46 [10.69]
0.4
95.84 [95.97]
11.38 [2.88]
78.45 [79.46]
10.89 [7.87]
78.66 [79.62]
12.62 [6.00]
51.71 [52.15]
10.90 [9.37]
0.6
95.60 [96.15]
11.32 [4.12]
78.66 [80.29]
10.44 [6.56]
79.54 [80.02]
12.80 [7.28]
53.03 [53.44]
10.16 [8.85]
1.0
95.69 [96.01]
14.70 [5.95]
79.16 [79.69]
12.57 [7.47]
79.69 [80.31]
14.96 [8.32]
52.83 [55.00]
11.69 [9.21]
Appendix
Table 4: Sensitivity study of the mixing-strength parameter αmix on CIFAR and CIFAR-C benchmarks using PreActResNet-18. Higher accuracy is better and lower ECE is better. Bracketed values report the corresponding method with DRO.
Department of Computer Science and Engineering, Southern University of Science and Technology · The University of Sydney · JD Explore Academy, JD.com Inc +1