Organizations: Engineering Research Center of Cyberspace and School of Software and AI, Yunnan University, Kunming, China · School of Information Science and Technology, Yunnan Normal University, Kunming, China
Adversarial distillation transfers robustness from high-capacity teachers to compact students. Existing adversarial distillation methods mainly use teacher predictions on clean or adversarial examples to supervise student learning. However, teacher-favorable supervision within the perturbation neighborhood remains underexplored in adversarial distillation. We therefore propose Collaboratively Guided Adversarial Robust Distillation (CGARD), which jointly optimizes distinct student-adversarial and teacher-collaborative examples within the same perturbation neighborhood. The teacher-collaborative example is constrained to incur no greater cross-entropy loss under the teacher than the clean input. CGARD combines collaborative teacher guidance with adversarial teacher supervision to improve robust knowledge transfer. Experiments on CIFAR-10 and CIFAR-100, including white-box evaluation and additional black-box transfer evaluation, demonstrate consistent robustness improvements over strong adversarial distillation baselines.
Figures & tables
Figure 1: Overview of CGARD. The joint inner search optimizes student-adversarial and teacher-collaborative examples through the discrepancy between teacher and student predictions, while dual alignment uses teacher supervision from both examples.
CIFAR-10
CIFAR-100
Type
Arch.
Method
Clean
PGD
CW ∞
AA
Clean
PGD
CW ∞
AA
Teacher
WRN
ST [ 12 ]
84.75
57.83
55.49
54.33
59.40
33.71
30.30
29.19
Student
RN-18
Natural
95.28
0.01
0.01
0.00
77.72
0.00
0.00
0.00
PGD-AT [ 13 ]
82.50
51.53
50.57
48.04
57.52
28.50
27.12
24.68
TRADES [ 19 ]
82.78
50.83
49.37
47.63
56.32
26.95
23.77
22.69
ST [ 12 ]
83.43
53.72
51.66
50.01
55.75
30.92
26.89
26.04
Table 1: White-box robustness comparison (%) on CIFAR-10 and CIFAR-100 under ℓ∞ attacks with ε=8/255 . Best and second-best student results are shown in bold and underlined, respectively.
Surrogate Model
WRN
VGG-19
MN-V2
Method
PGD
CW ∞
PGD
CW ∞
PGD
CW ∞
PGD-AT [ 13 ]
65.21
65.21
63.84
63.12
63.25
62.41
TRADES [ 19 ]
65.64
65.26
63.89
62.96
63.00
62.26
ST [ 12 ]
66.92
66.54
66.41
65.70
66.10
65.49
ARD [ 6 ]
66.25
66.06
65.57
64.76
64.85
64.19
RSLAD [ 22 ]
66.70
66.66
66.28
65.69
65.82
65.16
Table 2: Black-box transfer robustness (%) on CIFAR-10 with robust WRN, VGG-19, and MN-V2 surrogates and an RN-18 target. Best and second-best results are shown in bold and underlined.
Configuration
Clean
FGSM
PGD
CW ∞
AA
Sequential Search
84.25
60.97
54.92
52.36
51.01
Joint Search (CGARD)
82.96
61.12
56.47
53.69
52.46
Table 3: Comparison of sequential search and constrained joint inner search on CIFAR-10 with WRN as the teacher and RN-18 as the student. Best results are shown in bold.
Lcol
Ladv
Clean
FGSM
PGD
CW ∞
AA
✓
82.64
60.65
56.06
53.21
51.84
✓
83.13
60.71
55.95
53.07
51.83
✓
✓
82.96
61.12
56.47
53.69
52.46
Table 4: Ablation study of the dual-alignment objective on CIFAR-10 with WRN as the teacher and RN-18 as the student. Best results are shown in bold.
Adversarial Distillation aims to enhance student robustness by guiding the student with a robust teacher's soft labels within the min-max adversarial training framework, yet its success is notoriously inconsistent: a more robust teacher often fails to improve, or even harms, the student's robust generalization. In this paper, we identify a key mechanism of this teacher dependency: the misalignment between the teacher's supervisory confidence and the student's representational limitations on a consistent subset of training data -- the Robustly Unlearnable Set. We present a theoretical framework analyzing the feature learning dynamics of a two-layer neural network, demonstrating that this mismatch creates a dichotomy in distillation outcomes. We prove that when a teacher provides confident supervision on unlearnable samples, it compels the student to memorize spurious noise patterns that eventually overpower the learned robust signal, thereby driving robust overfitting. Conversely, a teacher that exhibits high uncertainty on these samples effectively suppresses noise memorization, allowing the student to rely solely on the learnable signal for robust generalization. We empirically validate our theory across both synthetic simulations and real-image classification datasets, confirming that robust overfitting is driven by the teacher's interaction with unlearnable samples. Finally, we demonstrate that a teacher's predictive entropy on unlearnable samples serves as a strong indicator of student robustness, validating our theoretical framework and offering a principled guideline for robust teacher selection.
Hongsin Lee, Hye Won Chung
School of Electrical Engineering, KAIST, Daejeon, Korea.
Deep neural networks (DNNs) have achieved remarkable success in classical machine learning problems. However, they are known to be vulnerable to adversarial attacks. Countermeasures proposed in the literature, notably Information Bottleneck Distillation (IBD) introduced by Kuang et al., degrade the classification accuracy on clean inputs while improving the robustness to adversarial inputs. In this work, we extend the IBD framework by introducing an extra teacher model (clean teacher) trained with only clean inputs, into the distillation process from a robust teacher model trained by adversarial training. The features of both clean and robust teachers are transferred to the student through a cross-layer attention matrix. Experimental results on the CIFAR-10 and CIFAR-100 datasets show that the proposed method improves classification accuracy on clean samples compared to the original IBD, while maintaining similar accuracy on adversarial samples. Furthermore, our methods are competitive with state-of-the-art approaches, including the recent dual-teacher distillation framework B-MTARD, particularly in terms of the harmonic mean between clean and robust accuracy. We also analyze the impact of different training settings that have different influences on the attention module.
Vincent Ryusuke Takahashi, Yoshinari Takeishi, Jun'ichi Takeuchi +1
Joint Graduate School of Mathematics for Innovation Kyushu University · Faculty of Information Science and Electrical Engineering Kyushu University · Université Savoie Mont Blanc
Dataset distillation (DD) compresses a large training set into a small synthetic set for efficient training, but most DD methods optimize only clean accuracy and leave robustness uncontrolled. Recent robust DD methods improve robustness, yet they often suffer from a poor accuracy-robustness trade-off because they (i) treat all adversarially perturbed examples uniformly, despite robust risk being dominated by near-zero robust margins, and (ii) do not explicitly increase inter-class separation in the decision boundary where attacks concentrate. We present Contrastive Curriculum for Robust Dataset Distillation (C2R), a framework that couples an attack-aware curriculum with a contrastive robustness objective. From a robust-margin perspective, we derive a perturbation score that approximates each sample's robust hinge, enabling a curriculum that prioritizes the smallest-margin adversaries that most directly drive robust error. In parallel, a class-balanced contrastive robustness loss enforces adversarial invariance while explicitly widening boundary separation across classes. Experiments on CIFAR-10/100, Tiny-ImageNet, and multiple ImageNet-1K subsets under six attacks show that C2R achieves the best robust accuracy, outperforming prior robust DD by 2.8% on average.
Muquan Li, Yingyi Ma, Yihong Huang +5
The Laboratory of Intelligent Collaborative Computing of UESTC, Chengdu, China · Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ), Shenzhen, China · Monash University, Melbourne, Australia