Adversarial training is one of the most reliable defenses against adversarial attacks, but its high computational cost must generally be paid anew for each task. Robust foundation models offer a promising alternative: adversarially pretrain a model once and then transfer its robustness to downstream tasks through lightweight adaptation. However, a fundamental question remains open: can robustness acquired during pretraining transfer to unseen tasks without further adversarial training? In this study, we answer this question affirmatively. A single model adversarially pretrained at scale can achieve optimal robustness on new tasks without additional task-specific training. Specifically, we show that, for a family of Gaussian-mixture classification tasks, a sufficiently deep linear transformer adversarially trained across tasks can asymptotically attain the robust Bayes error on previously unseen tasks through in-context learning from clean demonstrations. By contrast, a standardly trained model cannot. We further analyze convergence under gradient flow, an accuracy--robustness trade-off, and demonstration complexity.
Figures & tables
Figure 1: Robust zero-one errors for d=100 , r=10 , and ϵ=0.08 . Means and sample standard deviations are shown over three runs in the left panel and five runs in the other panels.
Training
CIFAR-100
Tiny ImageNet
Caltech-256
Standard
79.1 ± 2.0 / 26.8 ± 3.4
77.4 ± 3.4 / 27.6 ± 7.6
78.0 ± 3.6 / 23.8 ± 6.5
Adversarial
77.9 ± 1.7 / 63.2 ± 1.8
67.0 ± 6.0 / 52.8 ± 2.2
74.6 ± 3.9 / 62.4 ± 4.6
Table 1: Clean/robust accuracy (%) of six-layer transformers with softmax attention on real-world data, reported as mean ± sample standard deviation across five independent runs.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: Cosine similarity between the population-context effective direction of an adversarially trained transformer and the robust Bayes direction. The distribution and the three experimental conditions are the same as in Fig. 1 . The left panel varies the depth at T=40,000 and N=100,000 , the center panel varies T with L=10 and N=1,000 , and the right panel varies N with L=10 and T=1,000 . Means and sample standard deviations are shown over three runs in the left panel and five runs in the other panels.
Standard training
Adversarial training
N
Clean
Robust
Clean
Robust
10
84.6±4.2
63.5±4.5
84.8±0.7
67.8±0.3
50
95.1±4.3
52.8±20.6
92.4±0.3
75.5±0.1
100
98.0±0.6
45.1±15.9
93.6±0.3
76.6±0.1
500
98.7±0.2
39.6±7.0
94.1±0.3
77.5±0.1
1,000
98.4±0.6
45.1±13.0
94.0±0.1
77.6±0.1
Appendix
Table A1: Clean and robust accuracies (%), reported as mean ± sample standard deviation across five runs, for depth- 10 models trained on T=1,000 tasks. Each row uses the displayed number of demonstrations in both pretraining and evaluation. Robust accuracy uses ϵ=0.08 , and every model is evaluated on 40,000 test tasks with 200 queries per task.