cs.LGSep 30, 2026

Probabilistic Adversarial Training

Authors: Andi Zhang, Xingyu Zhao, Siddartha Khastgir

Organizations: WMG, University of Warwick · Wuhan University

Abstract

Building on a probabilistic perspective in which adversarial examples arise from the overlap between a distance-based distribution pdisp_{\mathrm{dis}} and a victim-classifier-induced distribution pvicp_{\mathrm{vic}}, we start from a simple intuition: adversarial examples become harder to generate when these two distributions are pushed apart, as their overlap becomes smaller, thereby increasing robustness. This intuition naturally motivates a KL-based robustness objective. We then prove that KL(pdis∥pvic)−log⁡Zvic\mathrm{KL}(p_{\mathrm{dis}}\|p_{\mathrm{vic}})-\log Z_{\mathrm{vic}} is a lower bound on probabilistic robustness (PR), where ZvicZ_{\mathrm{vic}} denotes the normalizing constant of pvicp_{\mathrm{vic}}. Since PR is generally intractable to compute directly, maximizing this KL-based lower bound provides a tractable surrogate objective for improving PR. We further show that this objective recovers a scaled form of adversarial training, offering a probabilistic interpretation of adversarial training and a principled route to robustness improvement. We call the resulting method probabilistic adversarial training. Experiments show that it consistently improves PR, and ablation studies demonstrate that the induced scaling factor can even enhance the PR of non-probabilistic adversarial training methods.

Figures & tables

Explore similar work

CardsList
  1. Improving Certified Robustness via Adversarial Distillation

    Jun 30, 2026Matteo Melis, Jesus Martinez Del Rincon, Vishal SharmaAdversarial TrainingCertification

  2. Robustness Meets Uncertainty: Evidential Adversarial Training for Robust Selective Classification

    Jul 3, 2026Nicolas Sournac, Ahmed Baha Ben Jmaa, Bertrand BraeckeveldtAdversarial TrainingInput Uncertainties

  3. Robust Alignment: Harmonizing Clean Accuracy and Adversarial Robustness in Adversarial Training

    Apr 29, 2026Yanyun Wang, Qingqing Ye, Li Liu +2Adversarial TrainingHarmonization