cs.LGFeb 6, 2025

How Vulnerable Is My Learned Policy? Universal Adversarial Perturbation Attacks On Modern Behavior Cloning Policies

Authors: Akansha Kalra, Basavasagar Patil, Guanhong Tao, Daniel S. Brown

Organizations: University of Utah

Abstract

Imitation learning, also known as learning from demonstrations, is a popular approach to train AI models; however, the vulnerability of these models to adversarial attacks remains underexplored. We present the first systematic study of adversarial attacks, across a range of both classic and recently proposed imitation learning algorithms, including Vanilla Behavior Cloning (Vanilla BC), LSTM-GMM, Implicit Behavior Cloning (IBC), Diffusion Policy (DP), and Vector-Quantized Behavior Transformer (VQ-BET). We study the vulnerability of these methods to white-box, grey-box and black-box adversarial perturbations. Our experiments reveal that most existing methods are highly vulnerable to these attacks, including black-box transfer attacks that transfer across algorithms. White-box attacks cause at least a 65% reduction in average task success across all evaluated tasks and algorithms, while the black-box transfer attacks reduce task success by up to 88% on Lift, 99% on Can, and 100% on Square. To the best of our knowledge, we are the first to study and compare the vulnerabilities of different popular imitation learning algorithms to both white-box and black-box attacks. Our findings highlight the vulnerabilities of modern imitation learning algorithms, paving the way for future work in addressing such limitations. Videos and code are available at https://sites.google.com/view/uap-attacks-on-bc.

Figures & tables

Explore similar work

CardsList
  1. Difference-Aware Retrieval Policies for Imitation Learning

    Jun 8, 2026Quinn Pfeifer, Ethan Pronovost, Paarth Shah +3Imitation LearningBehavior Cloning

  2. Provably Efficient Policy-Reward Co-Pretraining for Adversarial Imitation Learning

    Jun 20, 2026Tian Xu, Zexuan Chen, Zhilong Zhang +4Behavior Cloning

  3. Boosting Transferable Adversarial Attacks against Deep Reinforcement Learning

    Oct 5, 2026Zexin Li, Ruili Yao, Yiming Zeng +1