Imitation learning, also known as learning from demonstrations, is a popular approach to train AI models; however, the vulnerability of these models to adversarial attacks remains underexplored. We present the first systematic study of adversarial attacks, across a range of both classic and recently proposed imitation learning algorithms, including Vanilla Behavior Cloning (Vanilla BC), LSTM-GMM, Implicit Behavior Cloning (IBC), Diffusion Policy (DP), and Vector-Quantized Behavior Transformer (VQ-BET). We study the vulnerability of these methods to white-box, grey-box and black-box adversarial perturbations. Our experiments reveal that most existing methods are highly vulnerable to these attacks, including black-box transfer attacks that transfer across algorithms. White-box attacks cause at least a 65% reduction in average task success across all evaluated tasks and algorithms, while the black-box transfer attacks reduce task success by up to 88% on Lift, 99% on Can, and 100% on Square. To the best of our knowledge, we are the first to study and compare the vulnerabilities of different popular imitation learning algorithms to both white-box and black-box attacks. Our findings highlight the vulnerabilities of modern imitation learning algorithms, paving the way for future work in addressing such limitations. Videos and code are available at https://sites.google.com/view/uap-attacks-on-bc.
Figures & tables
Fig. 1 : Universal Adversarial Perturbation (UAP) Attack Pipeline for Behavior Cloning (BC) Algorithms. We craft adversarial attacks on pre-trained behavior cloning (BC) policies (illustrated here using Diffusion Policy πDP ). By applying learned universal adversarial attacks at test time via perturbed visual observations, we test the robustness of BC policies to both white-box and black-box attacks. In the white-box setting (bottom left), the attack degrades the performance of the source policy πDP . In the black-box setting (bottom right), an attack crafted for πDP is applied to different target policies (e.g., πBC and πVQ-BET ). Unattacked and attacked rollouts shown here illustrate successful black-box transfer attack to (πBC ), causing task failure, but limited transfer to (πVQ-BET) .
Fig. 2 : Testing Environments. We craft and evaluate Universal Adversarial Perturbation attacks to study adversarial robustness of modern behavior cloning algorithms under three RoboMimic [ 21 ] benchmark tasks.
TABLE I : Robustness of BC policies to UAP attacks : Task success rates measured as Mean ± Std over 3 seeds with 50 different environment initial conditions per seed (150 in total) under unattacked and UAP attack for Vanilla BC, LSTM-GMM, IBC, DP, and VQ-BET across all three tasks starting from left to right: Lift, Can, and Square task. Our results demonstrate that classical explicit BC methods (Vanilla BC and LSTM-GMM) are highly vulnerable, while implicit methods (IBC and DP) and recent transformer-based modern explicit BC algorithm (VQ-BET) maintain comparatively higher success on some of the simpler tasks, but are still highly vulnerable, especially for more complex tasks like Can and Square.
Fig. 3 : Qualitative and quantitative evaluation of universal adversarial perturbations under different perturbation budgets.
TABLE II : Inter-Algorithm Black-Box and Within-Algorithm Grey-Box Transferability of UAP attacks on Lift, Can, and Square : Each row shows perturbations crafted by the attacker algorithm, applied to target policies in columns under the same task. Each cell represents average task success rates of target policies under UAPs generated from different attacker policies within the same task. Diagonal entries quantify within-algorithm vulnerability under grey box setting, while off-diagonal entries measure cross-algorithm transferability under black box setting. We also include random attacks as a baseline to assess whether the observed performance degradation is due to cross-algorithm transferability rather than random perturbations. All transferability runs are evaluated on 10 rollouts for 6 seeds (total 60 runs) for diagonal grey box and for 9 seeds (total 90 runs) for each non-diagonal black box transfer. To save space, we abbreviated the names of the different BC algorithms across the second header. For this table only, we abbreviate Vanilla BC to BC and LSTM to LSTM-GMM, while keeping IBC, DP, and VQ-BET consistent with the rest of the paper. Lower values indicate more successful transfer attacks, as they correspond to lower target-policy success rates. Diffusion Policy and VQ-BET achieve high task success rates under unattacked settings across all 3 tasks, but are quite vulnerable to black-box UAP attacks.
Parametric imitation learning via behavior cloning can suffer from poor generalization to out-of-distribution states due to compounding errors during deployment. We show that reusing the training data during inference via a semi-parametric retrieval-based imitation learning approach can alleviate this challenge. We present Difference-Aware Retrieval Policies for Imitation Learning (DARP), a semi-parametric retrieval-based imitation learning approach that addresses this limitation by reparameterizing the imitation learning problem in terms of local neighborhood structure rather than direct state-to-action mappings. Instead of learning a global policy, DARP trains a model to predict actions based on k-nearest neighbors from expert demonstrations, their corresponding actions, and the relative distance vectors between neighbor states and query states. DARP requires no additional assumptions beyond those made for standard behavior cloning -- it does not require additional data collection, online expert feedback, or task-specific knowledge. We demonstrate consistent performance improvements of 15-46% over standard behavior cloning across diverse domains, including continuous control and robotic manipulation, and across different representations, including high-dimensional visual features. Code and demos are available at https://weirdlabuw.github.io/darp-site/.
Quinn Pfeifer, Ethan Pronovost, Paarth Shah +3
1Paul G. Allen School of Computer Science & Engineering, University of Washington · 2Toyota Research Institute · 3Google DeepMind +1
Adversarial imitation learning (AIL) achieves high-quality imitation compared to behavioral cloning (BC), but demands substantial online environment interaction. Recent empirical work has explored initializing AIL algorithms with BC pretrained policies to address this limitation, yet a rigorous theoretical understanding of pretraining's role in AIL remains elusive. This paper provides a systematic theoretical analysis and introduces principled pretraining algorithms for accelerating AIL. We begin by analyzing AIL with policy pretraining alone, identifying reward error as the dominant source of suboptimality. This reveals a critical and previously overlooked gap: the absence of reward pretraining. Motivated by this finding, we develop a principled policy-reward co-pretraining approach grounded in a reward shaping analysis. Our analysis uncovers a fundamental connection between expert policies and shaping rewards, which naturally gives rise to CoPT-AIL, an approach that jointly pretrains both policy and reward through a single BC procedure. We prove that CoPT-AIL achieves an improved imitation gap bound over standard AIL, establishing the first theoretical guarantee for the benefits of pretraining in AIL. Experimental results confirm CoPT-AIL's superior performance over existing AIL methods.
Tian Xu, Zexuan Chen, Zhilong Zhang +4
National Key Laboratory for Novel Software Technology and School of Artificial Intelligence, Nanjing University, China.
Most adversarial attacks on deep reinforcement learning (DRL) assume white-box access to the victim policy, which rarely holds in practice. This paper studies transfer-based black-box attacks on DRL: the attacker crafts observation perturbations on a white-box surrogate agent and feeds them to an unknown victim. We formulate the attack as return minimization under a per-step perturbation budget. We first show that transplanting transferable image-classification attacks (FGSM, MI-FGSM, and NI-FGSM) with a per-step objective yields perturbations that transfer but are no stronger than random noise of the same budget. We then propose a trajectory-level attack that optimizes a sequence of perturbations over a receding horizon through a differentiable model of the environment and a temperature-smoothed surrogate policy, with the same optimizers. On CartPole-v1 with ten DQN and DDQN agents and 100 surrogate--victim pairs, the trajectory-level attack outperforms per-step attacks and random noise in the white-box, cross-model, and cross-algorithm settings.
Zexin Li, Ruili Yao, Yiming Zeng +1
Nanyang Technological University · University of California, Riverside · University of Connecticut +1