Imitation learning, also known as learning from demonstrations, is a popular approach to train AI models; however, the vulnerability of these models to adversarial attacks remains underexplored. We present the first systematic study of adversarial attacks, across a range of both classic and recently proposed imitation learning algorithms, including Vanilla Behavior Cloning (Vanilla BC), LSTM-GMM, Implicit Behavior Cloning (IBC), Diffusion Policy (DP), and Vector-Quantized Behavior Transformer (VQ-BET). We study the vulnerability of these methods to white-box, grey-box and black-box adversarial perturbations. Our experiments reveal that most existing methods are highly vulnerable to these attacks, including black-box transfer attacks that transfer across algorithms. White-box attacks cause at least a 65% reduction in average task success across all evaluated tasks and algorithms, while the black-box transfer attacks reduce task success by up to 88% on Lift, 99% on Can, and 100% on Square. To the best of our knowledge, we are the first to study and compare the vulnerabilities of different popular imitation learning algorithms to both white-box and black-box attacks. Our findings highlight the vulnerabilities of modern imitation learning algorithms, paving the way for future work in addressing such limitations. Videos and code are available at https://sites.google.com/view/uap-attacks-on-bc.
Figures & tables
Fig. 1 : Universal Adversarial Perturbation (UAP) Attack Pipeline for Behavior Cloning (BC) Algorithms. We craft adversarial attacks on pre-trained behavior cloning (BC) policies (illustrated here using Diffusion Policy πDP ). By applying learned universal adversarial attacks at test time via perturbed visual observations, we test the robustness of BC policies to both white-box and black-box attacks. In the white-box setting (bottom left), the attack degrades the performance of the source policy πDP . In the black-box setting (bottom right), an attack crafted for πDP is applied to different target policies (e.g., πBC and πVQ-BET ). Unattacked and attacked rollouts shown here illustrate successful black-box transfer attack to (πBC ), causing task failure, but limited transfer to (πVQ-BET) .
Fig. 2 : Testing Environments. We craft and evaluate Universal Adversarial Perturbation attacks to study adversarial robustness of modern behavior cloning algorithms under three RoboMimic [ 21 ] benchmark tasks.
TABLE I : Robustness of BC policies to UAP attacks : Task success rates measured as Mean ± Std over 3 seeds with 50 different environment initial conditions per seed (150 in total) under unattacked and UAP attack for Vanilla BC, LSTM-GMM, IBC, DP, and VQ-BET across all three tasks starting from left to right: Lift, Can, and Square task. Our results demonstrate that classical explicit BC methods (Vanilla BC and LSTM-GMM) are highly vulnerable, while implicit methods (IBC and DP) and recent transformer-based modern explicit BC algorithm (VQ-BET) maintain comparatively higher success on some of the simpler tasks, but are still highly vulnerable, especially for more complex tasks like Can and Square.
Fig. 3 : Qualitative and quantitative evaluation of universal adversarial perturbations under different perturbation budgets.
TABLE II : Inter-Algorithm Black-Box and Within-Algorithm Grey-Box Transferability of UAP attacks on Lift, Can, and Square : Each row shows perturbations crafted by the attacker algorithm, applied to target policies in columns under the same task. Each cell represents average task success rates of target policies under UAPs generated from different attacker policies within the same task. Diagonal entries quantify within-algorithm vulnerability under grey box setting, while off-diagonal entries measure cross-algorithm transferability under black box setting. We also include random attacks as a baseline to assess whether the observed performance degradation is due to cross-algorithm transferability rather than random perturbations. All transferability runs are evaluated on 10 rollouts for 6 seeds (total 60 runs) for diagonal grey box and for 9 seeds (total 90 runs) for each non-diagonal black box transfer. To save space, we abbreviated the names of the different BC algorithms across the second header. For this table only, we abbreviate Vanilla BC to BC and LSTM to LSTM-GMM, while keeping IBC, DP, and VQ-BET consistent with the rest of the paper. Lower values indicate more successful transfer attacks, as they correspond to lower target-policy success rates. Diffusion Policy and VQ-BET achieve high task success rates under unattacked settings across all 3 tasks, but are quite vulnerable to black-box UAP attacks.