cs.LGOct 5, 2026

Boosting Transferable Adversarial Attacks against Deep Reinforcement Learning

Authors: Zexin Li, Ruili Yao, Yiming Zeng, Xiaoxue Gao

Organizations: Nanyang Technological University · University of California, Riverside · University of Connecticut · School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen

Abstract

Most adversarial attacks on deep reinforcement learning (DRL) assume white-box access to the victim policy, which rarely holds in practice. This paper studies transfer-based black-box attacks on DRL: the attacker crafts observation perturbations on a white-box surrogate agent and feeds them to an unknown victim. We formulate the attack as return minimization under a per-step perturbation budget. We first show that transplanting transferable image-classification attacks (FGSM, MI-FGSM, and NI-FGSM) with a per-step objective yields perturbations that transfer but are no stronger than random noise of the same budget. We then propose a trajectory-level attack that optimizes a sequence of perturbations over a receding horizon through a differentiable model of the environment and a temperature-smoothed surrogate policy, with the same optimizers. On CartPole-v1 with ten DQN and DDQN agents and 100 surrogate--victim pairs, the trajectory-level attack outperforms per-step attacks and random noise in the white-box, cross-model, and cross-algorithm settings.

Figures & tables

Explore similar work

CardsList
  1. Robust Policy Optimization via Adversarial Importance Sampling

    Sep 14, 2026Amine Andam, Jamal Bentahar, Mustapha HedabouOffline Reinforcement LearningAdversarial Training

  2. How Vulnerable Is My Learned Policy? Universal Adversarial Perturbation Attacks On Modern Behavior Cloning Policies

    Feb 6, 2025Akansha Kalra, Basavasagar Patil, Guanhong Tao +1

  3. RoAd-RL: A Unified Library and Benchmark for Robust Adversarial Reinforcement Learning

    Jun 29, 2026Adithya Mohan, Daniel Kriegl, Torsten SchönDeep Q-Networks