Progress in adversarial multi-agent reinforcement learning (MARL) for robotics has been hampered by a lack of shared, extensible infrastructure that supports heterogeneous agent morphologies in high-fidelity physics simulation. Existing frameworks either focus on cooperative tasks, rely on simplified physics engines, or provide isolated implementations that are difficult to extend. We present HARL-A, an open-source, actively maintained framework built on IsaacLab that enables scalable training and benchmarking of adversarial policies across morphologically diverse robot teams with any number of teams and any mix of robot morphologies per team. HARL-A extends the HARL algorithm library and IsaacLab with adversarial multi-agent support and contributes three components: (1) a modular software architecture that reduces the engineering overhead of defining new heterogeneous adversarial environments, (2) a suite of three benchmark environments---Sumo, Soccer, and 3D Galaga---spanning contact-rich pushing, ball-skill competition, and pursuit/evasion, (3) over ten pretrained policies spanning homogeneous and heterogeneous team configurations, released publicly on Hugging Face to enable immediate exploration of adversarial learning dynamics without retraining from scratch. We demonstrate the framework across multiple competitive scenarios, showing that it reliably produces learned adversarial policies and emergent role specialization. All code environments, trained policies, and documentation are openly available at https://github.com/DIRECTLab/IsaacLab-HARL.
Figures & tables
Fig. 1 : HARL-A environments in IsaacLab. (Top) Two heterogeneous teams—each comprising one Anymal C quadruped and one Leatherback rover—compete in an adversarial Sumo task. (Bottom) Multi-team configurations with three distinct agent morphologies illustrate the framework’s support for arbitrary team composition and scale.
Author (Year)
Environment
Adversarial
High-Fidelity Physics
Multi-Agent
Heterogeneous
Extensible Framework
Bansal et al. (2018) [ 8 ]
MuJoCo
✓
Gleave et al. (2020) [ 12 ]
MuJoCo
✓
Lowe et al. (2017) [ 13 ]
Particle Env.
✓
✓
Baker et al. (2020) [ 9 ]
MuJoCo
✓
✓
Schwarting et al. (2021) [ 14 ]
Visual RL
✓
✓
Liu et al. (2021) [ 15 ]
MuJoCo
✓
✓
TABLE I : Comparison of adversarial MARL works. HARL-A is the only framework that simultaneously satisfies all five properties.
Fig. 2 : Actor-critic training paradigms in adversarial reinforcement learning. Prior cooperative frameworks supported paradigm C (heterogeneous coordination with shared critic). HARL-A adds support for paradigms B (heterogeneous single-agent adversarial) and D (heterogeneous multi-team adversarial), enabling the full spectrum of competitive training configurations.
Fig. 3 : Evolution of the observation space across Sumo curriculum stages for the Anymal C robot. Zero-padded entries (shown as 0 ) are replaced with meaningful features as task complexity increases, following the zero-buffer strategy of [ 21 ] . This maintains a fixed observation dimension across all stages, enabling seamless policy transfer. This strategy was used for all benchmark environments presented in this paper.
Fig. 4 : Competitive behaviors in the heterogeneous Soccer environment (Anymal C vs. Leatherback). Top: Leatherback maintains possession and scores. Bottom: Leatherback possession followed by a turnover and an Anymal goal, illustrating the competitive dynamics that emerge from morphological asymmetry.
Fig. 5 : 3D Galaga: Anti-Aircraft Defense. MiniTanks fire arm-aligned laser-tag rays (2, 3) each control step. A drone (path shown in 1) is knocked out when it enters the radius of an active ray (4). The environment demonstrates HARL-A’s ability to compose pretrained locomotion primitives into novel adversarial scenarios.
Fig. 6 : Win rates of adversarially trained policies vs. a fixed baseline policy across the three benchmark environments (top: 3D Galaga, middle: Sumo, bottom: Soccer). Blue shows the agents trained adversarially, while orange represents agents trained against a fixed policy. For each saved checkpoint, we evaluate 1000 parallel environments against the initial baseline policy and average the win rate. Increasing win rate over training confirms that meaningful adversarial learning occurs across all three environments and task types.
Fig. 7 : Emergent role specialization in the heterogeneous Sumo environment (one Anymal C and one Leatherback per team). Without any explicit role assignment, the Leatherback learned to destabilize the opposing Anymal by targeting its legs (Top), while the Anymal learned to drag the opposing Leatherback out of the arena using its limbs (Bottom). These complementary strategies emerged consistently from morphological asymmetry.
Fig. 8 : FPS training distribution based on the number of agents in each environment. As can be seen in the figure, increased agent count decreases the FPS of the training.
Environment
Starting
Trained
Leatherback-Stage1-Soccer-v0
NO
YES
Leatherback-Stage2-Soccer-v0
YES
YES
AnymalC_Soccer_Hetero_By_Team-v0
YES
NO
Sumo-Stage2-Hetero-By-Team-v0
YES
NO
Sumo-Stage2-Hetero-Same-Critic-v0
YES
NO
Sumo-Stage2-Hetero-Same-Critic-No-Neg-v0
YES
NO
TABLE II : Pretrained policies available on Hugging Face. Environments span all three benchmark task categories, multiple curriculum stages, and both homogeneous and heterogeneous team configurations.
Deep Reinforcement Learning (DRL) has achieved significant success in robotics and autonomous systems, yet remains vulnerable to adversarial perturbations that can severely degrade performance. Research in adversarial reinforcement learning is often limited by fragmented implementations, inconsistent evaluation protocols, and poor reproducibility. To address these challenges, we present \textbf{RoAd-RL}, an open-source benchmarking framework that provides unified abstractions for policies, attacks, defenses, and robustness metrics, together with reproducible evaluation pipelines and seamless integration with Stable-Baselines3 and Gymnasium. We evaluate DQN, PPO, and SAC agents in LunarLander and Highway-v0 under 192 attack-defense configurations. Results reveal substantial variations in robustness across environments and show that some commonly used defenses can be more detrimental than the attacks they aim to mitigate, while temporal smoothing consistently achieves strong performance. RoAd-RL establishes a standardized benchmark for adversarial reinforcement learning research and is publicly available at https://pypi.org/project/road-rl.
Cooperation is central to multi-agent reinforcement learning (MARL), yet learned coordination can be fragile when external perturbations disrupt inter-agent interactions. Prior robust MARL methods have primarily considered value-oriented attacks, leaving a gap in robustness when interaction structures themselves are corrupted. In this paper, we propose an interaction-breaking adversarial learning (IBAL) framework that takes an information-theoretic view to construct attacks that impede coordination by perturbing agents' observations and actions, and trains agents to perform reliably under such disruptions. Empirically, our approach improves robustness over existing robust MARL baselines across diverse attack settings and yields stronger performance even under agent-missing scenarios. Our code is available at https://sunwoolee0504.github.io/IBAL.
Sunwoo Lee, Mingu Kang, Yonghyeon Jo +1
Graduate School of Artificial Intelligence, UNIST, Ulsan, South Korea
Reinforcement learning (RL) has become a powerful paradigm for robot learning, particularly in sim-to-real settings, but its broader adoption remains limited by the engineering pipeline surrounding the algorithms. Building tasks, shaping rewards, and tuning hyperparameters require substantial expert effort, making RL workflows costly and difficult to scale. We introduce HARBOR, an agentic framework that frames robot RL automation as a harness-engineering problem: given a simulator codebase and a task specification, it automates the workflow from environment setup to policy training in simulation. HARBOR decomposes such high-level objectives into bounded stages executed by specialized agents through standardized commands, persistent artifacts, executable gates, and reusable knowledge, and scales iteration via decentralized parallel trials and experience learning across runs. We evaluate HARBOR across 6 benchmarks and 16 tasks in total, spanning manipulation, locomotion, and bimanual dexterous control. We demonstrate that HARBOR automates the simulation RL workflow end-to-end, designs rewards, tunes algorithms to match or improve over default configurations, and reduces engineering effort at practical token and wall-clock cost; the resulting policies can also be transferred to real robots.
Zechu Li, Yufeng Jin, Xiaoyang Liu +4
TU Darmstadt 1 · Hessian.AI 7 · Honda Research Institute Europe 2 +4