Autonomous red team agents increasingly stress-test AI-enabled cyber defenses by planning strategy and executing multistage attacks. Reinforcement learning (RL) and large language models (LLMs) offer complementary mechanisms for the planning and execution such agents require, and prior work has combined them in hybrid hierarchies. Yet a given architecture is typically developed and evaluated within a single environment, leaving open whether an observed advantage reflects a generally stronger decision mechanism or merely alignment with a particular setting. We address this gap with a controlled cross-environment comparison of two homogeneous hierarchical red team architectures: an RL planner with an RL executor (RL+RL) and an LLM planner with an LLM executor (LLM+LLM). We evaluate both against expert autonomous defenders in CybORG CAGE-4 and in Cyberwheel at two network scales, across 18 configurations under one unified disruption metric. We find a pronounced environment-dependent inversion. RL+RL wins the compact, densely rewarded CAGE-4 (78.5% disruption success versus 18.0% for the strongest LLM configuration) and the 100-host Cyberwheel network (81.0% versus 50.5%), while a pretrained cybersecurity LLM agent wins the larger, escalation-gated 1010-host Cyberwheel network (55.0% versus 0.0% for RL). A kill-chain analysis explains the inversion through architecture-specific bottlenecks that aggregate success rates conceal.In the 1010-host Cyberwheel network, RL discovers and compromises hosts but stalls at privilege escalation, whereas in CAGE-4, LLM agents obtain privileged access but rarely convert it into operational impact. These results indicate that conclusions drawn in a single environment may not generalize, and that hybrid planner-executor designs should be motivated by specific failure modes rather than the assumption that one architecture is universally preferable.
Figures & tables
Figure 1: Both architectures share one two-level pipeline; only the paradigm filling each box changes: blue for RL+RL (PPO policies) and orange for LLM+LLM (served language models). The two levels communicate through a compact discrete intent schema ⟨ tactic, zone, risk ⟩ ; the planner is re-invoked every k steps, while the executor acts every step.
Setting
Hosts
Steps/episode
Defender
CybORG CAGE-4
multi-subnet
500 †
H-MARL expert
Cyberwheel-100
100
100
frozen HRL
Cyberwheel-1010
1010
150
frozen HRL
Table 1: Experimental setup. Each of the six red agent configurations is run in all three settings, giving 3×6=18 cells of n=200 episodes each (3,600 episodes in total). Each RL+RL row is a separate two-level PPO policy trained under its own replanning schedule; Foundation-Sec and DeepHat employ the same model in both roles.
Kill-chain reach (%)
CAGE-4 mechanics (mean/episode)
Configuration
Discover
Scan
Compromise
Root
Impact
Disruption
Degrade
OT-Impact
Zones
Detect
RL+RL (fixed)
100.0
100.0
100.0
76.5
65.0
68.5
61.5
2.27
0.95
0.11
RL+RL (event)
100.0
100.0
100.0
81.0
67.5
72.0
66.0
2.31
1.06
0.08
RL+RL (learned)
100.0
100.0
100.0
81.5
71.0
78.5
70.5
2.34
1.16
0.10
gpt-oss
100.0
100.0
100.0
100.0
2.0
2.0
0.5
0.02
0.02
0.28
Foundation-Sec
100.0
100.0
100.0
96.0
15.0
18.0
9.0
0.70
0.21
0.15
Table 2: CybORG CAGE-4 results (200 episodes per configuration). Kill-chain columns report the percentage of episodes reaching each stage, with Disruption corresponding to success rate. Degrade reports episodes achieving degradation, while OT-Impact, Zones, and Detect report successful impacts on operational technology, disrupted zones, and blue detections per episode. Best disruption is shown in bold.
Configuration
Discover
Scan
Compromise
Root
Impact
Disruption
100-host
RL+RL (fixed)
100.0
100.0
100.0
89.0
81.0
81.0
RL+RL (event)
100.0
100.0
100.0
65.5
62.5
62.5
RL+RL (learned)
100.0
100.0
100.0
60.0
51.0
51.0
gpt-oss
100.0
100.0
100.0
55.0
33.5
33.5
Foundation-Sec
100.0
100.0
100.0
85.0
50.5
50.5
Table 3: Cyberwheel results (200 episodes per configuration). Kill-chain columns report the percentage of episodes reaching each stage. Note that Cyberwheel has no Degrade action. CAGE-4-specific metrics are omitted. Best disruption per scale is shown in bold.
Figure 2: The winning architecture inverts with the environment: the best RL+RL schedule versus the best LLM+LLM model per setting: CAGE-4 and Cyberwheel (CW).
Figure 3: Kill-chain reach for all 18 configurations: percent of the 200 episodes reaching each stage (darker = reached). Every configuration reaches compromise in 100% of episodes (the solid left block) and the divergence is confined to root and impact. RL stalls at escalation on Cyberwheel-1010; gpt-oss roots on CAGE-4 but does not weaponize (100% root, 2% impact); RL+RL (fixed) carries Cyberwheel-100 and Foundation-Sec carries Cyberwheel-1010 furthest through impact (gpt-oss also reaches impact on both scales, at 33.5% and 19.5%).
Figure 4: Conditional conversion at the two escalation transitions, P(root ∣ compromise) and P(impact ∣ root), for the 13 configurations that reach root in at least 40 of 200 episodes (the rest reach root too rarely to define the conversion; Figure 3 ). RL clears both transitions on CAGE-4 and Cyberwheel-100, but on CAGE-4 the language models weaponize root only rarely (gpt-oss 2%); the same language models clear the second transition on Cyberwheel. Which transition breaks depends on the environment, not the architecture.
Figure 5: Each point is one configuration, where color marks the setting and shape the model. The CAGE-4 language models are the most expensive, 102k-217k tokens per episode, yet disrupt at most 18% (gpt-oss and DeepHat, the priciest, at most 2%), whereas Foundation-Sec leads the language models on both Cyberwheel scales at 29k-51k tokens. Capability and environment, not token budget, decide the outcome. Token counts partly reflect horizon: CAGE-4 episodes run 500 steps versus 100-150 for Cyberwheel.
An autonomous cyber defender trained with reinforcement learning (RL) is typically tied to the network on which it was trained, limiting its ability to generalize as network scale changes. Hierarchical RL reduces decision complexity by separating strategic targeting from tactical execution, but it does not eliminate this retraining dependence. We investigate whether frozen, zero-shot large language models (LLMs) can provide retraining-free control in hierarchical cyber defense and how performance changes as LLM control is extended from planning to execution. We formulate a controller-agnostic planner-executor hierarchy in which the planner selects a subnet to defend over a fixed horizon and the executor selects defensive actions within that subnet. Using the high fidelity Cyberwheel environment, with its built-in automated red team agent mapped to the MITRE ATT&CK framework, we compare RL+RL, LLM+RL, and LLM+LLM configurations using six models ranging from 3B to 70B parameters, including two cybersecurity-specialized models, across small, medium, and large networks. Replacing only the planner with an LLM yields limited gains as network size increases. In contrast, extending LLM control to execution produces notable improvements for sufficiently capable models. For instance, a frozen general purpose 70B model holds successful lateral movement to approximately 1% of steps and attacker impact near zero across all three network scales using the same model weights, while the RL baseline is retrained for each scale. Our results show that sufficiently capable frozen LLMs can maintain strong defensive performance across the evaluated network scales without task-specific retraining, while also indicating that strong tactical execution is important to realizing the benefits of LLM-based control.
Harshith Doppalapudi, Nathaniel D. Bastian, Ankit Shah
Indiana University Bloomington, IN, USA · Johns Hopkins University Baltimore, MD, USA
Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclusively against static, heuristic red agents, leaving their robustness against adaptive threats critically understudied. Meanwhile, recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have improved LLM reasoning, but their integration into cybersecurity remains elusive due to the absence of suitable benchmark environments and interaction datasets. To bridge this gap, we introduce Trident, an agentic LLM red teaming framework comprising three components: a dynamic benchmark with isolated sandbox servers spanning CybORG CAGE 4 and CyberWheel, a dataset comprises over 13,000 high-fidelity red-blue interaction trajectories for RLVR, and a ``Code-as-Policy'' RLVR agentic architecture Trident Agentic). The latter reformulates red agent training as a contextual bandit via a tripartite Log Summarizer--Planner--Coder design, where a trainable Planner generates complete attack strategies from compressed execution logs, which a frozen Coder translates into executable Python policies deployed against live DRL defenders. Empirical evaluations reveal a fundamental brittleness in existing defenses: with a single trainable 7B planner, Trident reduces blue agent defensive performance by an average of 522% compared to static red agent baselines while autonomously discovering emergent behaviors such as decoy avoidance and adaptive state prioritization that static heuristics entirely fail to uncover.
Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi +6
University of California, Irvine · Northeastern University · Johns Hopkins University
Automated red-team attacks and blue-team defenses for large language models (LLMs) are advancing quickly. However, attackers and defenders are built and tested in isolation, and the resulting scores are hard to trust. To tackle this, we present ACEA (Adversarial Co-Evolution Arena), a platform that connects a pluggable red-team adapter and a pluggable blue-team adapter to a shared target LLM and scores their attack and defense rates with an LLM judge. ACEA contributes four components. First, a pluggable, model-agnostic arena. Any red or blue project connects over a minimal HTTP protocol, which we call the ACEA Standard Adapter Protocol (ASAP). It can be written in any language, and a project that exposes nothing but the protocol is a full participant. Second, an evaluation methodology built for adversarial rounds. Seeding the target with canonical secrets gives verifiable ground truth that separates real leakage from hallucination. We also send each attack to the target even when the defense blocks it, which measures the attack's raw potency independently of whether it was stopped. Together these yield a per-round decomposition of attack strength and defense effectiveness. Third, a real-time, game-style visualization with a detailed end-of-battle report that localizes each failure. The evaluation thus becomes an actionable signal for improving a red or blue project. Fourth, an optional in-context improvement loop that turns each round's outcome into advisory hints for the next. An adapter can then adapt across rounds without keeping state, provided it reads the hints. We describe the design of ACEA and the metrics through which red and blue teams are scored head to head.