Crossing the Cyber Divide: Sim-to-Sim and Sim-to-Real Transfer for RL Agents
Organizations: The University of Texas at El Paso, El Paso, TX, USA
Abstract
Cyber attack agents are typically trained and evaluated within a single simulator, making it unclear whether learned policies transfer beyond the environments in which they were developed. This limitation hinders both deployment and fair comparison, as cyber simulators differ substantially in their state representations, observation models, and action spaces. In this paper, we study policy transfer across cyber environments and argue that simulator-to-simulator and simulator-to-real transfer can be viewed as instances of the same underlying alignment problem. We propose a framework that separates state alignment from action translation, enabling a policy trained in one environment to operate in another without retraining. We evaluate transfer across four cyber platforms, CyberBattleSim, NetSecGame, CyberWheel, and NASim, including emulated deployments in NASim. Our experiments show that zero-shot transfer is feasible, fully preserving source-policy performance in closely aligned environments and achieving 45.2% win rates when transferring policies whose source performance is 60.5%. In emulated virtual machine environments, transferred policies exhibit a Jensen-Shannon divergence of 0.085 from native policies, indicating strong behavioral similarity. Code and benchmarks are available at: https://anonymous.4open.science/r/RL-Transfer-between-env-4F47/.
Figures & tables
| Observation | Action | ||||||
|---|---|---|---|---|---|---|---|
| Env. | Train-to-converge | Dim | Decision-relevance | Space | Type | Task horizon | Configurable |
| CW | PPO steps (fast) | 78/512 | High (compact kill-chain) | 13 discrete | Informative + exploitative | 7 kill-chain phases/host | Yes (YAML hosts/subnets) |
| NSG | Fast ( CW) | 78 | High (all features relevant) | 12 slot actions | Informative + exploitative | Discovery impact, 25 steps | Fixed topology |
| CBS | Slow | 512 | Low (many irrelevant features) | 9 discrete | Exploitative (discovery supplied) | Short: actions, steps/8 nodes | Yes, harder than CW |
| Emulation (NASim) | n/a | — | n/a | Concrete tool invocations | Informative + exploitative | n/a | VMs, complex setup |
| Ph. | Label | CBS | NSG | NASim | Notes |
|---|---|---|---|---|---|
| 0 | Network disc. | local_exploit (frontier node) | ScanNetwork() | ServiceScan (subnet) | Slot has no effect; is constant |
| 1 | Service disc. | local_exploit (frontier node) | FindServices( ) | OSScan / ProcScan | Slot has no effect; is constant |
| 2 | Exploitation | Connect to target node | ExploitService( ) | SubnetScan | Held until host is controlled |
| 3 | Post-exploit | local_exploit (owned node) | FindData( ) | Exploit( ) | |
| 4 | Exfiltration | Cycles all vulnerabilities | ExfiltrateData( , C&C) | PrivEsc( ) Exfiltrate Data | Requires data present |
| 5+ | Terminal | — | No-op | No-op | NSG/NASim only |
| Condition | CW Win% | NSG Win% | Win Steps | Mean Return |
|---|---|---|---|---|
| Random Policy | — | 0.2% | — | |
| Zero-Shot Transfer | 21.7% | 13.5% | 7.3 | |
| Zero-Shot Transfer + Feature Eng. | 99.1% | 47.4% | 7.3 | |
| Zero-Shot Transfer + Feature Eng. + DAPN | 60.5% | 45.2% | 7.2 |
| Condition | Win Rate | Nodes Owned | Steps to 1st win | Mean Return |
|---|---|---|---|---|
| Random Policy | ||||
| Zero Shot + Feature Eng. | — | |||
| Zero Shot + Feature Eng. +DAPN |