Cyber attack agents are typically trained and evaluated within a single simulator, making it unclear whether learned policies transfer beyond the environments in which they were developed. This limitation hinders both deployment and fair comparison, as cyber simulators differ substantially in their state representations, observation models, and action spaces. In this paper, we study policy transfer across cyber environments and argue that simulator-to-simulator and simulator-to-real transfer can be viewed as instances of the same underlying alignment problem. We propose a framework that separates state alignment from action translation, enabling a policy trained in one environment to operate in another without retraining. We evaluate transfer across four cyber platforms, CyberBattleSim, NetSecGame, CyberWheel, and NASim, including emulated deployments in NASim. Our experiments show that zero-shot transfer is feasible, fully preserving source-policy performance in closely aligned environments and achieving 45.2% win rates when transferring policies whose source performance is 60.5%. In emulated virtual machine environments, transferred policies exhibit a Jensen-Shannon divergence of 0.085 from native policies, indicating strong behavioral similarity. Code and benchmarks are available at: https://anonymous.4open.science/r/RL-Transfer-between-env-4F47/.
Figures & tables
Figure 1: Main architecture of the proposed framework. A policy trained in a source cyber environment is transferred to a target environment through state alignment and action translation, enabling zero-shot execution across simulators and emulated systems.
Observation
Action
Env.
Train-to-converge
Dim
Decision-relevance
Space
Type
Task horizon
Configurable
CW
PPO steps (fast)
78/512
High (compact kill-chain)
13 discrete
Informative + exploitative
7 kill-chain phases/host
Yes (YAML hosts/subnets)
NSG
Fast ( ≈ CW)
78
High (all features relevant)
12 slot actions
Informative + exploitative
Discovery → impact, ≤ 25 steps
Fixed topology
CBS
Slow
512
Low (many irrelevant features)
9 discrete
Exploitative (discovery supplied)
Short: ∼3 actions, ∼8 steps/8 nodes
Yes, harder than CW
Emulation (NASim)
n/a
—
n/a
Concrete tool invocations
Informative + exploitative
n/a
VMs, complex setup
Table 1: Environments and their characteristics
Ph.
Label
CBS ψ
NSG ψ
NASim ψ
Notes
0
Network disc.
local_exploit (frontier node)
ScanNetwork()
ServiceScan (subnet)
Slot has no effect; ψ is constant
1
Service disc.
local_exploit (frontier node)
FindServices( h )
OSScan / ProcScan
Slot has no effect; ψ is constant
2
Exploitation
Connect to target node
ExploitService( h )
SubnetScan
Held until host is controlled
3
Post-exploit
local_exploit (owned node)
FindData( h )
Exploit( h )
4
Exfiltration
Cycles all vulnerabilities
ExfiltrateData( h , C&C)
PrivEsc( h ) Exfiltrate Data
Requires data present
5+
Terminal
—
No-op
No-op
NSG/NASim only
Table 2: Kill-chain phase mapping used by ψ for the CW → CBS, CW → NSG, and CW → NASim transfer pairs.
Figure 2: t-SNE projection of CW (blue) and NSG (orange) observations before (left) and after (right) the DAPN encoder. After encoding, the two domain clouds overlap substantially, confirming that the adversarial training closes the distributional gap between simulators.
Table 4: Three conditions evaluated on CBS chain-12, win =8 nodes owned, 20 episodes.
Figure 3: Action-type distribution over winning trajectories. Each panel compares a transferred policy against the NASim emulator policy across the five NASim action types with Jensen–Shannon divergence reported per panel. Translated CW → NASim (JS =0.087 ). Original CW vs. NASim emulator policy, no win (JS =0.111 ). Translated NSG → NASim (JS =0.085 ).
Due to limited resources and public safety concerns, deep reinforcement learning (RL) agents for many cyber-physical systems (e.g., autonomous vehicles) are first trained in simulators. However, when deployed in real world environments, they often suffer from performance degradation or safety violations because of the inevitable Sim2Real gap. Existing zero-shot approaches, such as robust safe RL and domain randomization, mitigate this issue but typically at the cost of degraded performance or residual safety risks when experiencing unmodeled system dynamics. To address these limitations, we propose a novel reinforcement learning framework that enables safe and efficient policy transfer via probabilistic latent embeddings and dynamic policy adaptation. We consider a family of Constrained Markov Decision Processes (CMDPs) under different environment contexts. By leveraging latent context variable in meta-RL, the proposed framework infers the latent representation of the environment from simulated experiences. Furthermore, it incorporates a distributional RL formulation, which allows risk levels of the deployed policy to be adjusted dynamically, based on the estimation accuracy of the latent context variable. This strategy promotes safety at the early deployment stage and improves efficiency through fast policy adaptation under the Sim2Real gap.
Gengyue Han, Yiheng Feng
Lyles School of Civil and Construction Engineering, Purdue University, West Lafayette, USA · Elmore Family School of Electrical and Computer Engineering, Purdue University, West Lafayette, USA
To mitigate the sample complexity of real-world reinforcement learning (RL), a common practice is to first train a policy in a simulator, where samples are cheap, and then deploy the learned policy in the real world with the hope that it generalizes effectively. Such direct sim-to-real transfer is not guaranteed to succeed: simulator-trained policies can be suboptimal in the real world due to sim-to-real mismatch. Correcting this mismatch requires collecting data from the real system, but in many applications, such as robotics and healthcare, this data-collection process is itself subject to safety constraints. This gives rise to the problem of safe sim-to-real transfer: how can an agent exploit an imperfect simulator while ensuring safe real-world data collection and learning a near-optimal feasible policy for the target system? We address this problem by formulating safe sim-to-real transfer within the framework of reward-free safe RL. We design a computationally efficient algorithm that exploits simulator information to provably reduce real-world interaction while ensuring safe exploration and enabling the computation of a near-optimal feasible policy for any potential reward function. Our real-world sample complexity bound characterizes the benefit of using the simulator in terms of the sim-to-real mismatch.
In recent years, reinforcement learning (RL) has shown remarkable success in robotics when a fast and accurate simulator is available for a given task. When using RL and simulation, more simulator realism is generally beneficial but becomes harder to obtain as robots are deployed in increasingly complex and widescale domains. In such settings, simulators will likely fail to model all relevant details of a given target task and this observation motivates the study of sim2real with simulators that leave out key task details. In this paper, we formalize and study the abstract sim2real problem: given an abstract simulator that models a target task at a coarse level of abstraction, how can we train a policy with RL in the abstract simulator and successfully transfer it to the real-world? Our first contribution is to formalize this problem using the language of state abstraction from the RL literature. This framing shows that an abstract simulator can be grounded to match the target task if the grounded abstract dynamics take the history of states into account. Based on the formalism, we then introduce a method that uses real-world task data to correct the dynamics of the abstract simulator. We then show that this method enables successful policy transfer both in sim2sim and sim2real evaluation.
Yunfu Deng, Yuhao Li, Josiah P. Hanna
Department of Computer Sciences, University of Wisconsin–Madison, Madison, WI 53706 USA · Manning College of Information and Computer Sciences, University of Massachusetts Amherst, Amherst, MA 01003 USA