Automated Penetration Testing

Latest papers 32

Mar 20, 2026cs.LG

NASimJax: A GPU-Accelerated Policy Learning Framework for Penetration Testing

Penetration testing - the practice of simulating cyberattacks to identify vulnerabilities - is a complex sequential decision-making task that is inherently partially observable and features large action spaces. Existing RL simulators for this domain are CPU-bound and fixed to narrow scenarios, making it infeasible to train policies that generalize across networks. We present NASimJax, a JAX-native framework that formulates penetration testing as a Contextual POMDP and introduces a network generation pipeline producing structurally diverse, guaranteed-solvable scenarios. The framework reaches up to 80×\times higher environment throughput than previous simulators, enabling experiments on larger networks and tractable hyperparameter searches. We provide PPO and PQN baselines and conduct the first systematic evaluation of unsupervised environment design for penetration testing. We find that Prioritized Level Replay and ACCEL handle dense training distributions better than Domain Randomization, and that training on sparser topologies yields an implicit curriculum that improves generalization - even to topologies denser than those seen during training. A recurrent PPO variant confirms that these distributional findings are not an artifact of feed-forward policies. The code is available at: https://github.com/raphsimon/NASimJax.
Sep 24, 2025cs.LG

Learning Robust Penetration Testing Policies under Partial Observability: A systematic evaluation

Penetration testing, the simulation of cyberattacks to identify security vulnerabilities, presents a sequential decision-making problem well-suited for reinforcement learning (RL) automation. Like many applications of RL to real-world problems, partial observability presents a major challenge, as it invalidates the Markov property present in Markov Decision Processes (MDPs). Partially Observable MDPs require history aggregation or belief state estimation to learn successful policies. We investigate stochastic, partially observable penetration testing scenarios over host networks of varying size, aiming to better reflect real-world complexity through more challenging and representative benchmarks. This approach leads to the development of more robust and transferable policies, which are crucial for ensuring reliable performance across diverse and unpredictable real-world environments. Using vanilla Proximal Policy Optimization (PPO) as a baseline, we compare a selection of PPO-based variants designed to mitigate partial observability, including frame-stacking, augmenting observations with historical information, and employing LSTM or TrXL architectures. We conduct a systematic empirical analysis of these algorithms across different host network sizes. We find that this task greatly benefits from history aggregation. Converging up to four times faster than other approaches. Manual inspection of the learned policies by the algorithms reveals clear distinctions and provides insights that go beyond quantitative results.