cs.CROct 7, 2026

Ask the Expert: LLM-Guided Reinforcement Learning for Autonomous Cyber Defense

Authors: Fernando Martinez, Abhishek Satyam, Tao Li, Junaid Farooq, Ying Wang, Juntao Chen

Organizations: Department of Computer and Information Sciences, Fordham University, New York, USA · Department of Systems Engineering, Stevens Institute of Technology, New Jersey, USA

Abstract

Policy-based reinforcement learning (RL) approaches have produced promising results for autonomous cyber defense; however, they are sample-inefficient in settings where defenders must respond under delayed, partial observations with actions from large action spaces. While large language models (LLMs) may reason semantically about security state space, high latency and trust assumptions prevent attractive in-line deployment models. We introduce Ask the Expert, a training-time guidance framework which first summarizes hard cyber-defense states, then intermittently queries an LLM for host-level defensive recommendations via a constrained action interface, and finally transforms those recommendations into tiered reward shaping for use with PPO. Because the LLM is discarded after training, deployment is a pure RL policy. Across TTCP CAGE CC1 and CC2 and both attacker types, this asymmetric design improves sample efficiency over PPO and outperforms the evaluated potential-based reward shaping (PBRS) baselines, while retaining the strongest terminal mean and requiring no LLM dependency at deployment time.

Figures & tables

Explore similar work

CardsList
  1. Distilling Knowledge from Large Language Models into Lightweight Reinforcement Learning Agents for Autonomous Cyber Operations

    Jul 30, 2026Konur Tholl, François Rivest, Mariam El Mezouar +2Cybersecurity

  2. Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution

    Sep 30, 2026Harshith Doppalapudi, Nathaniel D. Bastian, Ankit ShahHierarchical Reinforcement LearningRetraining

  3. Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

    Aug 5, 2026Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi +6Red-TeamingLarge Language Model Safety