Security Operations Centers (SOCs) for information technology and operational technology share one incident-response problem: a flood of correlated alerts and too few analysts. Large Language Models (LLMs) are increasingly proposed as reasoning engines that triage alerts and, in autonomous deployments, issue commands that block IPs, kill processes, or quarantine files on production hosts. This coupling introduces a new risk: a single adversarial alert can become a remote code path through the LLM's reasoning, leading it to recommend an action the SOC then executes. We present a constrained-action architecture with two coordinated layers: (i) a SIEM/XDR control plane that grounds remediation in correlated host events and confines the LLM's output to a closed intent vocabulary whose templated commands are executed by thin endpoint agents, backstopped by an argument validator; and (ii) a NeMo-Guardrails proxy that wraps the SOC-analyst LLM with input- and output-rail policies, evaluated out-of-the-box against a SOC-specific adversarial corpus we release. The stock proxy lifts injection recall from 25.0% to 94.5% at a 0.1% false-positive rate, and a live red-team exercise confirms that the closed intent vocabulary and argument validator contain the observed LLM failure modes before any command crosses the trust boundary. As an architectural fit (not yet a measured operational-technology deployment), the constrained-action property suits critical-infrastructure settings where a wrong remediation has physical, not merely operational, consequences. The loop is best run human-in-the-loop or delayed: the measured rail latency keeps inline control out of scope.
Figures & tables
Fig. 1: Trust-boundary view of the remediation pipeline. Shading encodes distrust: dark dashed = attacker-controlled, mid-gray = semi-trusted (LLM provider), light gray with a heavy border = policy boundary (the NeMo proxy, which screens all traffic to and from the LLM), white = fully trusted (SIEM/XDR control plane). Arrow labels name the data form crossing each boundary, as raw attacker bytes are progressively sanitised into a whitelisted command. The proxy blocks G1 (telemetry-borne injection) and G3 (leakage via the reply); the validator blocks G2 (attacker-induced action): the attacker goals of Section III .
Fig. 2: Request flow through the NeMo Guardrails proxy. The input rail classifies whether the user message is safe to forward; if it blocks, the main LLM is never invoked. The output rail re-screens the completion before it returns to the caller.
Category
Count
AdvBench (sampled) [ 25 ]
60
JailbreakBench [ 26 ]
40
HarmBench (sampled) [ 27 ]
50
SOC augmentation: homoglyph (Cyrillic → Latin)
20
SOC augmentation: Base64 wrapping
15
SOC augmentation: fake [SYSTEM] framing
10
TABLE I: Adversarial-corpus composition (200 variants). Rows 1–3 are sampled from published benchmarks; rows 4–7 are SOC-specific transformations.
Configuration
Recall
FPR
P50 (s)
P95 (s)
P99 (s)
Raw LLM (no defence)
25.0%
0.0%
0.31
0.48
1.20
NeMo Guardrails proxy
94.5%
0.1%
14.14
27.41
32.90
TABLE II: Head-to-head injection-defence comparison. Recall is over the 200 adversarial variants of Table I ; FPR is over the 1,000 benign LogHub lines. Latencies are end-to-end per call.
Category
n
Raw recall
NeMo recall
AdvBench (base)
60
21.7%
100.0%
JailbreakBench (base)
40
35.0%
100.0%
HarmBench (base)
50
20.0%
80.0%
SOC aug.: homoglyph
20
25.0%
95.0%
SOC aug.: Base64
15
13.3%
100.0%
SOC aug.: fake [SYSTEM]
10
40.0%
100.0%
TABLE III: Per-category recall on the 200 adversarial records (base prompts wrapped in HTTP/syslog/SSH log shapes; SOC rows further transformed).
Fig. 3: End-to-end latency CDF (log-scale x-axis), split by endpoint and label: the raw curve concentrates near ∼ 0.3 s; the NeMo curve is bimodal between fast-blocked and full-pipeline traffic.
Configuration
Recall
FPR
Calls/req
P50 (s)
Full NeMo proxy (Table II )
94.5%
0.10%
3.00
14.14
& deterministic Tier-0 (recall-safe)
95.0%
0.10%
2.88
14.11
Cheap-gate cascade (fast triage)
58.0%
0.00%
1.15
0.18
TABLE IV: Cascade triage over the full corpus (200 adversarial, 1,000 benign). The deterministic Tier-0 gate inverts the four SOC augmentations of Table I ; calls/req is the hardware-independent comparison (pass-path count for the proxy rows; measured mean, including the gate call, for the cascade).
Large language model (LLM) agents are increasingly proposed as autonomous SOC analysts, but two limitations make them unreliable at enterprise scale: a finite context window cannot hold a multi-thousand-host authentication graph, and free-form generation offers no guarantee that a recommended containment action is consistent with the topology it operates on. We present Sentinel-RL, an agentic-SOC architecture that decouples topological reasoning from semantic reasoning: a heterogeneous graph attention encoder summarizes the live authentication subgraph into a fixed-dimensional state, a Proximal Policy Optimization (PPO) policy maps this state to a constrained set of investigative actions, and an LLM agent loop is restricted to consuming the policy's recommendations and producing analyst-readable narratives gated by a critic. We instantiate the system on the LANL Comprehensive, Multi-Source Cyber-Security Events dataset and the Indiana University Quartz HPC cluster, reporting four results: (i) a two-phase CREATE ingestion pattern loads a 24M-edge authentication subgraph into Neo4j in 14.2 minutes on a single 32-core node, roughly 24x faster than the canonical MERGE-based pipeline; (ii) a sliding-window alert engine reliably trips a 25-event/10-second threshold in <=2.5 s across 50 trials; (iii) PPO training over 200 iterations converges to a mean episodic return of 8.74+/-0.31, with held-out precision of 0.91 and recall of 0.87 on labeled red-team events; and (iv) the integrated containment loop completes a full detect-investigate-recommend-human-approve cycle in a median of 6.3 s. We contribute a reusable engineering pattern (the hot-node deadlock workaround), a portable HPC deployment pattern (anchor-node co-location), and an enterprise-readiness analysis covering false-positive economics, reversibility guarantees, audit compliance, and the human-approval boundary.
Uday Vallabhaneni, Cassie L. Cagwin, David J. Wild
Luddy School of Informatics, Computing and Engineering · Indiana University, Bloomington, IN, USA
Large language model (LLM) agents increasingly issue API calls that mutate real systems, yet many current architectures pass stochastic model outputs directly to execution layers. We argue that this coupling creates a safety risk because model correctness, context awareness, and alignment cannot be assumed at execution time. We introduce Sovereign Agentic Loops (SAL), a control-plane architecture in which models emit structured intents with justifications, and the control plane validates those intents against true system state and policy before execution. SAL combines an obfuscation membrane, which limits model access to identity-sensitive state, with a cryptographically linked Evidence Chain for auditability and replay. We formalize SAL and show that, under the stated assumptions, it provides policy-bounded execution, identity isolation, and deterministic replay. In an OpenKedge prototype for cloud infrastructure, SAL blocks 93% of unsafe intents at the policy layer, rejects the remaining 7% via consistency checks, prevents unsafe executions in our benchmark, and adds 12.4 ms median latency.
The integration of Large Language Models (LLMs) into Security Operations Centers (SOCs) streamlines threat intelligence but introduces critical vulnerabilities, notably indirect prompt injection via log poisoning. Adversaries exploit this vector to execute multistep ``promptware'' kill chains by embedding malicious payloads within system logs to hijack the LLM's operational logic. Securing this pipeline presents a dichotomy: deterministic defenses are computationally efficient yet semantically blind, while purely neural evaluations introduce prohibitive latency and probabilistic flaws. To address this, we propose a novel neurosymbolic defense-in-depth architecture that ensures end-to-end pipeline integrity. The primary layer employs customized SIEM decoders as a deterministic pre-filter, performing immediate structural sanitization to neutralize volumetric padding and signature-based injections at the ingestion edge. The secondary layer leverages NeMo Guardrails to enforce strict semantic boundaries through self-checking validation on the structured SIEM alerts prior to LLM processing. Furthermore, the framework integrates a closed-loop telemetry system, providing critical Human-in-the-Loop (HITL) visibility into thwarted attacks directly within the SOC dashboard. We present a comprehensive experimental evaluation mapped to the MITRE ATLAS taxonomy, assessing the framework against diverse prompt injections. Our results demonstrate that this synergistic approach effectively dismantles the promptware kill chain - bounding LLM stochasticity with verifiable constraints, and delivering a resilient, highly observable defense mechanism for next-generation AI-SOCs.
Anna Gazani, Spyridon Kounoupidis, Panagiotis Katsaros +3
Aristotle University of Thessaloniki, Greece · Clone Systems, Cyprus