cs.CROct 6, 2026

Multi-Aspect Runtime Verification for Simulation-Based V&V of LLM-Enabled Autonomous Agents

Authors: Nikolaos Kekatos, Dimitrios Nikou, Anastasios Temperekidis, Alexios Lekidis, Nikolaos Kolokotronis, Panagiotis Katsaros, Stylianos Basagiannis

Organizations: Clone Systems, Cyprus · International Hellenic University, Greece · University of Thessaly, Greece · University of Peloponnese, Greece · Aristotle University of Thessaloniki, Greece

Abstract

LLM-based agents are entering decision-support roles in defence staff work, where the obligations they must respect are already written down and binding, and where retraining is not available as a control because models arrive as procured components. What can be placed under engineering control is the interface between the agent and the systems it acts on. Those obligations are at once spatial, temporal and text-semantic, and a violation typically lives in the composition of a multi-step interaction, which is why per-event guardrails miss sequential tool-attack chains. We present a multi-aspect runtime-verification framework that decomposes a natural-language policy clause into a typed spatial/temporal/semantic triple over one canonical event stream, checks each aspect with its own monitoring specification, and fuses the verdicts through a four-valued algebra that carries provenance. The spatial aspect is interpreted over a weighted two-sorted location graph in which mission geometry and information-release topology are one object; we show that these spatial obligations are not in general subsumed by a first-order temporal specification. The past-time aspect runs on the unmodified MonPoly engine, which agrees with our reference monitor at every time point. Across two mission domains, casualty evacuation and contested sustainment, and one civil domain, composition under the precautionary blocking policy drives attack success to zero with no observed false positives and microsecond-scale per-event cost, while every single aspect and every pair leaves a substantial share of attacks succeeding. In a closed-loop experiment a policy-naive planner reaches a violating state in most unshielded missions and in none when shielded, and four refused episodes in five still recover to a compliant outcome.

Figures & tables

Explore similar work

CardsList
  1. Mission-Level Runtime Assurance for LLM-Assisted ISR Swarms over a Verification-Aware Fabric

    Jul 26, 2026Nikolaos Kekatos, Panagiotis Katsaros, Alexios Lekidis +2Runtime EnforcementSwarms

  2. ActGov: Governing LLM Agent Actions via Policy-Constrained Validation

    Sep 21, 2026Kaiyuan Zhang, Yuke Peng, Ke Jiang +1AuthorizationRuntime Enforcement