Authors: Christoph Bühler, Matteo Biagiola, Luca Di Grazia, Guido Salvaneschi
Organizations: University of St. Gallen, St. Gallen, SG, Switzerland · University of St. Gallen and Università della Svizzera italiana (USI), St. Gallen and Lugano, SG and TI, Switzerland
AI agents built on large language models (LLMs) run shell commands, read and write files, and reach the network, typically with their user's privileges. However, what an agent does during an execution is difficult to understand: tests assert on the result, and the agent's trajectory records only what the agent reports about itself, which may omit behavior executed by its subprocesses. We present AgentSpy, an approach that observes an agent from outside the agent. AgentSpy runs the agent in an isolated environment, configured by a declarative specification, and records the system calls and network traffic of the agent and of every process it executes. Based on this monitoring, AgentSpy supports two families of analyses: conformance analyses, which measure obligations, i.e., what an agent execution should do, and safety analyses, which check prohibitions, i.e., what an agent execution must never do. We instantiate one analysis of each family. The reliability analysis uses rules to summarize each run by the environment resources the agent uses: the commands it executed, the files it accessed, and the hosts it contacted. The security analysis applies deterministic rules to the system calls of an execution. For reliability, we evaluated AgentSpy on 77 tasks with the codex harness and three recent LLMs, executing each task three times. Sets of repeated runs of the same task are more similar than sets that include runs of another task in 92.2% of the comparisons. Among tasks for which all three runs pass outcome-based tests, the agent performs task-unrelated activities in 18% of the cases, reads the grading files in 7%, and does not use the developers' guidance in 17%. For security, the generic rules of AgentSpy detect four of five attack categories we considered, with no false positives across 50 runs.
Figures & tables
#
Thought (NL)
Action (sh / PL)
Result (Output)
1
Pg Verify the citations; read the citation-management skill first.
sh cat .../SKILL.md; cat /root/test.bib
Pg truncated (1579 ln)
2
Pg Output truncated; list the entries and their titles instead.
sh rg ‘^@|title=’ /root/test.bib
Pg 27 entries; 3 look generic
3
Pg The three generic titles look suspicious; inspect the rest of their entries.
sh sed -n ‘80,360p’ /root/test.bib
Pg 10.5678, 10.1234 (fake prefixes)
4
Pg The suspects’ DOIs look invalid; check the titles against Crossref.
PL for t in titles: urlopen(‘crossref...’+t)
Pg reals match; 3 suspects fuzzy
5
Pg Try a plain web search as a fallback.
sh curl bing; curl duckduckgo
Pg pages returned; no titles parsed
6
Pg Use the skill’s Google Scholar script instead.
sh python .../search_google_scholar.py
Pg scholarly missing
Table 1 . Excerpt (7 of 15 steps) of a codex execution with glm-5.3-flash on citation-check , colored by modality: natural-language reasoning (NL) , shell commands (sh) , inline programming-language code (PL) , and unstructured tool output (Output) .
Figure 1 . Overview of AgentSpy. It takes as input a task, an agent, and an analysis type, and outputs a report summarizing the system-level analysis.
Figure 2 . Execution star graph of the codex trajectory shown in Table 1 . Each node is a resource that the agent uses: a command that it runs, a file that it reads or writes, or a host that it contacts. Each edge connects the execution root to a resource, and holds the number of times that the agent performs the action in the execution. Each grey [ k ] tag refers to the rows (#) of the trajectory in Table 1 the node comes from.
Commands
R1
R2
R3
Files
R1
R2
R3
Hosts
R1
R2
R3
python3.12
5
10
9
reads:test.bib
3
5
4
crossref.org
145
125
165
sed
2
2
8
reads:skill:SKILL.md
2
2
2
openalex.org
15
30
30
curl
–
–
32
reads:skill:gs.py
–
–
2
semanticscholar.org
15
20
45
head
–
–
7
writes:answer.json
1
1
1
dblp.org
–
–
3
python3.12:gs.py
–
–
1
duckduckgo.com
–
–
3
Other (18)
25
45
52
Other (8)
16
23
21
bing.com
–
–
15
Table 2 . Selected resources of the star graphs of three citation-check runs with codex and glm-5.3-flash . The bottom part shows the similarity of each pair of runs and the set for each resource type, and the skill coverage of each run. The symbol “–” indicates a resource that is not in the run. Bold values highlight differences in the runs. gs.py : search_google_scholar.py .
Figure 4 . RQ1 results on gpt-5.6-terra : similarity of runs, and ablation of the resource types.
glm
deepseek
gpt
Total
# Tasks with three passing runs
21
27
40
88
Q1: activity unrelated to the task?
# Tasks flagged
3
11
2
16
C1 environment exploration
1
4
0
5
C2 unrequested network access
2
4
2
8
C3 test file access
0
6
0
6
Table 3 . RQ2: Manual inspection of the tasks whose three runs all pass. A task is flagged when the answer to either of the questions (Q) is deemed true in any of the runs. The clusters (C) are non-exclusive.
Figure 5 . RQ3: AgentSpy security alerts across the 50 Skill-Inject runs.
Median ± std. dev.
Model
Task
With
Without
p -value
d
Reliability analysis: with and without auditd , laurel , and Zeek
gpt-5.6-terra
fastest
23.3 ± 6.5
26.0 ± 5.3
0.43
small
average
147.4 ± 37.4
148.1 ± 22.0
0.79
small
slowest
768.2 ± 151.1
767.3 ± 188.4
0.97
small
deepseek-v4-flash
fastest
77.7 ± 38.4
101.2 ± 51.3
0.62
small
Table 4 . RQ4: Time of the agent execution in seconds over ten runs with and without AgentSpy’s monitoring tools. p -value corresponds to the two-sided Mann-Whitney U test; d is the magnitude of Cohen’s d (small below 0.5 , medium below 0.8 , large otherwise).
Modern AI agents execute real-world side effects through tool calls such as file operations, shell commands, HTTP requests, and database queries. A single unsafe action, including accidental deletion, credential exposure, or data exfiltration, can cause irreversible harm. Existing defenses are incomplete: post-hoc benchmarks measure behavior after execution, static guardrails miss obfuscation and multi-step context, and infrastructure sandboxes constrain where code runs without understanding what an action means. We present AgentTrust, a runtime safety layer that intercepts agent tool calls before execution and returns a structured verdict: allow, warn, block, or review. AgentTrust combines a shell deobfuscation normalizer, SafeFix suggestions for safer alternatives, RiskChain detection for multi-step attack chains, and a cache-aware LLM-as-Judge for ambiguous inputs. We release a 300-scenario benchmark across six risk categories and an additional 630 independently constructed real-world adversarial scenarios. On the internal benchmark, the production-only ruleset achieves 95.0% verdict accuracy and 73.7% risk-level accuracy at low-millisecond end-to-end latency. On the 630-scenario benchmark, evaluated under a patched ruleset and not claimed as zero-shot, AgentTrust achieves 96.7% verdict accuracy, including about 93% on shell-obfuscated payloads. AgentTrust is released under the AGPL-3.0 license and provides a Model Context Protocol server for MCP-compatible agents.
Asynchronous monitoring, incident investigations, and compliance audits primarily rely on agent traces to reconstruct what happened. These analyses assume that LLM agents cannot tamper with their own execution traces. We show that local LLM agents such as Claude Code, Codex, Antigravity, Open Code and Grok Build fail to enforce this boundary. All tested harnesses, except Muse Code, allowed agents to delete their traces when asked, without triggering monitor guardrails. We also validate that external attackers can exploit this gap to induce trace deletion. Finally, we show that trace tampering behavior emerges naturally in frontier models, when agents try to improve their rewards. We advise practitioners to ensure trace logging happens through an independent interception mechanism outside of the agent's control, preserving trace integrity even in cases of full host compromise. Overall, our findings identify a concrete failure of trace integrity in agent infrastructure which can be used to conceal misaligned behaviors like scheming or sabotage.
Jeremy Qin, David Schmotz, Derck Prinzhorn +3
ELLIS Institute Tübingen · Max Planck Institute for Intelligent Systems · Tübingen AI Center +3
Third-party skills are becoming the package ecosystem for LLM agents. They package natural-language instructions, helper scripts, templates, documents, and service configuration into reusable workflows. This makes skills useful, but it also introduces a new security problem: a malicious skill does not need to ask the model to perform an obviously harmful action. Instead, it can disguise the harmful behavior as part of a routine workflow, relying on the agent to execute that workflow with high-value permissions and limited human supervision. We introduce AgentTrap, a dynamic benchmark for evaluating whether LLM agents can use third-party skills while resisting malicious runtime behavior. AgentTrap contains 141 tasks: 91 malicious tasks and 50 benign utility tasks, covering 16 security-impact dimensions grounded in agent-skill supply-chain threats. In each task, the agent receives an ordinary user request, runs with installed skills that may contain malicious workflow elements, and is executed in a sandboxed environment. AgentTrap then judges complete trajectories for attack success, blocked or refused behavior, attack-not-triggered cases, and no-attack-evidence outcomes. Our central finding is that the most informative failures are not simple jailbreaks. Models often complete the visible user task while treating unsafe side effects introduced by the skill as part of the normal workflow. This motivates runtime evaluation of the concrete model--framework--workspace environment in which users actually delegate work. Code and data are available at https://github.com/zhmzm/AgentTrap and https://huggingface.co/datasets/zhmzm/AgentTrap.
Haomin Zhuang, Hanwen Xing, Yujun Zhou +5
University of Notre Dame · University of Southern California · LMU Munich +1