AI Agent Monitoring

Momentum

20 papers in the last four weeks, up 186% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 170

All topics
CardsList
  1. SwarmReconGuard: Black-Box Detection of Distributed Collective Reconnaissance by Individually Benign-Looking Agent Populations

    Oct 6, 2026Vahid Tavakkoli, Kabeh Mohsenzadegan, Kyandoghere KyamakyaAI Agent SecurityNetwork Intrusion Detection

  2. A Case Study in Assuring AI-Written Software

    Oct 6, 2026Lindsey Ferris, Sierra BonillaAI Agent ReliabilityAI Assurance

  3. Before Agent Tells The Lie: Has Deception Already Been Represented?

    Oct 5, 2026Xinling Li, Dadi Guo, Qingyu Liu +5LLM InterpretabilityDeception in Language Models

  4. AgentPrivArena: Evaluating and Auditing Real-world AI Agent Privacy

    Oct 5, 2026Shouju Wang, Haopeng ZhangAI Agent AuditingLLM Agent Evaluation

  5. What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents

    Oct 5, 2026Zhongxiang Sun, Jiahao Yan, Hongkang Zhao +5Software Engineering AgentsHuman-in-the-Loop Evaluation

  6. Towards a Unified Misuse Monitoring Benchmark

    Oct 5, 2026Aniruddh Pramod, James Oldfield, Adel BibiAI Agent SecurityAI Agent Monitoring

  7. AgentSpy: Making AI Agent Behavior Observable

    Oct 5, 2026Christoph Bühler, Matteo Biagiola, Luca Di Grazia +1AI Agent ReliabilityAI Agent Auditing

  8. Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents

    Sep 30, 2026Haoyu Wang, Wei Zhao, Yedi Zhang +2AI Agent MonitoringLLM Agent Safety

  9. Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks

    Sep 30, 2026Zezhong Wang, Xueyang Tang, Rui Lian +2LLM Agent SecurityAI Agent Monitoring

  10. LLMs Learn to Evade Latent Monitors from Prior Feedback Alone

    Sep 29, 2026Hugo Lyons Keenan, Christopher Leckie, Sarah ErfaniLLM AuditingAdversarial Attacks on LLMs

  11. Improving scalable oversight with co-trained monitors

    Sep 28, 2026Joseph H. Rudoler, Kevin Tan, Benedict Tessler +2Scalable OversightAdversarial Robustness

  12. ResonAct: Streaming Metrics for Runtime Diagnosis and Self-Healing in Multi-Agent Systems

    Sep 28, 2026Tarun Chintada, Neelamadhav Gantayat, Ishaan Romil +3Agent Failure AnalysisAI Agent Monitoring

  13. AgentWare: Automating the Lifecycle of Agentic Applications across the Edge-to-Cloud Continuum

    Sep 28, 2026Michalis Kasioulis, Moysis Symeonides, George Pallis +1LLM Agent OrchestrationLLM Agent Evaluation

  14. LLM Agents Can Easily Tamper With Their Own Traces

    Sep 24, 2026Jeremy Qin, David Schmotz, Derck Prinzhorn +3AI Agent AuditingAI Agent Security

  15. On the Effectiveness of Kernel-Level Evidence for Agent Security

    Sep 24, 2026Spencer King, Zhilu Zhang, Mikhail Kuznetsov +3AI Agent SecurityAI Agent Security Benchmarks

  16. Compliant with Local Controls, Collectively Discriminatory. A Governance Architecture for Multi-Agent AI in Regulated Finance

    Sep 23, 2026Jose Manuel de la Chica Rodriguez, Juan Manuel Vera Diaz, Pablo Delgado RomeroFinancial ServicesAlgorithmic Fairness

  17. Monitorable Chart Reasoning Agents via Verifiable Process Rewards

    Sep 21, 2026Sanchit Sinha, Oana Frunza, Kashif Rasul +1AI Agent MonitoringReinforcement Learning with Verifiable Rewards

  18. Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents

    Sep 17, 2026Alex Remedios, Simon Storf, Fabien Roger +1AI Agent SecurityAI Agent Monitoring

  19. AI Safety: Not Optional, Not Later

    Sep 14, 2026Qinghua Lu, Yoshua BengioAI AssuranceAI Accountability

  20. DynSTEER: Dynamic Stage-wise Trajectory Evaluation and Execution-time Review for Agents

    Sep 13, 2026Zhichao Shi, Xuhui Jiang, Wenjie Zhang +5LLM Agent EvaluationAI Agent Evaluation

  21. DriftNet: A Dual-Head Trajectory Transformer for Detecting and Localizing Prompt Injection in LLM Agents

    Sep 12, 2026Asif Pinjari, Mithun Paul Saint-GermainAI Agent Security BenchmarksAI Agent Monitoring

  22. BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure

    Sep 12, 2026Shenghan Zheng, Zonglin Di, Yimin Liu +19Reward HackingLLM Agent Evaluation

  23. Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations

    Sep 10, 2026Priyanka Mary Mammen, Emil Joswin, Srujananjali MedicherlaAI Agent MonitoringRepresentation Probing

  24. SchemeArena: Factorized Stress Testing of Scheming in LLM Agents

    Sep 8, 2026Jie Ruan, Inderjeet Nair, Amy Liu +3LLM Agent EvaluationAI Agent Monitoring

  25. DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions

    Sep 3, 2026Junjie Pang, Zhenzhen Xie, Haoke Han +3Data ProvenanceAI Agent Auditing

  26. Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision

    Sep 2, 2026Sitong Pan, Yipeng Shen, Yilin Lu +3AI Agent MonitoringProcess Supervision