AI Agent Monitoring

Momentum

20 papers in the last four weeks, up 186% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 170

All topics
CardsList
  1. Agent Flight Recorder: Tamper-Evident Audit Trails with On-Chain Anchoring for Long-Horizon Tool-Using Agents

    Sep 1, 2026Laurent Bindschaedler, Quentin Botha, Christoph SiebenbrunnerData ProvenanceAlgorithmic Auditing

  2. CatchBench: When Can an Agent Failure Be Caught?

    Aug 24, 2026Yue Zhao, Mengyuan Li, Ruolin Li +5Agent Failure AnalysisBenchmark Auditing

  3. Agent Safety Should Be a Runtime Contract

    Aug 11, 2026Albus W. Ng, Yi Han, Jusheng Zhang +1AI Agent SafetyAI Agent Monitoring

  4. TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis?

    Aug 8, 2026Yuxuan Zhu, Peng PuSelective PredictionFault Detection

  5. Online Monitoring and Corrective Steering of Programming Agents

    Aug 7, 2026Shuyang Liu, Saman Dehghan, Ji Young Kim +3Software Engineering AgentsLLM Agent Workflow Optimization

  6. TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure

    Aug 7, 2026Joshua Zuniga, Srinivasan Subramanian, Ramya Madhuri Narapureddy +1Agent Failure AnalysisAI Agent Benchmarks

  7. ChainClaw: A Layered Agent Framework for Reliable On-Chain Execution

    Aug 6, 2026Jiacheng Wei, Zhaoxin Fan, Xin Wen +5LLM Agent OrchestrationAI Agent Monitoring

  8. Fail-Fast, Restart-Smart: Early Failure Prediction and Restart for SWE Agentic Tasks

    Aug 4, 2026Chenyu Wang, Yunbo Lyu, Junda He +4Software Engineering AgentsAI Agent Monitoring

  9. Beyond Component Testing: Validating Agentic AI Systems

    Jul 31, 2026Fabio Orazio Mirto, Luca D'Agati, Giuseppe Tricomi +4AI Agent EvaluationAI Agent Safety

  10. AgentGUI: An Interface for Observing and Steering Long-Running AI Agents

    Jul 28, 2026Xuan Zhao, Jiwoong Sohn, Qinyue Zheng +1GUI AgentsHuman-in-the-Loop AI

  11. Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study

    Jul 27, 2026Diego Fernandez Arias, Dev Prashant Mistry, Ren Wang +1Multi-Agent LLM SystemsLLM Backdoor Attacks

  12. Mission-Level Runtime Assurance for LLM-Assisted ISR Swarms over a Verification-Aware Fabric

    Jul 26, 2026Nikolaos Kekatos, Panagiotis Katsaros, Alexios Lekidis +2Swarm RoboticsAI Agent Monitoring

  13. Toward Continuous Assurance for the Democratization of AI Agent Creation in Industry

    Jul 23, 2026Natan Levy, Harel BergerAI Agent ReliabilityAI Agent Monitoring

  14. Clinical Pathways as Safety Specifications for Physical AI in Hospital Wards

    Jul 22, 2026Gabriele Franchini, Giulio Mallardi, Michele De Carolis +1Cyber-Physical SystemsAI Agent Monitoring

  15. ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

    Jul 21, 2026Lena Libon, Ben Rank, Jehyeok Yeon +5AI Agent MonitoringAI Control

  16. AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents

    Jul 21, 2026Kunlun Zhu, Xuyan Ye, Zhiguang Han +9Agent Failure AnalysisAI Agent Reliability

  17. ChainWatch: A Kill Chain-Aligned Sequential Detection Framework for Multi-Step Attacks in MCP-Based AI Agent Systems

    Jul 20, 2026Om Narayan, Rashmi Jyoti, Ramkinker SinghMCP SecurityAI Agent Security

  18. Operational Hallucination and Safety Drift in AI Agents

    Jul 20, 2026Shasha Yu, Fiona Carroll, Barry L. BentleyLLM Agent EvaluationAI Agent Safety

  19. Towards an Intention Abstraction Layer for Autonomous Industrial Systems

    Jul 16, 2026Artan Markaj, Raphael Höfer, Felix GehlhoffAI Agent MonitoringMulti-Agent Systems

  20. Traccia: An OpenTelemetry-Based Governance Platform for AI Systems

    Jul 15, 2026Nutan Kumar Naik, Aditya Kumar Saroj, Vijay Prasad Poudel +2AI AccountabilityAI Agent Monitoring