AI Agent Monitoring

Momentum

20 papers in the last four weeks, up 186% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 170

All topics
CardsList
  1. Step-level Optimization for Efficient Computer-use Agents

    Apr 29, 2026Jinbiao Wei, Kangqi Ni, Yilun Zhao +2Computer-Use AgentsAI Agent Monitoring

  2. Operating-Layer Controls for Onchain Language-Model Agents Under Real Capital

    Apr 28, 2026T. J. Barton, Chris Constantakis, Patti Hauseman +4AI Agent ReliabilityLLM Agent Evaluation

  3. Governing What You Cannot Observe: Adaptive Runtime Governance for Autonomous AI Agents

    Apr 27, 2026German Marin, Jatin ChaudharyAI Agent SafetyAI Agent Monitoring

  4. GAMMAF: A Common Framework for Graph-Based Anomaly Monitoring Benchmarking in LLM Multi-Agent Systems

    Apr 27, 2026Pablo Mateo-Torrejón, Alfonso Sánchez-MaciánMulti-Agent LLM SystemsAI Agent Security Benchmarks

  5. Omission Constraints Decay While Commission Constraints Persist in Long-Context LLM Agents

    Apr 22, 2026Yeran GamageAI Agent MonitoringLLM Security

  6. Auditing and Controlling AI Agent Actions in Spreadsheets

    Apr 22, 2026Sadra Sabouri, Zeinabsadat Saghi, Run Huang +4Human-in-the-Loop AIAI Agent Monitoring

  7. The Alignment Flywheel: A Governance-Centric Hybrid MAS for Architecture-Agnostic Safety

    Feb 28, 2026Elias Malomgré, Pieter SimoensAI Agent SafetyAI Agent Monitoring

  8. StepShield: When, Not Whether to Intervene on Rogue Agents

    Jan 29, 2026Gloria Felicia, Zitha Sasindran, Jinfeng He +3LLM GuardrailsAI Agent Security

  9. Measuring Harmfulness of Computer-Using Agents

    Jul 31, 2025Aaron Xuxiang Tian, Ruofan Zhang, Janet Tang +3Computer-Use Agent BenchmarksComputer-Use Agents

  10. Towards Mitigation of Hallucination for LLM-empowered Agents: Progressive Generalization Bound Exploration and Watchdog Monitor

    Jul 21, 2025Siyuan Liu, Wenjing Liu, Zhiwei Xu +3LLM Hallucination MitigationHallucination in Language Models

  11. A False Average: Pooled CoT-Monitor Accuracy Conceals a Reasoning-Dependent Fragility

    Date pendingShikhar Shiromani, Leo RichterReward HackingAdversarial Attacks on LLMs

  12. ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems

    Date pendingRakesh Sharma, Sydney Pugh, Cameron Beeche +14AI Agent SafetyAI Agent Monitoring

  13. Your Agent Says Yes: Interpreting Adversarial Market Behavior Beyond Individual Transactions

    Date pendingZelin Li, Yiyun Su, Matt White +3LLM Agent EvaluationAI Agent Security