AI Agent Safety

Momentum

22 papers in the last four weeks, up 144% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 207

All topics
CardsList
  1. Beyond Task Completion: Training Capable and Safe Computer-Use Agents

    Aug 27, 2026Zeyu Kang, Zhenyun Yin, Yang Zhang +5Computer-Use AgentsAI Agent Safety

  2. AI Agents Push Humans Out of the Loop

    Aug 24, 2026Margaret Mitchell, Avijit Ghosh, Samir PassiAI Agent SafetyHuman-in-the-Loop AI

  3. OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets

    Aug 14, 2026Dongsheng Chen, Xiangyu Zhao, Xin Yao +1AI Agent SafetyAI Agent Governance

  4. Agent Safety Should Be a Runtime Contract

    Aug 11, 2026Albus W. Ng, Yi Han, Jusheng Zhang +1AI Agent SafetyAI Agent Monitoring

  5. Multi-Agent AI Safety as an Institutional Design Problem

    Aug 10, 2026Abdullah XAI Agent SafetyMulti-Agent Systems

  6. ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents

    Aug 10, 2026Hongwei Yao, Yiming Liu, Meihui Chen +6AI Agent SecurityAI Agent Safety

  7. Software Engineering for and with GUI Agent

    Aug 10, 2026Shengcheng Yu, Yuchen Ling, Junyang Xing +3AI Agent ReliabilityGUI Agents

  8. Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents

    Aug 9, 2026Gabriele La Malfa, Lakmal Meegahapola, Edyta Bogucka +4AI Agent SafetyHuman-AI Collaboration

  9. Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution

    Aug 8, 2026Anton Razzhigaev, Andrei Gritsaev, Andrei Kaznacheev +3AI Coding AgentsAgent Harness Optimization

  10. StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection

    Aug 6, 2026Zhuoxin Zhan, Akbar Rafiey, Avery Ma +2Computer-Use AgentsAI Agent Safety

  11. Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents

    Aug 6, 2026Praphul Chandra, Sujit Gujar, Ganesh GhalmeAI Agent SafetyAI Agent Governance

  12. SafeCommit: Certifying When Memory-Grounded Agents May Safely Act

    Aug 4, 2026Mayur Akewar, Ravi RanjanAgent MemoryAI Agent Reliability

  13. Securing Agentic AI: From Per-Action Checks to Trajectory Assurance

    Aug 3, 2026Alireza Lotfi, Subangkar Karmaker Shanto, Imtiaz Karim +1AI Agent SecurityLLM Agent Security

  14. Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races

    Aug 2, 2026Phu Hoa Pham, Duy Minh Dao Sy, Trung Kiet Huynh +9LLM Agent EvaluationAI Agent Safety

  15. When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems

    Aug 1, 2026Neha Nagaraja, Amisha Bagari, Hayretdin BahsiMulti-Agent LLM SystemsLarge Language Model-Based Robot Planning

  16. Beyond Component Testing: Validating Agentic AI Systems

    Jul 31, 2026Fabio Orazio Mirto, Luca D'Agati, Giuseppe Tricomi +4AI Agent EvaluationAI Agent Safety

  17. LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents

    Jul 30, 2026Jingya Wang, Yuyang Gao, Liuzhenghao Lv +2Agent MemoryAI Agents for Scientific Discovery

  18. Towards a Systems Foundation for Agentic Cloud Management

    Jul 28, 2026Minghao Li, Ziqian Liu, Ziyu Mao +5Multi-Agent OrchestrationAI Agent Safety

  19. PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

    Jul 28, 2026Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou +11AI Agent EvaluationAI Agent Safety

  20. Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response

    Jul 28, 2026Abu Bakar SiddikAI Agent EvaluationCybersecurity

  21. SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems

    Jul 28, 2026Haowen Dai, Zonghao Ying, Wenfeng Li +10AI Agent SafetyInformation Flow Control

  22. GuardianAgentBench: Where Agents Fail and How to Guard Them

    Jul 23, 2026Vishal Ishwar Naik, Chenyu Xu, Donna Dong +5AI Agent ReliabilityLLM Guardrails

  23. OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills

    Jul 22, 2026Qiyuan Liu, Tingfeng Hui, Kun Zhan +2LLM Agent EvaluationAI Agent Safety

  24. JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety

    Jul 22, 2026Yuan Xiong, Linji Hao, Shizhu He +2AI Agent SafetyLLM Agent Safety

  25. Engineering Trustworthy Agentic AI for Critical Systems

    Jul 20, 2026Omar Al-Refai, Ibrahim Shahbaz, Adam Ali Husseinat +3AI AssuranceAgent Reliability

  26. Operational Hallucination and Safety Drift in AI Agents

    Jul 20, 2026Shasha Yu, Fiona Carroll, Barry L. BentleyLLM Agent EvaluationAI Agent Safety

  27. Specifying the Delegated-Autonomy Boundary: Requirements Engineering for Agentic AI

    Jul 19, 2026Chetan Arora, Andreas Vogelsang, Abbi SharmaAI Agent SafetyAI Agent Governance

  28. PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution

    Jul 18, 2026Yang Liu, Weixing Chen, Xinshuai Song +8LLM Agent OrchestrationAI Agent Safety