AI Agent Safety

Momentum

22 papers in the last four weeks, up 144% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 207

All topics
CardsList
  1. AI Safety Considerations for Agents With Limited Time to Act

    Oct 7, 2026Leo Zeitler, Jack Richings, Victoria NocklesAI Agent SafetyAI Safety

  2. RT-Safe: Benchmarking Agent Safety in Real-Time Embodied Environment

    Oct 7, 2026Tianruo Rose Xu, Jiawei Ren, Yichi Yang +2AI Agent SafetySafe Robot Navigation

  3. SpecGuard: Proving a Task Is Broken Before the Agent Cheats

    Oct 6, 2026Param Biyani, Krishnamurthy DvijothamAI Agent SafetySoftware Engineering Agents

  4. Learning to Report Unsafe Tasks in a Multi-Agent Game

    Oct 6, 2026Avyay M. Casheekar, Hariganesh TangiralaAI Agent SafetyMulti-Agent Systems

  5. Toward Evidence-Driven Human-Agent-Robot Teaming for Earth-Independent Anomaly Triage

    Oct 6, 2026Ignacio G Lopez-Francos, Alexis Gallagher, Samira ShalalHuman-Robot CollaborationSpace Robotics

  6. Benchmarking Jailbreak Guardrails for Embodied Agents

    Oct 5, 2026Xunguang Wang, Qingyue Wang, Yuguang Zhou +3LLM GuardrailsAI Agent Safety

  7. G-CARB: Graph-Localized Conformal Agent Risk Budget for Compositional Harm

    Oct 4, 2026Zijun Yu, Yu Gu, Vahid Partovi Nia +1AI Agent SafetyConformal Risk Control

  8. Chaining Skills to Hijack LLM Agents

    Oct 1, 2026Tian Dong, Zixuan Ma, Haodong Zhao +3LLM Agent SecurityAI Agent Safety

  9. Safety of Latent Communication in Multi-Agent Systems

    Sep 30, 2026Muhammad Huzaifa, Sina Mavali, Thorsten EisenhoferMulti-Agent LLM SystemsAI Agent Safety

  10. Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents

    Sep 30, 2026Serhii ZabolotniiAI Agent SafetyRuntime Enforcement for AI Agents

  11. A Competing-Hazards Systematization of Loss of Control in Autonomous Agents

    Sep 29, 2026Mohamed Aly BoukeAgent Failure AnalysisAI Agent Auditing

  12. Character Training for Risk-Averse Agents

    Sep 29, 2026Arav Dhoot, Punya Syon Pandey, Jamie Johnson +3Risk-Sensitive RLAI Agent Safety

  13. SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety

    Sep 28, 2026Jianxing Chen, Xiao Yu, Shipra Agrawal +1AI Agent SafetyAI Agent Safety Benchmarks

  14. SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents

    Sep 28, 2026Saswat Das, Parvati Viswanathan, Daniel Donnelly +3AI Agent SafetyAI Agent Safety Benchmarks

  15. Before Acting, Change the State: Prospective State Intervention for Web Agents under Deceptive Interfaces

    Sep 28, 2026Ruozhao Yang, Mingfei Cheng, Xiaofei XieAI Agent SafetyRuntime Enforcement for AI Agents

  16. AdaGuard: An Adaptive Guard Model with User-defined Policies

    Sep 28, 2026Yunhao Feng, Yifan Ding, Yuxiang Xie +4LLM GuardrailsAI Agent Safety

  17. Shutdown Sabotage Propensities in Multi-Agent Systems

    Sep 23, 2026Amelie Knecht, Ulysse Schaller, Christopher Summerfield +1AI Agent SafetyMulti-Agent System Security

  18. Finite-Sample Probabilistic Safety Certification for AI-Based Grid-Edge Coordination

    Sep 23, 2026Yihong Zhou, Hanbin Yang, Thomas MorstynAI Agent Safety

  19. Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness

    Sep 22, 2026Varshini Elangovan, James Wedgwood, Chhavi Yadav +3AI Agent SafetyHuman-AI Interaction

  20. ClashBench: Conflicts Leading Agents to Seize and Harm

    Sep 17, 2026Yuejin Xie, Yu Li, Dadi Guo +6AI Agent SecurityAI Agent Safety

  21. HazardAuditor: From Executable Threats to Safer Computer-Use Agents

    Sep 14, 2026Yunhao Feng, Ruixiao Lin, Ming Wen +5Language Model Safety EvaluationAI Agent Auditing

  22. Safety Signals to Verify NetOps Agents with Action-Level Granularity

    Sep 13, 2026Tobias Labarta, Frederik Pahde, Novak Boškov +5Long-Horizon Agent EvaluationAI Agent Safety

  23. Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents

    Sep 3, 2026Zhaoyuan Huang, Tianjie Ju, Pengzhou Cheng +7GUI AgentsAI Agent Safety

  24. PACE: Towards Surfacing Hidden Conflicts in User Requests

    Sep 3, 2026Yoojin Kim, Jihyoung Jang, Hyounghun KimAI Agent SafetyLLM Refusal Behavior

  25. Beyond Task Completion: Training Capable and Safe Computer-Use Agents

    Aug 27, 2026Zeyu Kang, Zhenyun Yin, Yang Zhang +5Computer-Use AgentsAI Agent Safety

  26. AI Agents Push Humans Out of the Loop

    Aug 24, 2026Margaret Mitchell, Avijit Ghosh, Samir PassiAI Agent SafetyHuman-in-the-Loop AI

  27. OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets

    Aug 14, 2026Dongsheng Chen, Xiangyu Zhao, Xin Yao +1AI Agent SafetyAI Agent Governance

  28. Agent Safety Should Be a Runtime Contract

    Aug 11, 2026Albus W. Ng, Yi Han, Jusheng Zhang +1AI Agent SafetyAI Agent Monitoring

  29. Multi-Agent AI Safety as an Institutional Design Problem

    Aug 10, 2026Abdullah XAI Agent SafetyMulti-Agent Systems

  30. ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents

    Aug 10, 2026Hongwei Yao, Yiming Liu, Meihui Chen +6AI Agent SecurityAI Agent Safety

  31. Software Engineering for and with GUI Agent

    Aug 10, 2026Shengcheng Yu, Yuchen Ling, Junyang Xing +3AI Agent ReliabilityGUI Agents

  32. Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents

    Aug 9, 2026Gabriele La Malfa, Lakmal Meegahapola, Edyta Bogucka +4AI Agent SafetyHuman-AI Collaboration

  33. Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution

    Aug 8, 2026Anton Razzhigaev, Andrei Gritsaev, Andrei Kaznacheev +3AI Coding AgentsAgent Harness Optimization

  34. StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection

    Aug 6, 2026Zhuoxin Zhan, Akbar Rafiey, Avery Ma +2Computer-Use AgentsAI Agent Safety

  35. Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents

    Aug 6, 2026Praphul Chandra, Sujit Gujar, Ganesh GhalmeAI Agent SafetyAI Agent Governance

  36. SafeCommit: Certifying When Memory-Grounded Agents May Safely Act

    Aug 4, 2026Mayur Akewar, Ravi RanjanAgent MemoryAI Agent Reliability

  37. Securing Agentic AI: From Per-Action Checks to Trajectory Assurance

    Aug 3, 2026Alireza Lotfi, Subangkar Karmaker Shanto, Imtiaz Karim +1AI Agent SecurityLLM Agent Security

  38. Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races

    Aug 2, 2026Phu Hoa Pham, Duy Minh Dao Sy, Trung Kiet Huynh +9LLM Agent EvaluationAI Agent Safety

  39. When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems

    Aug 1, 2026Neha Nagaraja, Amisha Bagari, Hayretdin BahsiMulti-Agent LLM SystemsLarge Language Model-Based Robot Planning

  40. Beyond Component Testing: Validating Agentic AI Systems

    Jul 31, 2026Fabio Orazio Mirto, Luca D'Agati, Giuseppe Tricomi +4AI Agent EvaluationAI Agent Safety

  41. LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents

    Jul 30, 2026Jingya Wang, Yuyang Gao, Liuzhenghao Lv +2Agent MemoryAI Agents for Scientific Discovery

  42. Towards a Systems Foundation for Agentic Cloud Management

    Jul 28, 2026Minghao Li, Ziqian Liu, Ziyu Mao +5Multi-Agent OrchestrationAI Agent Safety

  43. PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

    Jul 28, 2026Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou +11AI Agent EvaluationAI Agent Safety

  44. Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response

    Jul 28, 2026Abu Bakar SiddikAI Agent EvaluationCybersecurity

  45. SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems

    Jul 28, 2026Haowen Dai, Zonghao Ying, Wenfeng Li +10AI Agent SafetyInformation Flow Control

  46. GuardianAgentBench: Where Agents Fail and How to Guard Them

    Jul 23, 2026Vishal Ishwar Naik, Chenyu Xu, Donna Dong +5AI Agent ReliabilityLLM Guardrails

  47. OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills

    Jul 22, 2026Qiyuan Liu, Tingfeng Hui, Kun Zhan +2LLM Agent EvaluationAI Agent Safety

  48. JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety

    Jul 22, 2026Yuan Xiong, Linji Hao, Shizhu He +2AI Agent SafetyLLM Agent Safety

  49. Engineering Trustworthy Agentic AI for Critical Systems

    Jul 20, 2026Omar Al-Refai, Ibrahim Shahbaz, Adam Ali Husseinat +3AI AssuranceAgent Reliability

  50. Operational Hallucination and Safety Drift in AI Agents

    Jul 20, 2026Shasha Yu, Fiona Carroll, Barry L. BentleyLLM Agent EvaluationAI Agent Safety

  51. Specifying the Delegated-Autonomy Boundary: Requirements Engineering for Agentic AI

    Jul 19, 2026Chetan Arora, Andreas Vogelsang, Abbi SharmaAI Agent SafetyAI Agent Governance

  52. PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution

    Jul 18, 2026Yang Liu, Weixing Chen, Xinshuai Song +8LLM Agent OrchestrationAI Agent Safety