AI Agent Safety

Momentum

22 papers in the last four weeks, up 175% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 201

All topics
CardsList
  1. Benchmarking Jailbreak Guardrails for Embodied Agents

    Oct 5, 2026Xunguang Wang, Qingyue Wang, Yuguang Zhou +3LLM GuardrailsAI Agent Safety

  2. G-CARB: Graph-Localized Conformal Agent Risk Budget for Compositional Harm

    Oct 4, 2026Zijun Yu, Yu Gu, Vahid Partovi Nia +1AI Agent SafetyConformal Risk Control

  3. Chaining Skills to Hijack LLM Agents

    Oct 1, 2026Tian Dong, Zixuan Ma, Haodong Zhao +3LLM Agent SecurityAI Agent Safety

  4. Safety of Latent Communication in Multi-Agent Systems

    Sep 30, 2026Muhammad Huzaifa, Sina Mavali, Thorsten EisenhoferMulti-Agent LLM SystemsAI Agent Safety

  5. Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents

    Sep 30, 2026Serhii ZabolotniiAI Agent SafetyRuntime Enforcement for AI Agents

  6. A Competing-Hazards Systematization of Loss of Control in Autonomous Agents

    Sep 29, 2026Mohamed Aly BoukeAgent Failure AnalysisAI Agent Auditing

  7. Character Training for Risk-Averse Agents

    Sep 29, 2026Arav Dhoot, Punya Syon Pandey, Jamie Johnson +3Risk-Sensitive RLAI Agent Safety

  8. SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety

    Sep 28, 2026Jianxing Chen, Xiao Yu, Shipra Agrawal +1AI Agent SafetyAI Agent Safety Benchmarks

  9. SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents

    Sep 28, 2026Saswat Das, Parvati Viswanathan, Daniel Donnelly +3AI Agent SafetyAI Agent Safety Benchmarks

  10. Before Acting, Change the State: Prospective State Intervention for Web Agents under Deceptive Interfaces

    Sep 28, 2026Ruozhao Yang, Mingfei Cheng, Xiaofei XieAI Agent SafetyRuntime Enforcement for AI Agents

  11. AdaGuard: An Adaptive Guard Model with User-defined Policies

    Sep 28, 2026Yunhao Feng, Yifan Ding, Yuxiang Xie +4LLM GuardrailsAI Agent Safety

  12. Shutdown Sabotage Propensities in Multi-Agent Systems

    Sep 23, 2026Amelie Knecht, Ulysse Schaller, Christopher Summerfield +1AI Agent SafetyMulti-Agent System Security

  13. Finite-Sample Probabilistic Safety Certification for AI-Based Grid-Edge Coordination

    Sep 23, 2026Yihong Zhou, Hanbin Yang, Thomas MorstynAI Agent Safety

  14. Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness

    Sep 22, 2026Varshini Elangovan, James Wedgwood, Chhavi Yadav +3AI Agent SafetyHuman-AI Interaction

  15. ClashBench: Conflicts Leading Agents to Seize and Harm

    Sep 17, 2026Yuejin Xie, Yu Li, Dadi Guo +6AI Agent SecurityAI Agent Safety

  16. HazardAuditor: From Executable Threats to Safer Computer-Use Agents

    Sep 14, 2026Yunhao Feng, Ruixiao Lin, Ming Wen +5Language Model Safety EvaluationAI Agent Auditing

  17. Safety Signals to Verify NetOps Agents with Action-Level Granularity

    Sep 13, 2026Tobias Labarta, Frederik Pahde, Novak Boškov +5Long-Horizon Agent EvaluationAI Agent Safety

  18. Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents

    Sep 3, 2026Zhaoyuan Huang, Tianjie Ju, Pengzhou Cheng +7GUI AgentsAI Agent Safety

  19. PACE: Towards Surfacing Hidden Conflicts in User Requests

    Sep 3, 2026Yoojin Kim, Jihyoung Jang, Hyounghun KimAI Agent SafetyLLM Refusal Behavior

  20. Beyond Task Completion: Training Capable and Safe Computer-Use Agents

    Aug 27, 2026Zeyu Kang, Zhenyun Yin, Yang Zhang +5Computer-Use AgentsAI Agent Safety

  21. AI Agents Push Humans Out of the Loop

    Aug 24, 2026Margaret Mitchell, Avijit Ghosh, Samir PassiAI Agent SafetyHuman-in-the-Loop AI

  22. Agent Safety Should Be a Runtime Contract

    Aug 11, 2026Albus W. Ng, Yi Han, Jusheng Zhang +1AI Agent SafetyAI Agent Monitoring

  23. Multi-Agent AI Safety as an Institutional Design Problem

    Aug 10, 2026Abdullah XAI Agent SafetyMulti-Agent Systems

  24. ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents

    Aug 10, 2026Hongwei Yao, Yiming Liu, Meihui Chen +6AI Agent SecurityAI Agent Safety

  25. Software Engineering for and with GUI Agent

    Aug 10, 2026Shengcheng Yu, Yuchen Ling, Junyang Xing +3AI Agent ReliabilityGUI Agents

  26. Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents

    Aug 9, 2026Gabriele La Malfa, Lakmal Meegahapola, Edyta Bogucka +4AI Agent SafetyHuman-AI Collaboration

  27. Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution

    Aug 8, 2026Anton Razzhigaev, Andrei Gritsaev, Andrei Kaznacheev +3AI Coding AgentsAgent Harness Optimization

  28. StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection

    Aug 6, 2026Zhuoxin Zhan, Akbar Rafiey, Avery Ma +2Computer-Use AgentsAI Agent Safety

  29. Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents

    Aug 6, 2026Praphul Chandra, Sujit Gujar, Ganesh GhalmeAI Agent SafetyAI Agent Governance

  30. SafeCommit: Certifying When Memory-Grounded Agents May Safely Act

    Aug 4, 2026Mayur Akewar, Ravi RanjanAgent MemoryAI Agent Reliability

  31. Securing Agentic AI: From Per-Action Checks to Trajectory Assurance

    Aug 3, 2026Alireza Lotfi, Subangkar Karmaker Shanto, Imtiaz Karim +1AI Agent SecurityLLM Agent Security

  32. Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races

    Aug 2, 2026Phu Hoa Pham, Duy Minh Dao Sy, Trung Kiet Huynh +9LLM Agent EvaluationAI Agent Safety

  33. When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems

    Aug 1, 2026Neha Nagaraja, Amisha Bagari, Hayretdin BahsiMulti-Agent LLM SystemsLarge Language Model-Based Robot Planning

  34. Beyond Component Testing: Validating Agentic AI Systems

    Jul 31, 2026Fabio Orazio Mirto, Luca D'Agati, Giuseppe Tricomi +4AI Agent EvaluationAI Agent Safety

  35. LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents

    Jul 30, 2026Jingya Wang, Yuyang Gao, Liuzhenghao Lv +2Agent MemoryAI Agents for Scientific Discovery

  36. Towards a Systems Foundation for Agentic Cloud Management

    Jul 28, 2026Minghao Li, Ziqian Liu, Ziyu Mao +5Multi-Agent OrchestrationAI Agent Safety

  37. PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

    Jul 28, 2026Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou +11AI Agent EvaluationAI Agent Safety

  38. Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response

    Jul 28, 2026Abu Bakar SiddikAI Agent EvaluationCybersecurity

  39. SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems

    Jul 28, 2026Haowen Dai, Zonghao Ying, Wenfeng Li +10AI Agent SafetyInformation Flow Control

  40. GuardianAgentBench: Where Agents Fail and How to Guard Them

    Jul 23, 2026Vishal Ishwar Naik, Chenyu Xu, Donna Dong +5AI Agent ReliabilityLLM Guardrails

  41. OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills

    Jul 22, 2026Qiyuan Liu, Tingfeng Hui, Kun Zhan +2LLM Agent EvaluationAI Agent Safety

  42. JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety

    Jul 22, 2026Yuan Xiong, Linji Hao, Shizhu He +2AI Agent SafetyLLM Agent Safety

  43. Engineering Trustworthy Agentic AI for Critical Systems

    Jul 20, 2026Omar Al-Refai, Ibrahim Shahbaz, Adam Ali Husseinat +3AI AssuranceAgent Reliability

  44. Operational Hallucination and Safety Drift in AI Agents

    Jul 20, 2026Shasha Yu, Fiona Carroll, Barry L. BentleyLLM Agent EvaluationAI Agent Safety

  45. Specifying the Delegated-Autonomy Boundary: Requirements Engineering for Agentic AI

    Jul 19, 2026Chetan Arora, Andreas Vogelsang, Abbi SharmaAI Agent SafetyAI Agent Governance

  46. PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution

    Jul 18, 2026Yang Liu, Weixing Chen, Xinshuai Song +8LLM Agent OrchestrationAI Agent Safety

  47. A Self-Evolving Agent for Longitudinal Personal Health Management

    Jul 15, 2026Haoran Li, Jiebi Deng, Tong Jin +10Agent MemoryHealthcare

  48. ANCHOR: Automated Alignment Auditing for CLI Agents on Real-World Harm

    Jul 11, 2026Kefan Song, Yanjun QiAI Agent AuditingLLM Auditing

  49. Neuro-Agentic Control: A Deep Learning-based LLM-Powered Agentic AI Framework for Controlling Security Controls

    Jul 10, 2026Saroj Gopali, Bipin Chhetri, Deepika Giri +2Language Model-Based ControlSafety-Critical Control