AI Agent Safety

Momentum

22 papers in the last four weeks, up 144% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 207

All topics
CardsList
  1. The Authorization-Execution Gap Is a Major Safety and Security Problem in Open-World Agents

    May 10, 2026Baoyuan Wu, Qingshan Liu, Adel Bibi +2Agent Failure AnalysisAI Agent Security

  2. A Prompt-Aware Structuring Framework for Reliable Reuse of AI-Generated Content in the Agentic Web

    May 10, 2026Shusaku Egami, Masahiro HamasakiAI Agent SafetyLLM Agent Reliability

  3. Containment Verification: AI Safety Guarantees Independent of Alignment

    May 9, 2026Royce Moon, Lav R. VarshneyAI Agent SafetyFormal Verification

  4. Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI

    May 8, 2026Xuanqiang Angelo Huang, Charlie Tharas, Samuele Marro +4Mechanism DesignAI Agent Safety

  5. Why Does Agentic Safety Fail to Generalize Across Tasks?

    May 7, 2026Yonatan Slutzky, Yotam Alexander, Tomer Slor +2AI Agent SafetyLLM Agent Safety

  6. Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors

    May 7, 2026Jonas Wiedermann-Möller, Leonard Dung, Maksym AndriushchenkoAI Agent EvaluationAI Agent Safety

  7. Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges

    May 7, 2026Shihao Weng, Yang Feng, Xiaofei XieLLM-as-a-JudgeAI Agent Safety

  8. TRACE: A Metrologically-Grounded Engineering Framework for Trustworthy Agentic AI Systems in Operationally Critical Domains

    May 5, 2026Serhii ZabolotniiAI Agent SafetyAI Agent Governance

  9. An Empirical Study of Agent Skills for Healthcare: Practice, Gaps, and Governance

    May 4, 2026Gelei Xu, Ningzhi Tang, Xueyang Li +4HealthcareAI Agent Safety

  10. Executor-Side Progressive Risk-Gated Actuation for Agentic AI in Wireless Supervisory Control

    May 4, 2026Zhenyu Liu, Yi Ma, Rahim TafazolliSafety-Critical ControlAI Agent Safety

  11. MILD: Mediator Agent System with Bidirectional Perception and Multi-Layered Alignment for Human-Vehicle Collaboration

    May 2, 2026Jiyao Wang, Yunbiao Wang, Yubo Jiao +6AI AlignmentAI Agent Safety

  12. Position: Safety and Fairness in Agentic AI Depend on Interaction Topology, Not on Model Scale or Alignment

    May 1, 2026Tanav Singh Bajaj, Nikhil Singh, Karan Anand +1Opinion DynamicsAI Agent Safety

  13. AgenticAITA: A Proof-Of-Concept About Deliberative Multi-Agent Reasoning for Autonomous Trading Systems

    May 1, 2026Ivan LetteriMulti-Agent OrchestrationAI Agent Safety

  14. Causal Foundations of Collective Agency

    Apr 30, 2026Frederik Hytting Jørgensen, Sebastian Weichwald, Lewis HammondAI Agent SafetyCausal Abstraction

  15. Ambient Persuasion in a Deployed AI Agent: Unauthorized Escalation Following Routine Non-Adversarial Content Exposure

    Apr 29, 2026Diego F. Cuadros, Abdoul-Aziz MaigaAgent Failure AnalysisAI Agent Security

  16. Governing What You Cannot Observe: Adaptive Runtime Governance for Autonomous AI Agents

    Apr 27, 2026German Marin, Jatin ChaudharyAI Agent SafetyAI Agent Monitoring

  17. Evaluating whether AI models would sabotage AI safety research

    Apr 27, 2026Robert Kirk, Alexandra Souly, Kai Fronsdal +2Language Model Safety EvaluationLLM Auditing

  18. OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents

    Apr 27, 2026Zheng Wu, Yi Hua, Zhaoyuan Huang +8Computer-Use AgentsAI Agent Evaluation

  19. Benchmarking Web Agent Safety under E-commerce Deceptive Interfaces

    Apr 26, 2026Zijing Shi, Meng Fang, Ling ChenWeb Agent BenchmarksAI Agent Safety

  20. Vision-Language-Action Safety: Threats, Challenges, Evaluations, and Mechanisms

    Apr 26, 2026Qi Li, Bo Yin, Weiqi Huang +6AI Agent SafetyAdversarial Attacks on VLA Models

  21. Interval POMDP Shielding for Imperfect-Perception Agents

    Apr 22, 2026William Scarbro, Ravi MangalSafety-Critical ControlAI Agent Safety