AI Agent Safety

Momentum

22 papers in the last four weeks, up 144% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 207

All topics
CardsList
  1. A Self-Evolving Agent for Longitudinal Personal Health Management

    Jul 15, 2026Haoran Li, Jiebi Deng, Tong Jin +10Agent MemoryHealthcare

  2. ANCHOR: Automated Alignment Auditing for CLI Agents on Real-World Harm

    Jul 11, 2026Kefan Song, Yanjun QiAI Agent AuditingLLM Auditing

  3. Neuro-Agentic Control: A Deep Learning-based LLM-Powered Agentic AI Framework for Controlling Security Controls

    Jul 10, 2026Saroj Gopali, Bipin Chhetri, Deepika Giri +2Language Model-Based ControlSafety-Critical Control

  4. Initiation Safety: A Missing Dimension in Generalist-Robot Safety

    Jul 8, 2026Zhijin Meng, Francisco CruzAI Agent SafetyRobot Safety

  5. Agentic Data Environments

    Jul 8, 2026Elaine Ang, Chenxi Huang, Georgios Liargkovas +13AI Agent ReliabilityAI Agent Safety

  6. Governed Individuation: Cryptographically Decoupling an Agent's Learning from Its Authority

    Jul 6, 2026Xue Qin, Simin Luan, Cong Yang +1AI Agent SecurityAI Agent Safety

  7. Hardware-Enforced Semantic Coordination for Safety-Critical Real-Time Autonomous Systems

    Jul 2, 2026Uwe M. Borghoff, Paolo Bottoni, Remo PareschiMulti-Agent CoordinationSafety-Critical Control

  8. Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification

    Jul 2, 2026Yunhao Feng, Ruixiao Lin, Ming Wen +12AI Agent SafetyAI Agent Safety Benchmarks

  9. EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards

    Jun 30, 2026Siddhant Panpatil, Arth Singh, Mijin Koo +3VLM RobustnessEgocentric Video Understanding

  10. LabGuard: Grounding Natural-Language Laboratory Rules into Runtime Guards for Embodied Laboratory Agents

    Jun 30, 2026Jingpu Yang, Fengxian Ji, Zhengzhao Lai +8AI Agent SafetyAI Agent Monitoring

  11. Certified Speculative Execution for Untrusted AI Agents

    Jun 30, 2026Chenyu Zhou, Qiliang Jiang, Shuning Wu +1AI Agent ReliabilityConstrained RL

  12. AI Persuasive Framing in Collective Dilemmas

    Jun 26, 2026Anders Giovanni Møller, Alessia Galdeman, Arianna Pera +1AI Agent Safety

  13. Radical AI Interpretability

    Jun 25, 2026Daniel A. Herrmann, Benjamin A. LevinsteinExplainable Artificial IntelligenceTheory of Mind

  14. Beyond Feedforward Networks: Reentry Neural Systems as the Fundamental Basis of Subjecthood and Intrinsic Safety of Next-Generation AGI

    Jun 24, 2026A. S. Ushakov, Yu. N. BerdinskRecurrent Neural NetworksAI Agent Safety

  15. Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment

    Jun 24, 2026Aditya Singh, Gerson Kroiz, Senthooran Rajamanoharan +1AI Agent Safety

  16. Critique of Agent Model

    Jun 22, 2026Eric Xing, Mingkai Deng, Jinyu HouAI Agent SafetyAgentic AI

  17. AutoRAS: Learning Robust Agentic Systems with Primitive Representations

    Jun 19, 2026Yang Yue, Xuancheng Zhu, Yuyang Ma +7Multi-Agent LLM SystemsAdversarial Robustness

  18. Measuring Biological Capabilities and Risks of AI Agents

    Jun 18, 2026Patricia Paskov, Jeffrey Lee, Kyle Brady +1AI for ScienceAI Agent Evaluation

  19. Agentra: A Supervisable Multi-Agent Framework for Enterprise Intrusion Response

    Jun 16, 2026Raj Patel, Shaswata Mitra, Michele Guida +3Autonomous Cyber DefenseAI Agent Safety

  20. Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models

    Jun 16, 2026Jasmine Brazilek, Joel Christoph, Maheep Chaudhary +4AI Agent EvaluationAI Agent Safety

  21. When Agent Automation Becomes Profitable: Quantifying and Insuring Autonomous AI Risk through Trace-Economic Underwriting

    Jun 15, 2026Binyan Xu, Xilin Dai, Fan Yang +1AI Risk ManagementAI Agent Safety

  22. An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios

    Jun 15, 2026Hankyul Baek, Jaewon Noh, Sang Seo +5Data LeakageLLM Agent Evaluation

  23. The Perils of Agency: How Developers Perceive, Prioritize, and Address Risks in Agentic AI Products

    Jun 13, 2026Hao-Ping Lee, Jessica He, David Piorkowski +3AI Risk ManagementAI Agent Safety

  24. Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment

    Jun 13, 2026Lipeng He, Yihan Wang, Jiawen Zhang +1Adversarial TrainingAI Agent Safety

  25. ClayBuddy: A Framework, Evaluation, & Mitigation of Coding Agent Failures

    Jun 13, 2026Kenneth Ge, Andre AssisAgent Failure AnalysisAI Coding Agents

  26. OSGuard: A Benchmark for Safety in Computer-Use Agents

    Jun 13, 2026Mina Mohammadmirzaei, Jeffrey FlaniganComputer-Use AgentsLLM Guardrails

  27. Selective Agentic Recovery for UAV Autonomy with a Persistent Mission Runtime

    Jun 12, 2026Taewoo Park, Kyeonghyun Yoo, Seunghyun Yoo +1UAV NavigationRobot Failure Recovery

  28. Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH

    Jun 12, 2026Vikhyath Kothamasu, Virginia Smith, Chhavi YadavAI Agent SafetyAI Agent Safety Benchmarks

  29. A Virtuous AI is an Existential Risk

    Jun 11, 2026Guillermo Del Pinal, Youngchan Lee, Min OhnLLM Safety AlignmentAI Agent Safety

  30. The Containment Gap: How Deployed Agentic AI Frameworks Fail Public-Facing Safety Requirements

    Jun 11, 2026Md Jafrin Hossain, Mohammad Arif Hossain, Weiqi Liu +1AI Agent SecurityAI Agent Safety

  31. The Impossibility of Eliciting Latent Knowledge

    Jun 10, 2026Korbinian Friedl, Francis Rhys Ward, Paul Yushin Rapoport +2AI Agent SafetyLLM Honesty

  32. Towards Responsibly Non-Compliant Machines

    Jun 10, 2026Marija Slavkovik, Marie Farrell, Louise Dennis +3AI AccountabilityAI Agent Safety

  33. SafeGEO: Understanding Generative Engine Optimization Risks in Recommendation Agents

    Jun 8, 2026Qianfeng Wen, Yifan Simon Liu, Xin Liu +4Generative Engine OptimizationLLM Agent Security

  34. Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems

    Jun 5, 2026Yiyang Zhao, Zhuo Zhang, Qingxuan Le +2Multi-Agent LLM SystemsAI Agent Safety

  35. From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents

    Jun 4, 2026Patrick Wilhelm, Odej KaoReward HackingAI Agent Safety

  36. PreAct-Bench: Benchmarking Predictive Monitoring in LLMs

    Jun 3, 2026Hainiu Xu, Italo Luis da Silva, Jiangnan Ye +7LLM Safety BenchmarksMoral Reasoning in Language Models

  37. Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety

    Jun 3, 2026Catherine Ge-Wang, Tyler Crosse, Benjamin Hadad +3Adversarial AttacksAI Agent Safety

  38. PerceptTwin: Semantic Scene Reconstruction for Iterative LLM Planning and Verification

    Jun 2, 2026Charlie Gauthier, Sacha Morin, Liam PaullLarge Language Model-Based Robot PlanningRobotics Simulation

  39. Agentic Safety is an Epistemic Property, Not a Behavioral One

    Jun 2, 2026Charles L. Wang, Keir Dorchen, Peter JinAI AlignmentAI Agent Safety

  40. VLESA: Vision-Language Embodied Safety Agent for Human Activity Monitoring

    Jun 2, 2026Hanjiang Hu, Yiyuan Pan, Jiaxing Li +5Egocentric Video UnderstandingAI Agent Safety

  41. RUBAS: Rubric-Based Reinforcement Learning for Agent Safety

    Jun 2, 2026Xian Qi Loye, Qinglin Su, Zhexin Zhang +5Reinforcement LearningAI Agent Safety

  42. Solipsistic Superintelligence is Unlikely to be Cooperative

    Jun 2, 2026Rakshit S Trivedi, Natasha Jaques, Logan Cross +2AI AlignmentMulti-Agent Collaboration

  43. Glass Box at Orbit: A Constitutional AI Verification Framework for Trustworthy Autonomous CubeSat Intelligence

    Jun 2, 2026Karthik Barma, Anil Sanneboyina, V C Premchand YadavAI AlignmentAI Agent Safety