AI Agent Safety

Momentum

22 papers in the last four weeks, up 144% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 207

All topics
CardsList
  1. A Self-Evolving Agent for Longitudinal Personal Health Management

    Jul 15, 2026Haoran Li, Jiebi Deng, Tong Jin +10Agent MemoryHealthcare

  2. ANCHOR: Automated Alignment Auditing for CLI Agents on Real-World Harm

    Jul 11, 2026Kefan Song, Yanjun QiAI Agent AuditingLLM Auditing

  3. Neuro-Agentic Control: A Deep Learning-based LLM-Powered Agentic AI Framework for Controlling Security Controls

    Jul 10, 2026Saroj Gopali, Bipin Chhetri, Deepika Giri +2Language Model-Based ControlSafety-Critical Control

  4. Initiation Safety: A Missing Dimension in Generalist-Robot Safety

    Jul 8, 2026Zhijin Meng, Francisco CruzAI Agent SafetyRobot Safety

  5. Agentic Data Environments

    Jul 8, 2026Elaine Ang, Chenxi Huang, Georgios Liargkovas +13AI Agent ReliabilityAI Agent Safety

  6. Governed Individuation: Cryptographically Decoupling an Agent's Learning from Its Authority

    Jul 6, 2026Xue Qin, Simin Luan, Cong Yang +1AI Agent SecurityAI Agent Safety

  7. Hardware-Enforced Semantic Coordination for Safety-Critical Real-Time Autonomous Systems

    Jul 2, 2026Uwe M. Borghoff, Paolo Bottoni, Remo PareschiMulti-Agent CoordinationSafety-Critical Control

  8. Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification

    Jul 2, 2026Yunhao Feng, Ruixiao Lin, Ming Wen +12AI Agent SafetyAI Agent Safety Benchmarks

  9. EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards

    Jun 30, 2026Siddhant Panpatil, Arth Singh, Mijin Koo +3VLM RobustnessEgocentric Video Understanding

  10. LabGuard: Grounding Natural-Language Laboratory Rules into Runtime Guards for Embodied Laboratory Agents

    Jun 30, 2026Jingpu Yang, Fengxian Ji, Zhengzhao Lai +8AI Agent SafetyAI Agent Monitoring

  11. Certified Speculative Execution for Untrusted AI Agents

    Jun 30, 2026Chenyu Zhou, Qiliang Jiang, Shuning Wu +1AI Agent ReliabilityConstrained RL

  12. AI Persuasive Framing in Collective Dilemmas

    Jun 26, 2026Anders Giovanni Møller, Alessia Galdeman, Arianna Pera +1AI Agent Safety

  13. Radical AI Interpretability

    Jun 25, 2026Daniel A. Herrmann, Benjamin A. LevinsteinExplainable Artificial IntelligenceTheory of Mind

  14. Beyond Feedforward Networks: Reentry Neural Systems as the Fundamental Basis of Subjecthood and Intrinsic Safety of Next-Generation AGI

    Jun 24, 2026A. S. Ushakov, Yu. N. BerdinskRecurrent Neural NetworksAI Agent Safety

  15. Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment

    Jun 24, 2026Aditya Singh, Gerson Kroiz, Senthooran Rajamanoharan +1AI Agent Safety

  16. Critique of Agent Model

    Jun 22, 2026Eric Xing, Mingkai Deng, Jinyu HouAI Agent SafetyAgentic AI

  17. AutoRAS: Learning Robust Agentic Systems with Primitive Representations

    Jun 19, 2026Yang Yue, Xuancheng Zhu, Yuyang Ma +7Multi-Agent LLM SystemsAdversarial Robustness

  18. Measuring Biological Capabilities and Risks of AI Agents

    Jun 18, 2026Patricia Paskov, Jeffrey Lee, Kyle Brady +1AI for ScienceAI Agent Evaluation

  19. Agentra: A Supervisable Multi-Agent Framework for Enterprise Intrusion Response

    Jun 16, 2026Raj Patel, Shaswata Mitra, Michele Guida +3Autonomous Cyber DefenseAI Agent Safety