AI Agent Safety

Momentum

22 papers in the last four weeks, up 144% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 207

All topics
CardsList
  1. AgentSOC: A Multi-Layer Agentic AI Framework for Security Operations Automation

    Apr 22, 2026Joyjit Roy, Samaresh Kumar SinghAutonomous Cyber DefenseCybersecurity

  2. SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models

    Apr 21, 2026Josue Torres-Fonseca, Naihao Deng, Yinpei Dai +5LLM Safety BenchmarksLarge Language Model-Based Robot Planning

  3. AI-Native Network Controller: A Modular Framework for Safe Agentic Control of Multi-Domain Network Infrastructure

    Apr 20, 2026Merim DzaferagicAI Agent Safety

  4. Owner-Harm: A Missing Threat Model for AI Agent Safety

    Apr 20, 2026Dongcheng Zhang, Yiqing JiangAI Agent SecurityAI Agent Safety

  5. From Craft to Kernel: A Governance-First Execution Architecture and Semantic ISA for Agentic Computers

    Apr 20, 2026Xiangyu Wen, Yuang Zhao, Xiaoyu Xu +9AI Agent SecurityAI Agent Safety

  6. Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs

    Apr 19, 2026Carissa Cullen, Harry Garland, Alexander Roman +3Agent ReliabilityAI Agent Safety

  7. Governance by Design: Architecting Agentic AI for Organizational Learning and Scalable Autonomy

    Apr 17, 2026Nelly Dux, Cristina Alaimo, Philippe Roussiere +1AI AccountabilityAI Agent Safety

  8. Agentic Microphysics: A Manifesto for Generative AI Safety

    Apr 16, 2026Federico Pierucci, Matteo Prandi, Marcantonio Bracale Syrnikov +2AI Agent SafetyAI Safety

  9. HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark

    Apr 15, 2026Jiacheng Wang, Jinchang Hou, Fabian Wang +3AI Agent AuditingLong-Horizon Agent Evaluation

  10. AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents

    Apr 3, 2026Yunhao Feng, Yifan Ding, Yingshui Tan +6LLM Safety BenchmarksComputer-Use Agent Benchmarks

  11. Peer-Preservation in Frontier Models

    Mar 30, 2026Yujin Potter, Nicholas Crispino, Vincent Siu +2AI Agent SafetyLLM Safety Evaluation

  12. The Alignment Flywheel: A Governance-Centric Hybrid MAS for Architecture-Agnostic Safety

    Feb 28, 2026Elias Malomgré, Pieter SimoensAI Agent SafetyAI Agent Monitoring

  13. When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents

    Feb 9, 2026Yuting Ning, Jaylen Jones, Zhehao Zhang +5LLM Self-CorrectionComputer-Use Agents

  14. Agentic AI for Commercial Insurance Underwriting with Adversarial Self-Critique

    Jan 21, 2026Joyjit Roy, Samaresh Kumar SinghLLM Self-CorrectionAI Agent Safety

  15. When Assisting One Disempowers Another

    Nov 6, 2025Claire Yang, Claire Jie Zhang, Maya Cakmak +1AI Agent SafetyHuman-Centered AI

  16. Measuring Harmfulness of Computer-Using Agents

    Jul 31, 2025Aaron Xuxiang Tian, Ruofan Zhang, Janet Tang +3Computer-Use Agent BenchmarksComputer-Use Agents

  17. Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity?

    Date pendingAnmol Goel, Iryna GurevychData LeakageContextual Integrity

  18. ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems

    Date pendingRakesh Sharma, Sydney Pugh, Cameron Beeche +14AI Agent SafetyAI Agent Monitoring

  19. Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best

    Date pendingKevin Baum, Rūta Binkytė, Felix JahnAI AlignmentLLM Safety Alignment