LLM Agent Safety

LLM: Large Language Model

Momentum

24 papers in the last four weeks, up 50% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 193

All topics
CardsList
  1. Web Agents Should Adopt the Plan-Then-Execute Paradigm

    May 14, 2026Julien Piet, Annabella Chow, Yiwei Hou +5LLM Agent SafetyLLM Agent Planning

  2. Auditing Agent Harness Safety

    May 14, 2026Chengzhi Liu, Yichen Guo, Yepeng Liu +8AI Agent AuditingLLM Auditing

  3. No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills

    May 13, 2026Ying Li, Hongbo Wen, Yanju Chen +3LLM GuardrailsLLM Agent Verification

  4. Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety

    May 12, 2026Muhammad Bilal, Jon Crowcroft, Ruizhi Wang +2LLM Agent EvaluationAI Agent Security

  5. No Action Without a NOD: A Heterogeneous Multi-Agent Architecture for Reliable Service Agents

    May 12, 2026Zixu Yang, Hang Zheng, Nan Jiang +5AI Agent ReliabilityMulti-Agent LLM Systems

  6. SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces

    May 12, 2026Chang Jin, An Wang, Zeming Wei +7LLM Agent SecurityAI Agent Security Benchmarks

  7. On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment

    May 12, 2026Bo Yin, Qi Li, Xinchao WangAgentic RLLLM Safety Alignment

  8. LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments

    May 11, 2026Chiyu Zhang, Huiqin Yang, Bendong Jiang +8AI Agent Security BenchmarksLLM Jailbreak Attacks

  9. Safe Multi-Agent Behavior Must Be Maintained, Not Merely Asserted: Constraint Drift in LLM-Based Multi-Agent Systems

    May 11, 2026Tianxiao Li, Yixing Ma, Haiquan Wen +4Multi-Agent LLM SystemsAI Agent Governance

  10. NEXUS: Continual Learning of Symbolic Constraints for Safe and Robust Embodied Planning

    May 10, 2026Tiehan Cui, Peipei Liu, Yanxu Mao +3Continual Learning for LLM AgentsLarge Language Model-Based Robot Planning

  11. AI-Care: A Conversational Agentic System for Task Coordination in Alzheimer's Disease Care

    May 8, 2026Preyash Yadav, Michelle Cohn, Priyanka Koppolu +5Alzheimer's DiseaseTool-Using Agents

  12. Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents

    May 8, 2026Zhengyang Tang, Yi Zhang, Chenxin Li +18Computer-Use AgentsAgent Evaluation

  13. Why Does Agentic Safety Fail to Generalize Across Tasks?

    May 7, 2026Yonatan Slutzky, Yotam Alexander, Tomer Slor +2AI Agent SafetyLLM Agent Safety

  14. Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors

    May 7, 2026Jonas Wiedermann-Möller, Leonard Dung, Maksym AndriushchenkoAI Agent EvaluationAI Agent Safety

  15. TheraAgent: Self-Improving Therapeutic Agent for Precise and Comprehensive Treatment Planning

    May 7, 2026Junkai Li, Yunghwei Lai, Tianyi Zhu +3LLM Self-RefinementHealthcare

  16. SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety

    May 7, 2026Zhe Liu, Zonghao Ying, Wenxin Zhang +5LLM GuardrailsLLM Agent Security

  17. Safeguarding LLM Agents from Misalignment through Provenance Analysis

    May 1, 2026Yining She, Yiliang Liang, Eunsuk KangLLM GuardrailsLLM Agent Safety

  18. Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control

    Apr 29, 2026Mahiro Nakao, Kazuhiro TakemotoLanguage Model-Based ControlHealthcare

  19. Think Before You Act -- A Neurocognitive Governance Model for Autonomous AI Agents

    Apr 28, 2026Eranga Bandara, Ross Gore, Asanga Gunaratna +12Cognitive Architectures for AI AgentsAI Agent Governance

  20. The Interlocutor Effect: Why LLMs Leak More Personal Data to Agents Than Humans

    Apr 26, 2026Faouzi El Yagoubi, Godwin Badu-Marfo, Ranwa Al MallahData LeakageMulti-Agent LLM Systems

  21. Semantic Denial of Service in LLM-controlled robots

    Apr 25, 2026Jonathan Steinberg, Oren GalLanguage Model Safety EvaluationAdversarial Attacks on LLMs

  22. LLM-Guided Safety Agent for Edge Robotics with an ISO-Compliant Perception-Compute-Control Architecture

    Apr 22, 2026Xu Huang, Ruofan Zhang, Lu Cheng +9RoboticsRobot Safety

  23. Omission Constraints Decay While Commission Constraints Persist in Long-Context LLM Agents

    Apr 22, 2026Yeran GamageAI Agent MonitoringLLM Security

  24. SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models

    Apr 21, 2026Josue Torres-Fonseca, Naihao Deng, Yinpei Dai +5LLM Safety BenchmarksLarge Language Model-Based Robot Planning

  25. Human-Guided Harm Recovery for Computer Use Agents

    Apr 20, 2026Christy Li, Sky CH-Wang, Andi Peng +1Computer-Use AgentsAI Agent Security Benchmarks