LLM Agent Safety

LLM: Large Language Model

Momentum

24 papers in the last four weeks, up 50% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 193

All topics
CardsList
  1. POLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM Agents

    Oct 6, 2026Yunju Kang, Seonghyeon Cho, Irene Li +2LLM GuardrailsRuntime Enforcement for AI Agents

  2. DecepEval: A Benchmark for Evaluating Deception in LLM Agents

    Oct 6, 2026Yiming Xu, Hongyue Yu, Beihua Yang +8Deception in Language ModelsLLM Agent Evaluation

  3. BazaarBench: Delegation Safety in Decentralized C2C Marketplaces Run by LLM Agents

    Oct 5, 2026Ziyan Wang, Shuqing Shi, James Oldfield +5AI Agent Safety BenchmarksLLM Agent Safety

  4. G-CARB: Graph-Localized Conformal Agent Risk Budget for Compositional Harm

    Oct 4, 2026Zijun Yu, Yu Gu, Vahid Partovi Nia +1AI Agent SafetyConformal Risk Control

  5. Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents

    Sep 30, 2026Haoyu Wang, Wei Zhao, Yedi Zhang +2AI Agent MonitoringLLM Agent Safety

  6. ActionGuard: Tool Call Authorization under Poisoned Skills

    Sep 30, 2026Jihun Han, Yejin Jang, Byung Il Kwak +1Runtime Enforcement for AI AgentsPrompt Injection Defense

  7. SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time

    Sep 29, 2026Yu Cheng, Yongkang Hu, Shuaijie Ma +12Continual Test-Time AdaptationLLM Agents

  8. SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety

    Sep 28, 2026Jianxing Chen, Xiao Yu, Shipra Agrawal +1AI Agent SafetyAI Agent Safety Benchmarks

  9. Tool Mediation Alters Refusal Mechanisms in Large Language Models

    Sep 28, 2026Abel Rodríguez, Giuseppe Garofalo, Lieven Desmet +1Language Model Safety EvaluationLLM Agent Safety

  10. PROACT-Agent: Progressive Runtime Oversight and Active Circuit-breaking for Real-Time Safety

    Sep 28, 2026Ding Jia, Wei Liu, Xianglong Du +5LLM GuardrailsRuntime Enforcement for AI Agents

  11. AgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM Agents

    Sep 27, 2026Tianzhuo Yang, Zirui Mi, Yantao Huang +4LLM Agent EvaluationLLM Agents

  12. Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure

    Sep 24, 2026David Schmotz, Derck Prinzhorn, Luca Beurer-Kellner +3Runtime Enforcement for AI AgentsAI Agent Safety Benchmarks

  13. SkinAgent AI: A Safety-Grounded Multimodal Agentic Framework for Non-Diagnostic Skincare Support

    Sep 24, 2026Muhammad Muhtasim Shahriar, Abdullah Mohammad Sayem, Tze Hui Liew +2LLM Agent OrchestrationLLM Agent Evaluation

  14. PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety

    Sep 23, 2026Jiapeng Sun, Yujin Zhou, Han Zhu +4AI Agent Safety BenchmarksInference-Time Intervention

  15. Emergent Collusion in Long-Horizon LLM Agent Interaction

    Sep 21, 2026Xinrui Shi, Yanzhe Zhang, Diyi YangMulti-Agent CoordinationLLM Agent Safety

  16. Indirect tipping: a social attack surface in AI agent populations

    Sep 21, 2026Ariel Flint, Luca Maria Aiello, Sara M. Constantino +2Multi-Agent CoordinationAI Agent Security

  17. MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs

    Sep 16, 2026Albert Wu, Nicholas Roberts, Tzu-Heng Huang +5Multi-Agent LLM SystemsFormal Verification

  18. Symbolic Temporal Supervision of LLM Agents Using Contracts

    Sep 16, 2026Yifeng Xiao, Pierluigi NuzzoRuntime Enforcement for AI AgentsLLM Agent Safety

  19. BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents

    Sep 14, 2026Sadia Asif, Mohammad Mohammadi Amiri, Momin Abbas +2Long-Horizon Agent EvaluationAI Agent Safety Benchmarks

  20. ActGuard: Pre-execution Action Auditing against Indirect Prompt Injection in LLM Agents

    Sep 14, 2026Bingzheng Wang, Xiaoyan Gu, Wentao Wang +3Runtime Enforcement for AI AgentsPrompt Injection Defense

  21. SchemeArena: Factorized Stress Testing of Scheming in LLM Agents

    Sep 8, 2026Jie Ruan, Inderjeet Nair, Amy Liu +3LLM Agent EvaluationAI Agent Monitoring

  22. HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals

    Sep 3, 2026Jasmine Brazilek, Miles Tidmarsh, Matthias Endres +2Moral Reasoning in Language ModelsLLM Safety Alignment

  23. SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment

    Sep 2, 2026Qinghua Mao, Wanying Qu, Dadi Guo +8LLM Safety AlignmentLLM Agent Safety