AI Agent Safety Benchmarks

Momentum

22 papers in the last four weeks, up 83% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 143

All topics
CardsList
  1. Right Screen, Wrong Transition: World Models as Verifiers for GUI Agents

    Oct 8, 2026Jiaming Zhang, Xuan Wang, Fuyao Zhang +3AI Agent MonitoringLatent World Models

  2. Safe Actions Alone Do Not Ensure Safe Agents: Identifying Unfulfilled Obligations with Guard Models

    Oct 8, 2026Youwei Feng, Yitong Zhang, Yuetong Liu +1LLM Agent SafetyLLM Guardrails

  3. Workerville: Towards an Organizational Behavior Account of Agent Safety

    Oct 8, 2026Hanjun Luo, Junting Mao, Yuhan Lu +5LLM Agent SafetyAgentic AI

  4. RT-Safe: Benchmarking Agent Safety in Real-Time Embodied Environment

    Oct 7, 2026Tianruo Rose Xu, Jiawei Ren, Yichi Yang +2AI Agent SafetySafe Robot Navigation

  5. DecepEval: A Benchmark for Evaluating Deception in LLM Agents

    Oct 6, 2026Yiming Xu, Hongyue Yu, Beihua Yang +8Deception in Language ModelsLLM Agent Evaluation

  6. From Evidence to Action: How Tool-Using Agents Fail

    Oct 6, 2026Hongzhan Lin, Shidong Cao, Ziyang Luo +3Agent Failure AnalysisEvidence-Grounded Reasoning

  7. BazaarBench: Delegation Safety in Decentralized C2C Marketplaces Run by LLM Agents

    Oct 5, 2026Ziyan Wang, Shuqing Shi, James Oldfield +5AI Agent Safety BenchmarksLLM Agent Safety

  8. Inspect Robots: Evaluating the Capabilities and Safety of Embodied AI

    Oct 5, 2026Christopher Leet, Achu Menon, Sravanthi Machcha +12AI Agent EvaluationRobotic Policy Evaluation

  9. Towards a Unified Misuse Monitoring Benchmark

    Oct 5, 2026Aniruddh Pramod, James Oldfield, Adel BibiAI Agent SecurityAI Agent Monitoring

  10. Benchmarking Jailbreak Guardrails for Embodied Agents

    Oct 5, 2026Xunguang Wang, Qingyue Wang, Yuguang Zhou +3LLM GuardrailsAI Agent Safety

  11. The Backdrop Exposes What the World Around an Agent Costs It

    Sep 29, 2026Nusrat Jahan Lia, Shubhashis Roy DiptaAI Agent ReliabilityPrompt Injection Attacks on AI Agents

  12. CheatBench: Measuring Reward Gaming in AI Agents

    Sep 28, 2026Long Phan, Stephen K. Yang, Jason J. Lim +10Reward HackingAI Agent Benchmarks

  13. SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety

    Sep 28, 2026Jianxing Chen, Xiao Yu, Shipra Agrawal +1AI Agent SafetyAI Agent Safety Benchmarks

  14. SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents

    Sep 28, 2026Saswat Das, Parvati Viswanathan, Daniel Donnelly +3AI Agent SafetyAI Agent Safety Benchmarks

  15. PROACT-Agent: Progressive Runtime Oversight and Active Circuit-breaking for Real-Time Safety

    Sep 28, 2026Ding Jia, Wei Liu, Xianglong Du +5LLM GuardrailsRuntime Enforcement for AI Agents

  16. AdaGuard: An Adaptive Guard Model with User-defined Policies

    Sep 28, 2026Yunhao Feng, Yifan Ding, Yuxiang Xie +4LLM GuardrailsAI Agent Safety

  17. AgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM Agents

    Sep 27, 2026Tianzhuo Yang, Zirui Mi, Yantao Huang +4LLM Agent EvaluationLLM Agents

  18. Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure

    Sep 24, 2026David Schmotz, Derck Prinzhorn, Luca Beurer-Kellner +3Runtime Enforcement for AI AgentsAI Agent Safety Benchmarks

  19. PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety

    Sep 23, 2026Jiapeng Sun, Yujin Zhou, Han Zhu +4AI Agent Safety BenchmarksInference-Time Intervention

  20. ClashBench: Conflicts Leading Agents to Seize and Harm

    Sep 17, 2026Yuejin Xie, Yu Li, Dadi Guo +6AI Agent SecurityAI Agent Safety

  21. PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?

    Sep 16, 2026Mika Okamoto, Ansel Kaplan ErolAI Agent AuditingLLM Agent Evaluation

  22. BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents

    Sep 14, 2026Sadia Asif, Mohammad Mohammadi Amiri, Momin Abbas +2Long-Horizon Agent EvaluationAI Agent Safety Benchmarks

  23. HazardAuditor: From Executable Threats to Safer Computer-Use Agents

    Sep 14, 2026Yunhao Feng, Ruixiao Lin, Ming Wen +5Language Model Safety EvaluationAI Agent Auditing

  24. MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents

    Sep 14, 2026Jianhua Jiang, Dongbo Yuan, Weihua LiLong-Horizon Agent EvaluationAI Agent Safety Benchmarks

  25. ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

    Sep 12, 2026Yizhan Li, Jianxin You, Mengyang Xiong +5Physical ReasoningRobotics Simulation