AI Agent Security Benchmarks

Momentum

21 papers in the last four weeks, up 91% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 200

All topics
CardsList
  1. If you're waiting for a sign... that might not be it! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems

    Apr 21, 2026Jiamin Chang, Minhui Xue, Ruoxi Sun +3VLM RobustnessAdversarial Attacks on VLMs

  2. How Adversarial Environments Mislead Agentic AI?

    Apr 20, 2026Zhonghao Zhan, Huichi Zhou, Zhenhao Li +3Adversarial RobustnessAI Agent Evaluation

  3. Human-Guided Harm Recovery for Computer Use Agents

    Apr 20, 2026Christy Li, Sky CH-Wang, Andi Peng +1Computer-Use AgentsAI Agent Security Benchmarks

  4. Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories

    Apr 19, 2026Ivan Bercovich, Ivgeni Segal, Kexun Zhang +3Reward HackingAI Agent Security Benchmarks

  5. Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks

    Apr 18, 2026Tyler H. Merves, Michael H. Conaway, Joseph M. Escobar +2LLM Agent EvaluationCybersecurity

  6. MemEvoBench: Benchmarking Safety Risks from Memory Misevolution in LLM Agents

    Apr 17, 2026Weiwei Xie, Shaoxiong Guo, Fan Zhang +5Long-Horizon Agent EvaluationAI Agent Security Benchmarks

  7. HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?

    Apr 16, 2026Yukun Jiang, Yage Zhang, Michael Backes +2LLM Agent SecurityAI Agent Security Benchmarks

  8. Detecting Multi-Agent Collusion Through Multi-Agent Interpretability

    Apr 1, 2026Aaron Rose, Carissa Cullen, Sahar Abdelnabi +3AI Agent Security BenchmarksMulti-Agent System Security

  9. Quantifying Frontier LLM Capabilities for Container Sandbox Escape

    Mar 1, 2026Rahul Marchand, Art O Cathain, Jerome Wynne +5LLM Safety BenchmarksLLM Agent Security

  10. Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent Defense

    Feb 9, 2026Jiacheng Liu, Yaxin Luo, Jiacheng Cui +3CybersecurityAI Agent Security

  11. StepShield: When, Not Whether to Intervene on Rogue Agents

    Jan 29, 2026Gloria Felicia, Zitha Sasindran, Jinfeng He +3LLM GuardrailsAI Agent Security

  12. Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks

    Dec 2, 2025Songwen Zhao, Danqing Wang, Kexun Zhang +3AI Coding AgentsSoftware Security

  13. Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks

    Oct 16, 2025Trilok Padhi, Pinxian Lu, Abdulkadir Erol +5Adversarial Attacks on LLMsAI Agent Security Benchmarks

  14. Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks

    Oct 1, 2025Shoumik Saha, Jifan Chen, Sam Mayers +3AI Coding AgentsAI Agent Security

  15. Measuring Harmfulness of Computer-Using Agents

    Jul 31, 2025Aaron Xuxiang Tian, Ruofan Zhang, Janet Tang +3Computer-Use Agent BenchmarksComputer-Use Agents

  16. Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents

    Jul 3, 2025Sizhe Chen, Arman Zharmagambetov, David Wagner +1LLM Agent SecurityAI Agent Security Benchmarks

  17. Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity?

    Date pendingAnmol Goel, Iryna GurevychData LeakageContextual Integrity