AI Agent Security Benchmarks

Momentum

21 papers in the last four weeks, up 91% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 200

All topics
CardsList
  1. AutoDojo: A Generative Benchmark for Evaluating Prompt Injection Defenses in LLM Agents

    Jun 13, 2026Xinhang Ma, Taoran Li, Chaowei Xiao +3AI Agent Security BenchmarksIndirect Prompt Injection

  2. Same-Origin Policy for Agentic Browsers

    Jun 12, 2026Xilong Wang, Xiaoxing Chen, Patrick Li +2Data LeakageAI Agent Security

  3. SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents

    Jun 11, 2026Rui Melo, Riccardo Fogliato, Sean Zhou +2Software Vulnerability DetectionAI Agent Security

  4. Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents

    Jun 11, 2026Zihao Wang, Yiming Li, Yutong Wu +8AI Agent EvaluationLLM Agent Security

  5. The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems

    Jun 11, 2026Jiaqi Luo, Jiarun Dai, Zhile Chen +8LLM EvaluationAI Agent Security Benchmarks

  6. MAStrike: Shapley-Guided Collusive Red-Teaming on Multi-Agent Systems

    Jun 11, 2026Chejian Xu, Zhaorun Chen, Jingyang Zhang +5AI Agent Security BenchmarksShapley Value Attribution

  7. ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity

    Jun 9, 2026Andrew Bo Liu, Samira Nedungadi, Bryce Cai +3LLM Agent SecurityAI Agent Security Benchmarks

  8. CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs

    Jun 9, 2026Joachim Schaeffer, Alexander Panfilov, Thomas Jiralerspong +4Language Model Safety EvaluationAI Agent Security Benchmarks

  9. RedAct: Redacting Agent Capability Traces for Procedural Skill Protection

    Jun 9, 2026Shuwen Xu, Zhitao He, Yi R. FungAI Agent SecurityAI Agent Security Benchmarks

  10. Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation

    Jun 9, 2026Yuchen Ling, Shengcheng Yu, Zhenyu Chen +1Prompt InjectionLLM Agent Security

  11. Assessing Automated Prompt Injection Attacks in Agentic Environments

    Jun 9, 2026David Hofer, Edoardo Debenedetti, Florian TramèrAdversarial Prompt GenerationLLM Agent Security

  12. Context-Fractured Decomposition Attacks on Tool-Using LLM Agents: Exploiting Artifact Provenance Gaps

    Jun 8, 2026Xiaofeng Lin, Yukai Yang, Daniel Guo +3Jailbreak AttacksLLM Agent Security

  13. Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops

    Jun 8, 2026Ziqian Zhong, Ivgeni Segal, Ivan Bercovich +3Reward HackingAI Agent Security Benchmarks

  14. GitInject: Real-World Prompt Injection Attacks in AI-Powered CI/CD Pipelines

    Jun 7, 2026Jafar Isbarov, Umid Suleymanov, Ilia Shumailov +1AI Agent SecurityAI Agent Security Benchmarks

  15. Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety

    Jun 3, 2026Catherine Ge-Wang, Tyler Crosse, Benjamin Hadad +3Adversarial AttacksAI Agent Safety

  16. CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity Capabilities

    Jun 3, 2026Tianneng Shi, Robin Rheem, Dongwei Jiang +13Automated Program RepairSoftware Vulnerability Detection

  17. What If Prompt Injection Never Left? Exploring Cross-Session Stored Prompt Injection in Agentic Systems

    Jun 3, 2026Yuanbo Xie, Tianyun Liu, Yingjie Zhang +4Prompt InjectionAI Agent Security

  18. From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents

    Jun 3, 2026Pritam Dash, Tongyu Ge, Aditi Jain +2AI Agent Security BenchmarksAgent Memory Poisoning

  19. SkillHarm: Lifecycle-Aware Skill-Based Attacks via Automated Construction

    Jun 1, 2026Yuting Ning, Zhehao Zhang, Yash Kumar Lal +8AI Agent SecurityAI Agent Security Benchmarks

  20. SeClaw: Spec-Driven Security Task Synthesis for Evaluating Autonomous Agents

    Jun 1, 2026Hao Cheng, Changtao Miao, Tianle Song +21AI Agent EvaluationLLM Agent Security

  21. Benchmarking Security Risk Detection and Verification in Open Agentic Skill Ecosystems

    May 30, 2026Ismail Hossain, Sai Puppala, Zhuoran Lu +2AI Agent SecurityAI Agent Security Benchmarks

  22. "I Strongly Suspect This Website Is a Scam": Benchmarking PII Leakage and Detection without Defense in Autonomous Web Agents

    May 30, 2026Soham Roy, Sarthakbrata Halder, Arya Bharaty +5Data LeakageLLM Agent Security

  23. When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems

    May 30, 2026Su Wang, Pin Qian, Yihang Chen +6Software SecurityAI Agent Security

  24. From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Backdoors

    May 29, 2026Jiejun Tan, Zhicheng Dou, Xinyu Yang +4Prompt InjectionLLM Agent Security