AI Agent Security Benchmarks

Momentum

21 papers in the last four weeks, up 91% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 200

All topics
CardsList
  1. SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems

    Sep 1, 2026Rui Yang, Junjie Xu, Zhengyu Liu +4AI Agent Security BenchmarksMulti-Agent System Security

  2. Beyond the Payload: How User Invocation Shapes Coding Agent Vulnerability to Repository Poisoning

    Aug 31, 2026Fukang Zhu, Binbin Zhao, Ruixiao Lin +3Prompt SensitivityAI Coding Agents

  3. EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolving Agents

    Aug 31, 2026Doyun Kim, Chanwoo Kim, Sugyeong Eo +2AI Agent SecurityAI Agent Security Benchmarks

  4. Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation

    Aug 13, 2026Rana Muhammad Ahmed, Sabahat AbbasMCP SecurityLLM Agent Security

  5. ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

    Aug 12, 2026Yutao Mou, Pengfei Yang, Zhe Yin +6AI Agent EvaluationAdversarial Scenario Generation

  6. The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark

    Aug 11, 2026Jeremy Spence, Nicholas Assaderaghi, Feng Xiao +9Software Reverse EngineeringAI Agent Security Benchmarks

  7. Toward Metacognitive One-Shot Indirect Prompt Injection: Strategy Abstraction Via Outcome-Conditioned Reflection

    Aug 9, 2026Sihan Hou, Xinmeng Hou, Zhijun Zhang +5AI Agent Security BenchmarksIndirect Prompt Injection

  8. SkillsMetric: Mapping the Detection Boundary of Static Analysis for Malicious Agent Skills

    Aug 9, 2026Xinze Chen, Chi Zhang, Ping Ji +1Static Code AnalysisAI Agent Security Benchmarks

  9. Breadcrumbing Search Agents

    Aug 5, 2026Xuebin Li, Hanqing Zhao, Siyuan Liang +4Adversarial AttacksWeb Search Agents

  10. Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

    Aug 5, 2026Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi +6Autonomous Cyber DefenseLLM Red Teaming

  11. When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills

    Aug 4, 2026Yongli Xiang, Zhifang Zhang, Bojun Yang +4Privacy Leakage in Language ModelsAI Agent Security Benchmarks

  12. DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction

    Aug 4, 2026Xuyang Liu, Yibin Han, Zhenwei Zhang +8Retrieval-Augmented GenerationLLM Agent Evaluation

  13. WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks

    Aug 4, 2026Prince Zizhuang Wang, Aojie Yuan, Haiyue Zhang +3Multi-Agent CollaborationAI Agent Security Benchmarks

  14. Invisible Ink Threats: Adversarial Goals Behind Legitimate Tasks in Computer-Use Agents

    Aug 3, 2026Jia-Chen Zhang, Ze-Yu Zhang, Kai-Wei ZhangAI Agent Security BenchmarksPrompt Injection Defense

  15. SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

    Jul 29, 2026Lehan Wang, Boli Chen, Ruixue Ding +7LLM Agent EvaluationCybersecurity

  16. MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents

    Jul 26, 2026Belal S. Alsinglawi, Weizheng Wang, Junyi Wu +3Aerial RoboticsAI Agent Security Benchmarks

  17. Where Is the Cost of Third-Party API Routers in Agentic Software Development?

    Jul 26, 2026Donghao Fu, Jingxin Li, Xue Jiang +1Coding AgentsAI Agent Security

  18. False Prophets: On the Security of World Models in Agentic Systems

    Jul 25, 2026Erik Imgrund, Anna Wimbauer, Klim Kireev +1Adversarial AttacksAI Agent Security

  19. Agent Security Needs Redefinition through a Holistic Framework

    Jul 24, 2026Vincent Siu, Jingxuan He, Kyle Montgomery +3AI Agent SecurityAI Agent Security Benchmarks

  20. Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

    Jul 23, 2026Tencent WorkBuddy Bench Team, Siqi Cai, Shaopeng Chen +35Benchmark ContaminationSoftware Engineering Benchmarks

  21. IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests

    Jul 22, 2026Ankur Singh, Jinqiu Yang, Tse-Hsun ChenAI Coding AgentsAI Agent Security

  22. Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents

    Jul 22, 2026Or Zion Eliav, Eyal Lenga, Shir Bernstien +1AI Agent SecurityAI Agent Security Benchmarks

  23. Cross-Agent Campaign Attribution: Linking Asynchronous Attacks Across LLM Agents

    Jul 21, 2026SangJin Park, Myungsub Choi, Jineok Kim +1LLM Agent EvaluationAI Agent Security

  24. The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI

    Jul 20, 2026Om Narayan, Ramkinker Singh, Praveen BaskarAdversarial AttacksAI Agent Security