AI Agent Security Benchmarks

Momentum

21 papers in the last four weeks, up 91% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 200

All topics
CardsList
  1. Constrained-Action AI Remediation for SIEM/XDR via a NeMo-Guardrails Proxy

    Oct 7, 2026Georgios Koutidis, Nikolaos Kekatos, Tom Nianios +1LLMs for CybersecurityTool Access Control for LLM Agents

  2. SwarmReconGuard: Black-Box Detection of Distributed Collective Reconnaissance by Individually Benign-Looking Agent Populations

    Oct 6, 2026Vahid Tavakkoli, Kabeh Mohsenzadegan, Kyandoghere KyamakyaAI Agent SecurityNetwork Intrusion Detection

  3. Surviving the Router: Optimizing Skill Injections for Retrieval and Execution

    Oct 6, 2026Haneen Najjar, Luca Scionis, Haritz Puerto +1AI Agent Security BenchmarksLLM Agent Skill Retrieval

  4. ANT: A Multi-Granularity Network Traffic Dataset and Benchmark for Agents Behavior Auditing

    Oct 5, 2026Fan Li, Xiangyu Gao, Zixuan Liu +4AI Agent AuditingAI Agent Security Benchmarks

  5. Runaway Reaction: When Benign Skills Compose into Malicious Behavior

    Oct 5, 2026Zunlong Zhou, Ziyuan Yang, Mengyu Sun +1AI Agent SecurityLLM Agent Security

  6. Can CaMeLs Talk? Securing Multi-Agent Systems Against Indirect Prompt Injection Attacks

    Oct 5, 2026James Peters-Gill, Avi Semler, Henning Bartsch +2AI Agent Security BenchmarksInformation Flow Control

  7. AgentDoxx: Agentic Re-identification of Anonymized Text with Web Search

    Oct 4, 2026Jianing Wen, Tianshi LiPrivacy Leakage in Language ModelsAI Agent Security Benchmarks

  8. Blocking at the Boundary: Auditing Long-Horizon Agents against Staged Prompt Injection

    Oct 4, 2026Jingkai Liu, Yufei Han, Xiaoting Lyu +2Prompt InjectionAI Agent Auditing

  9. OverAct: Measuring and Mitigating Proactive Over-Authorization in LLM Tool-Calling Agents

    Oct 1, 2026Taolin Zhang, Jiuheng Wan, Hanyu Wang +2Privacy Leakage in Language ModelsAI Agent Security Benchmarks

  10. No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents

    Sep 30, 2026Ayan Javeed Shaikh, Arunesh Sinha, Nathaniel D. Bastian +1LLM Red TeamingAI Agent Evaluation

  11. MADBench: Benchmarking the Security of Multi-Agent Debate

    Sep 30, 2026Yuwan Liu, Jiaming Zhang, Yue Huang +1Adversarial Attacks on LLMsMulti-Agent Debate

  12. CoSec: Benchmarking Agent Security in Communities

    Sep 28, 2026Hao Chen, Wenhui Dong, Ye Chen +12LLM Agent SecurityAI Agent Security Benchmarks

  13. SEAD: A State-Based Perspective on Attack and Defense in Tool-Using Agents

    Sep 28, 2026Xinjie Shen, Junran Wang, Rongzhe Wei +1AI Agent SecurityAI Agent Security Benchmarks

  14. SecProbe: Adaptive Evaluation of Coding Agents on Cybersecurity Vulnerabilities

    Sep 27, 2026Xiaonan Luo, Yue Huang, Kehan Guo +9Software Vulnerability DetectionAI Agent Evaluation

  15. On the Effectiveness of Kernel-Level Evidence for Agent Security

    Sep 24, 2026Spencer King, Zhilu Zhang, Mikhail Kuznetsov +3AI Agent SecurityAI Agent Security Benchmarks

  16. Persistent Billable State: Denial-of-Wallet Attacks and Defenses in Tool-Calling LLM Agents

    Sep 23, 2026Jinqian Zhang, Haojun Xia, Shujiang Wu +4AI Agent SecurityLLM Agent Security

  17. Evaluating Coding Agents on Kernel Exploit Generation

    Sep 22, 2026Junyoung Jang, Gwanhyun Lee, Hwiwon Lee +4AI Coding AgentsSoftware Security

  18. ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions

    Sep 16, 2026Guosen Wu, Huizhen Huang, Guoxiong Long +2Privacy AuditingLLM Agent Evaluation

  19. Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale

    Sep 14, 2026Aman Priyanshu, Supriti Vijay, Kimia Majd +8Software Vulnerability DetectionAI Agent Security

  20. DriftNet: A Dual-Head Trajectory Transformer for Detecting and Localizing Prompt Injection in LLM Agents

    Sep 12, 2026Asif Pinjari, Mithun Paul Saint-GermainAI Agent Security BenchmarksAI Agent Monitoring

  21. An Experimental Evaluation of Multimodal Prompt Injection Attacks on Agentic AI Frameworks

    Sep 10, 2026Viet K. Nguyen, Mohammad I. HusainAI Agent SecurityAI Agent Security Benchmarks

  22. VEX-Bench: Benchmarking LLM Agents for Assessing Exploitability of Software Supply Chain Vulnerabilities

    Sep 7, 2026Jiahao Shi, Edward Tsien, Yifeng Di +10LLM Agent EvaluationSoftware Security

  23. Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning

    Sep 5, 2026Zhongan Bi, Qiwen Wang, Jianrong Jiang +11LLM Agent EvaluationWeb Search Agents

  24. PatchBench: Evaluating AI Agents for Vulnerability Patching

    Sep 3, 2026Chihao Shen, Jiacheng Li, Aastha Mahajan +3Automated Program RepairNeural Network Memorization

  25. Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents

    Sep 1, 2026Xiaofang Yang, Ziqi Miao, Dianbo Sui +2AI Agent SecurityAI Agent Security Benchmarks

  26. SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems

    Sep 1, 2026Rui Yang, Junjie Xu, Zhengyu Liu +4AI Agent Security BenchmarksMulti-Agent System Security

  27. Beyond the Payload: How User Invocation Shapes Coding Agent Vulnerability to Repository Poisoning

    Aug 31, 2026Fukang Zhu, Binbin Zhao, Ruixiao Lin +3Prompt SensitivityAI Coding Agents

  28. EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolving Agents

    Aug 31, 2026Doyun Kim, Chanwoo Kim, Sugyeong Eo +2AI Agent SecurityAI Agent Security Benchmarks

  29. Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation

    Aug 13, 2026Rana Muhammad Ahmed, Sabahat AbbasMCP SecurityLLM Agent Security

  30. ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

    Aug 12, 2026Yutao Mou, Pengfei Yang, Zhe Yin +6AI Agent EvaluationAdversarial Scenario Generation

  31. The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark

    Aug 11, 2026Jeremy Spence, Nicholas Assaderaghi, Feng Xiao +9Software Reverse EngineeringAI Agent Security Benchmarks

  32. Toward Metacognitive One-Shot Indirect Prompt Injection: Strategy Abstraction Via Outcome-Conditioned Reflection

    Aug 9, 2026Sihan Hou, Xinmeng Hou, Zhijun Zhang +5AI Agent Security BenchmarksIndirect Prompt Injection

  33. SkillsMetric: Mapping the Detection Boundary of Static Analysis for Malicious Agent Skills

    Aug 9, 2026Xinze Chen, Chi Zhang, Ping Ji +1Static Code AnalysisAI Agent Security Benchmarks

  34. Breadcrumbing Search Agents

    Aug 5, 2026Xuebin Li, Hanqing Zhao, Siyuan Liang +4Adversarial AttacksWeb Search Agents

  35. Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

    Aug 5, 2026Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi +6Autonomous Cyber DefenseLLM Red Teaming

  36. When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills

    Aug 4, 2026Yongli Xiang, Zhifang Zhang, Bojun Yang +4Privacy Leakage in Language ModelsAI Agent Security Benchmarks

  37. DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction

    Aug 4, 2026Xuyang Liu, Yibin Han, Zhenwei Zhang +8Retrieval-Augmented GenerationLLM Agent Evaluation

  38. WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks

    Aug 4, 2026Prince Zizhuang Wang, Aojie Yuan, Haiyue Zhang +3Multi-Agent CollaborationAI Agent Security Benchmarks

  39. Invisible Ink Threats: Adversarial Goals Behind Legitimate Tasks in Computer-Use Agents

    Aug 3, 2026Jia-Chen Zhang, Ze-Yu Zhang, Kai-Wei ZhangAI Agent Security BenchmarksPrompt Injection Defense

  40. SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

    Jul 29, 2026Lehan Wang, Boli Chen, Ruixue Ding +7LLM Agent EvaluationCybersecurity

  41. MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents

    Jul 26, 2026Belal S. Alsinglawi, Weizheng Wang, Junyi Wu +3Aerial RoboticsAI Agent Security Benchmarks

  42. Where Is the Cost of Third-Party API Routers in Agentic Software Development?

    Jul 26, 2026Donghao Fu, Jingxin Li, Xue Jiang +1Coding AgentsAI Agent Security

  43. False Prophets: On the Security of World Models in Agentic Systems

    Jul 25, 2026Erik Imgrund, Anna Wimbauer, Klim Kireev +1Adversarial AttacksAI Agent Security

  44. Agent Security Needs Redefinition through a Holistic Framework

    Jul 24, 2026Vincent Siu, Jingxuan He, Kyle Montgomery +3AI Agent SecurityAI Agent Security Benchmarks

  45. Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

    Jul 23, 2026Tencent WorkBuddy Bench Team, Siqi Cai, Shaopeng Chen +35Benchmark ContaminationSoftware Engineering Benchmarks

  46. IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests

    Jul 22, 2026Ankur Singh, Jinqiu Yang, Tse-Hsun ChenAI Coding AgentsAI Agent Security

  47. Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents

    Jul 22, 2026Or Zion Eliav, Eyal Lenga, Shir Bernstien +1AI Agent SecurityAI Agent Security Benchmarks

  48. Cross-Agent Campaign Attribution: Linking Asynchronous Attacks Across LLM Agents

    Jul 21, 2026SangJin Park, Myungsub Choi, Jineok Kim +1LLM Agent EvaluationAI Agent Security

  49. The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI

    Jul 20, 2026Om Narayan, Ramkinker Singh, Praveen BaskarAdversarial AttacksAI Agent Security