AI Agent Security Benchmarks

Momentum

21 papers in the last four weeks, up 91% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 200

All topics
CardsList
  1. VisualLeakBench: Reproducible Action-Boundary Propagation Failures in Vision-Language Agents

    May 29, 2026Youting Wang, Yuan Tang, Yitian Qian +1Data LeakageTool Use in VLMs

  2. PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say

    May 29, 2026Mingxuan Zhang, Jiahui Han, Dadi Guo +5Data LeakagePrivacy Auditing

  3. Send a SCOUT First: Pre-hoc Reasoning for Adaptive Detector Allocation in Prompt-Injection Defense

    May 29, 2026Shuhao Zhang, Jiarui Li, Qi Cao +2AI Agent Security BenchmarksLLM Security

  4. MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents

    May 29, 2026Alexander Gurung, Spandana Gella, Alexandre Drouin +3Data LeakageDeep Research Agents

  5. The Surface You Test Is Not the Surface That Breaks

    May 28, 2026Syed Nazmus Sakib, Nafiul Haque, Shahrear Bin Amin +1LLM Agent SecurityAI Agent Security Benchmarks

  6. Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

    May 28, 2026Aditya Nawal, Manit Baser, Mohan GurusamyAI Agent Security BenchmarksLLM Safety Evaluation

  7. Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents

    May 27, 2026Yongxiang Li, Moxin Li, Zhixin Ma +2LLM Agent SecurityAI Agent Security Benchmarks

  8. SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents

    May 27, 2026Yubin Qu, Yi Liu, Gelei Deng +4Coding AgentsAI Agent Security

  9. MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content

    May 27, 2026Ruoqi Guo, Yi Liu, Gelei Deng +7Mobile GUI AutomationGUI Agents

  10. SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?

    May 26, 2026Hwiwon Lee, Jiawei Liu, Dongjun Kim +3Software Vulnerability DetectionSoftware Security

  11. When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Agents

    May 22, 2026Shi Liu, Xuehai Tang, Xikang Yang +4MCP SecurityLLM Agent Security

  12. Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

    May 21, 2026Piercosma Bisconti, Matteo Prandi, Federico Pierucci +11LLM Safety BenchmarksComputer-Use Agent Benchmarks

  13. Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard

    May 21, 2026Sahar Abdelnabi, Chris Hicks, Konrad Rieck +1AI Agent EvaluationAI Agent Security

  14. Benchmarking Autonomous Agents against Temporal, Spatial, and Semantic Evasions

    May 21, 2026Jianan Ma, Xiaohu Du, Ruixiao Lin +8LLM Agent EvaluationLLM Agent Security

  15. POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents

    May 18, 2026Qiaoyuan Zheng, Yiqu Yang, Qi Gao +1LLM Agent EvaluationLLM Agent Security

  16. Overeager Coding Agents: Measuring Out-of-Scope Actions on Benign Tasks

    May 18, 2026Yubin Qu, Ying Zhang, Yanjun Zhang +4Computer-Use Agent BenchmarksCoding Agents

  17. LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection

    May 18, 2026Lei Zhao, Abhay Bhaskar, Edgar DobribanAI Agent SecurityAI Agent Safety

  18. Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback

    May 17, 2026Lecheng Yan, Ruizhe Li, Xicheng Han +5AI Agent ReliabilityLLM Agent Evaluation

  19. ADR: An Agentic Detection System for Enterprise Agentic AI Security

    May 17, 2026Chenning Li, Pan Hu, Justin Xu +9MCP SecurityAI Agent Security

  20. ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents

    May 17, 2026Udari Madhushani Sehwag, Zhengyang Shan, Heming Liu +3LLM Agent SecurityAI Agent Security Benchmarks

  21. SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors

    May 15, 2026Elle Najt, Colin Toft, Tyler Tracy +2Software Engineering AgentsAdversarial Attacks

  22. Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks

    May 14, 2026Karthik Raghu Iyer, Yazdan Jamshidi, Nicholas Bray +1LLM Safety BenchmarksBenchmark Auditing

  23. Do Coding Agents Understand Least-Privilege Authorization?

    May 14, 2026Zheng Yan, Jingxiang Weng, Charles Chen +9Software Engineering AgentsAI Agent Security

  24. Auditing Agent Harness Safety

    May 14, 2026Chengzhi Liu, Yichen Guo, Yepeng Liu +8AI Agent AuditingLLM Auditing

  25. ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents

    May 13, 2026Seunghyun Lee, David BrumleyAI Agent Security BenchmarksLLMs for Cybersecurity

  26. AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills

    May 13, 2026Haomin Zhuang, Hanwen Xing, Yujun Zhou +5AI Agent SecurityLLM Agent Security

  27. The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems

    May 12, 2026Tanzim Ahad, Ismail Hossain, Md Jahangir Alam +3AI Agent Security BenchmarksAgent Memory Poisoning