LLM Agent Security

LLM: Large Language Model

Latest papers 234

All topics
CardsList
  1. The Surface You Test Is Not the Surface That Breaks

    May 28, 2026Syed Nazmus Sakib, Nafiul Haque, Shahrear Bin Amin +1LLM Agent SecurityAI Agent Security Benchmarks

  2. Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents

    May 27, 2026Yongxiang Li, Moxin Li, Zhixin Ma +2LLM Agent SecurityAI Agent Security Benchmarks

  3. Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems

    May 26, 2026Aman Priyanshu, Supriti Vijay, Esha PahwaData LeakageMulti-Agent LLM Systems

  4. SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?

    May 26, 2026Hwiwon Lee, Jiawei Liu, Dongjun Kim +3Software Vulnerability DetectionSoftware Security

  5. ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation

    May 26, 2026Xiaochong Jiang, Shiqi Yang, Ziwei Li +3LLM Agent SecurityCapability-Based Security

  6. Security of OpenClaw Agents: Fundamentals, Attacks, and Countermeasures

    May 25, 2026Yuntao Wang, Jianle Ba, Han Liu +5AI Agent SecurityLLM Agent Security

  7. Evo-Attacker: Memory-Augmented Reinforcement Learning for Long-Horizon Tool Attacks on LLM-MAS

    May 25, 2026Bingyu Yan, Xiaoming Zhang, Jinyu Hou +4Adversarial Attacks on LLMsLLM Agent Security

  8. MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning

    May 24, 2026Xuanye Zhang, Yongsen Zheng, Zhuqin Xu +5LLM Agent SecurityTool-Using Agents

  9. IterInject: Indirect Prompt Injection Against LLM Agents via Feedback-Guided Iterative Optimization

    May 23, 2026Zixuan Chen, Jiaxiang Chen, Li Luo +4LLM Agent SecurityIndirect Prompt Injection

  10. When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Agents

    May 22, 2026Shi Liu, Xuehai Tang, Xikang Yang +4MCP SecurityLLM Agent Security

  11. Benchmarking Autonomous Agents against Temporal, Spatial, and Semantic Evasions

    May 21, 2026Jianan Ma, Xiaohu Du, Ruixiao Lin +8LLM Agent EvaluationLLM Agent Security

  12. PocketAgents: A Manifest-Driven Library of Autonomous Defense Agents

    May 20, 2026Sidnei Barbieri, Ágney Lopes Roth Ferraz, Lourenço Alves Pereira JúniorAutonomous Cyber DefenseLLM Agent Security

  13. POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents

    May 18, 2026Qiaoyuan Zheng, Yiqu Yang, Qi Gao +1LLM Agent EvaluationLLM Agent Security

  14. OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences

    May 18, 2026Kaixiang Wang, Jiong Lou, Zhaojiacheng Zhou +1LLM Agent SecurityAgent Memory Poisoning

  15. Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback

    May 17, 2026Lecheng Yan, Ruizhe Li, Xicheng Han +5AI Agent ReliabilityLLM Agent Evaluation

  16. ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents

    May 17, 2026Udari Madhushani Sehwag, Zhengyang Shan, Heming Liu +3LLM Agent SecurityAI Agent Security Benchmarks

  17. Who Owns This Agent? Tracing AI Agents Back to Their Owners

    May 15, 2026Ruben Chocron, Doron Jonathan Ben Chayim, Eyal Lenga +3AI AccountabilityLLM Agent Security

  18. Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces

    May 14, 2026William Lugoloobi, Samuelle Marro, Jabez Magomere +2Language Model FingerprintingLLM Agent Security

  19. MemLineage: Lineage-Guided Enforcement for LLM Agent Memory

    May 14, 2026Ciyan Ouyang, Rui HouLLM Agent MemoryLLM Agent Security

  20. AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills

    May 13, 2026Haomin Zhuang, Hanwen Xing, Yujun Zhou +5AI Agent SecurityLLM Agent Security

  21. SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces

    May 12, 2026Chang Jin, An Wang, Zeming Wei +7LLM Agent SecurityAI Agent Security Benchmarks

  22. Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems

    May 12, 2026Zhaojiacheng ZhouLLM Agent SecurityAutomated Red Teaming

  23. Can a Single Message Paralyze the AI Infrastructure? The Rise of AbO-DDoS Attacks through Targeted Mobius Injection

    May 12, 2026Zi Liang, Ronghua Li, Yanyun Wang +2AI Agent SecurityLLM Agent Security

  24. Comment and Control: Hijacking Agentic Workflows via Context-Grounded Evolution

    May 11, 2026Neil Fendley, Zhengyu Liu, Aonan Guan +2Adversarial Attacks on LLMsAgentic Workflows

  25. Engineering Robustness into Personal Agents with the AI Workflow Store

    May 11, 2026Roxana Geambasu, Mariana Raykova, Pierre Tholoniat +3Agentic WorkflowsLLM Agent Security