LLM Agent Security

LLM: Large Language Model

Latest papers 234

All topics
CardsList
  1. AgentSpy: Making AI Agent Behavior Observable

    Oct 5, 2026Christoph Bühler, Matteo Biagiola, Luca Di Grazia +1AI Agent ReliabilityAI Agent Auditing

  2. Runaway Reaction: When Benign Skills Compose into Malicious Behavior

    Oct 5, 2026Zunlong Zhou, Ziyuan Yang, Mengyu Sun +1AI Agent SecurityLLM Agent Security

  3. The Innocent Courier: Covert Exfiltration Through Legitimate LLM Web Fetching

    Oct 1, 2026Alessandro Pegoraro, Daryan Merx, Phillip Rieger +1LLM Agent SecurityPrivacy Leakage in Language Models

  4. Chaining Skills to Hijack LLM Agents

    Oct 1, 2026Tian Dong, Zixuan Ma, Haodong Zhao +3LLM Agent SecurityAI Agent Safety

  5. PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents

    Oct 1, 2026Fengpeng Li, Qizhou Wang, Yuke Hu +5LLM Agent SecurityRuntime Enforcement for AI Agents

  6. Memetic Trojans: Social Contagions as Carriers of Adversarial Payloads in Agent Networks

    Sep 30, 2026Birk Torpmann-Hagen, Finn Schwall, Leon MoonenLLM Agent SecurityMulti-Agent System Security

  7. Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks

    Sep 30, 2026Zezhong Wang, Xueyang Tang, Rui Lian +2LLM Agent SecurityAI Agent Monitoring

  8. Hiding in Plain Sight: Decoupling Pretext from Actuation for Skill Poisoning in LLM Agents

    Sep 30, 2026Wenxin Wu, Lingyong Yan, Lei Sha +2Adversarial Attacks on LLMsLLM Agent Security

  9. Can Agents Trust Their Skills? Uncovering Unsafe Chains of Trust in Skill-Based LLM Agents

    Sep 30, 2026Yan Wang, Zhihao Zhang, Ke Chen +5LLM Agent SecurityPrompt Injection Attacks on AI Agents

  10. SKILLLITE: Evidence-Guided Malicious Skill Auditing with Compact LLMs

    Sep 29, 2026Haoran Ou, Gelei Deng, Xuanye Zhang +3Small Language ModelsLLM Agent Security

  11. When Valid Tool Calls Change Meaning: Formation-Consistent Dispatch for LLM Agents

    Sep 28, 2026Geonwoo Kim, Brent ByungHoon KangLLM Agent SecurityRuntime Enforcement for AI Agents

  12. CoSec: Benchmarking Agent Security in Communities

    Sep 28, 2026Hao Chen, Wenhui Dong, Ye Chen +12LLM Agent SecurityAI Agent Security Benchmarks

  13. CoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-based Agents

    Sep 28, 2026Xiao Yang, Yangchen Ou, Yuhan Gao +3Adversarial TrainingLLM Agent Security

  14. Cheap to Hypothesize, Costly to Verify: The Defense Surface of Agentic Vulnerability Discovery

    Sep 28, 2026Kaikai Zhang, Zihan Zhang, Yuchong Xie +4Software Vulnerability DetectionCybersecurity

  15. Certified Multi-Source Integrity for Structured Agent Actions

    Sep 28, 2026Anmol Pandey, Aditya Jain, Liang Chen +2Data ProvenanceLLM Agent Security

  16. When Consent Outlives Context: Residual Authority Replay in Long-Lived Agents

    Sep 27, 2026Zhihao Zhang, Chao Wang, Rujia Li +3LLM Agent SecurityPrompt Injection Attacks on AI Agents

  17. Calibrated Decision Models for Autonomous Penetration-Testing Harnesses: JEV and Laya as System One Decision Layers for LLM-Driven Pentest Agents

    Sep 24, 2026Joas Antonio dos Santos BarbosaProbability CalibrationCybersecurity

  18. Persistent Billable State: Denial-of-Wallet Attacks and Defenses in Tool-Calling LLM Agents

    Sep 23, 2026Jinqian Zhang, Haojun Xia, Shujiang Wu +4AI Agent SecurityLLM Agent Security

  19. ActGov: Governing LLM Agent Actions via Policy-Constrained Validation

    Sep 21, 2026Kaiyuan Zhang, Yuke Peng, Ke Jiang +1LLM Agent SecurityRuntime Enforcement for AI Agents

  20. ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions

    Sep 16, 2026Guosen Wu, Huizhen Huang, Guoxiong Long +2Privacy AuditingLLM Agent Evaluation

  21. The Stochastic Deputy: Structural Tenant Isolation for Tool-Using LLM Agents

    Sep 13, 2026Mirza Samad Ahmed Baig, Syeda Anshrah Gillani, Asher Ali +1LLM Agent SecurityTool Access Control for LLM Agents

  22. SkillAtlas: An Attack Trace Library for Agent Skills

    Sep 11, 2026Yuxin Tian, Zenghao Duan, Liang Pang +2LLM Agent SecurityAgent Skill Security Auditing

  23. AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents

    Sep 10, 2026Zhihao Liu, Hongyu Sun, Zhiyuan Fu +7Adversarial Patch AttacksComputer-Use Agents