4 papers in the last four weeks, level with the four weeks before. 0.0% of all new papers.
Oct 7, 2026·Wanjing Han, Levi Taiji Li, Mu Zhang +2AI Agent SecurityTool-Using Vision-Language Agents
University of Utah
Oct 4, 2026·Tongyan Hu, Hao Li, Xiaogeng Liu +8Test-Time TrainingLLM Jailbreak Attacks
Johns Hopkins University · National University of Singapore · Washington University in St. Louis +2
Sep 30, 2026·Ayan Javeed Shaikh, Arunesh Sinha, Nathaniel D. Bastian +1LLM Red TeamingAI Agent Evaluation
Indiana University Bloomington, IN, USA · Rutgers University Piscataway, NJ, USA · Johns Hopkins University Baltimore, MD, USA
Sep 17, 2026·Alex Remedios, Simon Storf, Fabien Roger +1AI Agent SecurityAI Agent Monitoring
Anthropic Fellows Program · Anthropic
Sep 10, 2026·Divyanshu Kumar, Nitin Aravind Birur, Tanay Baswa +2AI Agent SecurityAI Agent Security Benchmarks
Enkrypt AI
Aug 31, 2026·Doyun Kim, Chanwoo Kim, Sugyeong Eo +2AI Agent SecurityAI Agent Security Benchmarks
Soongsil University · Yonsei University · Jeju National University
Aug 31, 2026·Chen Xiong, Zhiyuan He, Pin-Yu Chen +2Computer-Use AgentsIndirect Prompt Injection
The Chinese University of Hong Kong · IBM Research · University of Zagreb, FER & Radboud University +1
Aug 11, 2026·Lukasz Olejnik, Wenchao Dong, Jonas R. Kunst +4Agent-Based ModelingAutomated Red Teaming
Department of War Studies, King’s College London, London, United Kingdom · Independent researcher · Max Planck Institute for Security and Privacy, Bochum, Germany +4
Aug 10, 2026·Berkay Ozcam, Irem Onen, Mehmet Fatih Amasyali +1Adversarial Prompt GenerationAdversarial Attacks on LLMs
Cybersecurity R&D Turkcell Istanbul, Turkey · Department of Computer Engineering Yıldız Technical University Istanbul, Turkey
Aug 5, 2026·Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi +6Autonomous Cyber DefenseLLM Red Teaming
University of California, Irvine · Northeastern University · Johns Hopkins University
Aug 3, 2026·Zheng Lin, Yuzhe Huang, Zhenxing Niu +2Agent Memory PoisoningAutomated Red Teaming
Xidian University Xi’an, Shaanxi, China
Jul 28, 2026·Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal +15Adversarial TrainingLLM Red Teaming
OpenAI
Jul 15, 2026·Genglin Liu, Muye Zhang, Krishnamurthy Viswanathan +5Multimodal RobustnessAdversarial Scenario Generation
1UCLA · *Work done as a student researcher at Google. · 2Google
Jul 13, 2026·Xutao Mao, Xiang Zheng, Cong WangLLM Agent SecurityAutomated Red Teaming
City University of Hong Kong
Jul 13, 2026·Praneeth Narisetty, Shiva Nagendra Babu KoreAutomated Penetration TestingAutomated Red Teaming
LaunchSafe
Jun 22, 2026·Yarin Yerushalmi Levi, Roy Betser, Amit Giloni +5AI Agent SecurityAI Agent Security Benchmarks
1Fujitsu Research of Europe (FRE) · 2Fujitsu Research of India Pvt. Ltd. (FRIPL)
Jun 11, 2026·Chejian Xu, Zhaorun Chen, Jingyang Zhang +5AI Agent Security BenchmarksShapley Value Attribution
University of Illinois, Urbana-Champaign · Virtue AI · University of Chicago +3
Jun 8, 2026·Ziqian Zhong, Ivgeni Segal, Ivan Bercovich +3Reward HackingAI Agent Security Benchmarks
Carnegie Mellon University · Fewshot Corp · Independent Researcher
May 12, 2026·Hao Wang, Hanchen Li, Qiuyang Mang +3Computer-Use Agent BenchmarksReward Hacking
UC Berkeley
May 12, 2026·Zhaojiacheng ZhouLLM Agent SecurityAutomated Red Teaming
Department of Computer Science and Engineering Shanghai Jiao Tong University Shanghai 200240, China
May 12, 2026·Chia-Pei, Chen, Kentaroh Toyoda +2AI Agent SecurityAI Agent Security Benchmarks
Janet · Vulcan Research, AIFT
May 12, 2026·Cristian Morasso, Anisa Halimi, Muhammad Zaid Hameed +1Adversarial TrainingAdversarial Prompt Generation
IBM Research · Trinity College Dublin, Ireland
May 11, 2026·Neil Fendley, Zhengyu Liu, Aonan Guan +2Adversarial Attacks on LLMsAgentic Workflows
Johns Hopkins University Maryland, USA · Wyze Labs Washington, USA
May 10, 2026·Monika Jotautaitė, Maria Angelica Martinez, Ollie Matthews +1Software Engineering AgentsAI Agent Security Benchmarks
Independent · Redwood Research
May 7, 2026·Wesley Hanwen Deng, Mingxi Yan, Sunnie S. Y. Kim +5Adversarial Prompt GenerationLLM Red Teaming
Carnegie Mellon University Pittsburgh, Pennsylvania, USA · Apple Seattle, Washington, USA · Apple Cupertino, California, USA +1
May 6, 2026·Zhaorun Chen, Xun Liu, Haibo Tong +14AI Agent EvaluationAI Agent Security
Virtue AI · University of Chicago · University of Illinois, Urbana-Champaign +4
May 5, 2026·Raja Sekhar Rao Dheekonda, Will Pearce, Nick LandersAdversarial Attacks on LLMsLLM Red Teaming
1Dreadnode, USA
Apr 24, 2026·Aishwarya Padmakumar, Leon Derczynski, Traian Rebedea +1Adversarial Attacks on LLMsLLM Red Teaming
NVIDIA
Apr 23, 2026·Tanmay Gautam, Alireza Bahramali, Sandeep AtluriLarge Language Model-Guided OptimizationProgram Synthesis
Responsible AI, Microsoft
Apr 22, 2026·Jesse Zymet, Andy Luo, Swapnil Shinde +2Adversarial Prompt GenerationLLM Red Teaming
Capital One, AI Foundations