Automated Red Teaming

Latest papers 35

All topics
CardsList
  1. Red-TTT: Test-Time Training for Automated Jailbreaking Large Language Models

    Oct 4, 2026Tongyan Hu, Hao Li, Xiaogeng Liu +8Test-Time TrainingLLM Jailbreak Attacks

  2. No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents

    Sep 30, 2026Ayan Javeed Shaikh, Arunesh Sinha, Nathaniel D. Bastian +1LLM Red TeamingAI Agent Evaluation

  3. Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents

    Sep 17, 2026Alex Remedios, Simon Storf, Fabien Roger +1AI Agent SecurityAI Agent Monitoring

  4. EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolving Agents

    Aug 31, 2026Doyun Kim, Chanwoo Kim, Sugyeong Eo +2AI Agent SecurityAI Agent Security Benchmarks

  5. SIR: Self-improving Red-teaming for Compute Use Agents

    Aug 31, 2026Chen Xiong, Zhiyuan He, Pin-Yu Chen +2Computer-Use AgentsIndirect Prompt Injection

  6. IO Factory: Simulating AI-Enabled Influence Campaigns at Scale

    Aug 11, 2026Lukasz Olejnik, Wenchao Dong, Jonas R. Kunst +4Agent-Based ModelingAutomated Red Teaming

  7. Generating Attacks for LLMs with GFlowNets

    Aug 10, 2026Berkay Ozcam, Irem Onen, Mehmet Fatih Amasyali +1Adversarial Prompt GenerationAdversarial Attacks on LLMs

  8. Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

    Aug 5, 2026Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi +6Autonomous Cyber DefenseLLM Red Teaming

  9. Salami Attack: Stealthy Collusive Memory Poisoning against OpenClaw

    Aug 3, 2026Zheng Lin, Yuzhe Huang, Zhenxing Niu +2Agent Memory PoisoningAutomated Red Teaming

  10. GPT-Red: Automated Red Teaming via Self-Play at Scale

    Jul 28, 2026Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal +15Adversarial TrainingLLM Red Teaming

  11. Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation

    Jul 15, 2026Genglin Liu, Muye Zhang, Krishnamurthy Viswanathan +5Multimodal RobustnessAdversarial Scenario Generation

  12. Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming

    Jul 13, 2026Xutao Mao, Xiang Zheng, Cong WangLLM Agent SecurityAutomated Red Teaming

  13. RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems

    Jun 22, 2026Yarin Yerushalmi Levi, Roy Betser, Amit Giloni +5AI Agent SecurityAI Agent Security Benchmarks

  14. MAStrike: Shapley-Guided Collusive Red-Teaming on Multi-Agent Systems

    Jun 11, 2026Chejian Xu, Zhaorun Chen, Jingyang Zhang +5AI Agent Security BenchmarksShapley Value Attribution

  15. Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops

    Jun 8, 2026Ziqian Zhong, Ivgeni Segal, Ivan Bercovich +3Reward HackingAI Agent Security Benchmarks

  16. Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems

    May 12, 2026Zhaojiacheng ZhouLLM Agent SecurityAutomated Red Teaming

  17. Persona-Conditioned Adversarial Prompting: Multi-Identity Red-Teaming for Adversarial Discovery and Mitigation

    May 12, 2026Cristian Morasso, Anisa Halimi, Muhammad Zaid Hameed +1Adversarial TrainingAdversarial Prompt Generation

  18. Comment and Control: Hijacking Agentic Workflows via Context-Grounded Evolution

    May 11, 2026Neil Fendley, Zhengyu Liu, Aonan Guan +2Adversarial Attacks on LLMsAgentic Workflows

  19. MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring

    May 10, 2026Monika Jotautaitė, Maria Angelica Martinez, Ollie Matthews +1Software Engineering AgentsAI Agent Security Benchmarks

  20. PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI

    May 7, 2026Wesley Hanwen Deng, Mingxi Yan, Sunnie S. Y. Kim +5Adversarial Prompt GenerationLLM Red Teaming

  21. DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents

    May 6, 2026Zhaorun Chen, Xun Liu, Haibo Tong +14AI Agent EvaluationAI Agent Security

  22. Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours

    May 5, 2026Raja Sekhar Rao Dheekonda, Will Pearce, Nick LandersAdversarial Attacks on LLMsLLM Red Teaming

  23. Training a General Purpose Automated Red Teaming Model

    Apr 24, 2026Aishwarya Padmakumar, Leon Derczynski, Traian Rebedea +1Adversarial Attacks on LLMsLLM Red Teaming

  24. Adaptive Instruction Composition for Automated LLM Red-Teaming

    Apr 22, 2026Jesse Zymet, Andy Luo, Swapnil Shinde +2Adversarial Prompt GenerationLLM Red Teaming