LLM Red Teaming

LLM: Large Language Model

Latest papers 47

All topics
CardsList
  1. No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents

    Sep 30, 2026Ayan Javeed Shaikh, Arunesh Sinha, Nathaniel D. Bastian +1LLM Red TeamingAI Agent Evaluation

  2. Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning

    Sep 27, 2026Chenlong Yin, Xiaolong Jin, Wei Zou +2Adversarial Attacks on LLMsLLM Red Teaming

  3. REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems

    Aug 11, 2026Zixing Chen, Xingyuan Liu, Jie Zhu +6LLM Red TeamingLLM Agents

  4. Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

    Aug 5, 2026Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi +6Autonomous Cyber DefenseLLM Red Teaming

  5. GPT-Red: Automated Red Teaming via Self-Play at Scale

    Jul 28, 2026Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal +15Adversarial TrainingLLM Red Teaming

  6. AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation

    Jul 13, 2026Yi Ting Shen, Kentaroh Toyoda, Alex LeungLLM Safety BenchmarksAdversarial Attacks on LLMs

  7. CONTRA: Red-Teaming Configurations of Personalizable Agents

    Jul 3, 2026Jonathan Nöther, Adish Singla, Goran RadanovicLLM Red TeamingAI Agent Security

  8. MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG

    Jun 25, 2026Inderjeet Singh, Andrés Murillo, Motoyoshi Sekiya +2Multimodal RAGTool-Augmented Language Model Agents

  9. Distributed Quality-Diversity Search for Toxicity in Large Language Models

    Jun 23, 2026Onkar Shelar, Travis DesellAdversarial Prompt GenerationLLM Red Teaming

  10. OTTER: A Red-Teaming System for Toxicity-Evading Jailbreak Prompt Optimization

    Jun 19, 2026Jerry Wang, Hsin-Ling Hsu, Yi-Cheng Lai +2Adversarial Prompt GenerationAdversarial Attacks on LLMs

  11. NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms

    Jun 18, 2026Hanwool Lee, Dasol Choi, Bokyeong Kim +2Multi-Agent LLM SystemsSafety-Critical Control

  12. FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming

    Jun 18, 2026Chaeyun Kim, Daeyoung Park, Junghwan Kim +4LLM Safety BenchmarksFinancial Services

  13. A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models

    Jun 16, 2026Nicola FrancoLanguage Model Safety EvaluationNeural Network Robustness

  14. Securing Multi-Agent GIS Systems: Risk Evaluation and Prompt Hardening Optimization

    Jun 13, 2026Kyle Gao, Pranavi Kotta, Linlin Xu +2Adversarial RobustnessLLM Red Teaming

  15. PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections

    Jun 10, 2026Pengfei He, Lesly Miculicich, Vishesh Sharma +5AI Agent AuditingLLM Auditing

  16. Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO

    Jun 8, 2026Blake Bullwinkel, Eugenia Kim, Amanda Minnich +1Adversarial TrainingAdversarial Attacks on LLMs

  17. VATS: Exploiting Implicit Authority in Error-Path Injection via Systematic Mutation

    Jun 6, 2026Harshil Patel, Kunal PaiLLM Red TeamingAI Agent Security

  18. CHASE: Adversarial Red-Blue Teaming for Improving LLM Safety using Reinforcement Learning

    Jun 4, 2026Rahul Markasserithodi, Aditya Joshi, Yuekang Li +3Adversarial TrainingAdversarial Prompt Generation

  19. Understanding Safety-Sensitive Expert Behavior in Mixture-of-Experts LLMs

    May 28, 2026Zhibo Zhang, Yuxi Li, Zhen Ouyang +2Mixture-of-Experts Language ModelsLLM Auditing

  20. Constitutional Arms Races in the Public Goods Game: Co-Evolving LLM Constitutions Under Cooperation-Defection Pressure

    May 26, 2026Ujwal Kumar, Arth Singh, Hershraj Niranjani +5LLM Red TeamingMulti-Agent LLMs