LLM Red Teaming

LLM: Large Language Model

Latest papers 47

All topics
CardsList
  1. How Far Will They Go? Red-Teaming Online Influence with Large Language Models

    May 20, 2026Daniel C. Ruiz, Anna Serbina, Ashwin Rao +2Language Model SteeringLLM Auditing

  2. Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution

    May 13, 2026Xiaozhe Zhang, Chaozhuo Li, Hui Liu +4Continual Learning for LLMsAdversarial Attacks on LLMs

  3. Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization

    May 11, 2026Huilin Zhou, Jian Zhao, Yilu Zhong +7Inference-Time OptimizationAdversarial Attacks on LLMs

  4. OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents

    May 9, 2026Xinyu Li, Ronghui Mu, Lin Li +2Adversarial Attacks on LLMsLLM Red Teaming

  5. LoopTrap: Termination Poisoning Attacks on LLM Agents

    May 7, 2026Huiyu Xu, Zhibo Wang, Wenhui Zhang +4Adversarial Attacks on LLMsLLM Red Teaming

  6. PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI

    May 7, 2026Wesley Hanwen Deng, Mingxi Yan, Sunnie S. Y. Kim +5Adversarial Prompt GenerationLLM Red Teaming

  7. Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours

    May 5, 2026Raja Sekhar Rao Dheekonda, Will Pearce, Nick LandersAdversarial Attacks on LLMsLLM Red Teaming

  8. ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming

    May 4, 2026Mario Rodríguez Béjar, Francisco J. Cortés-Delgado, S. Braghin +1Evolutionary OptimizationLLM Red Teaming

  9. Training a General Purpose Automated Red Teaming Model

    Apr 24, 2026Aishwarya Padmakumar, Leon Derczynski, Traian Rebedea +1Adversarial Attacks on LLMsLLM Red Teaming

  10. Adaptive Instruction Composition for Automated LLM Red-Teaming

    Apr 22, 2026Jesse Zymet, Andy Luo, Swapnil Shinde +2Adversarial Prompt GenerationLLM Red Teaming

  11. ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System

    Apr 20, 2026Jiacheng Liang, Yao Ma, Tharindu Kumarage +5Reward ModelingAdversarial Attacks on LLMs

  12. T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search

    Mar 21, 2026Hyomin Lee, Sangwoo Park, Yumin Choi +3Adversarial Prompt GenerationLLM Red Teaming

  13. Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment

    Sep 6, 2025Yuchong Xie, Mingyu Luo, Zesen Liu +7Prompt InjectionAI Coding Agents

  14. Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

    Jul 30, 2025Jiazhen Pan, Bailiang Jian, Paul Hager +20LLM Safety BenchmarksLanguage Model Safety Evaluation

  15. RedCoder: Automated Multi-Turn Red Teaming for Code LLMs

    Jun 25, 2025Wenjie Jacky Mo, Qin Liu, Xiaofei Wen +5Adversarial Prompt GenerationLLM Red Teaming

  16. Learning diverse attacks on large language models for robust red-teaming and safety tuning

    May 28, 2024Seanie Lee, Minsu Kim, Lynn Cherif +8Adversarial Attacks on LLMsLLM Red Teaming