Red-Teaming

Recent momentum

-33%

4 papers in the last 28 days · 0.1% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

1 new paper

A weekly snapshot of new work published in Red-Teaming.

Period ending 2026-09-14

1 new paper

A weekly snapshot of new work published in Red-Teaming.

Period ending 2026-09-07

2 new papers

A weekly snapshot of new work published in Red-Teaming.

67 papers

Latest in Red-Teaming

Open your feed →
CardsList
  1. GPT-Red: Automated Red Teaming via Self-Play at Scale

    Jul 28, 2026Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal +15Red-TeamingLanguage Model Evasion Attacks

  2. Code Monitor Red Teaming for Public-Test-Passing Code

    Jul 23, 2026Junchi Liao, Jiawen Deng, Fuji RenRed-TeamingVerifier

  3. CONTRA: Red-Teaming Configurations of Personalizable Agents

    Jul 3, 2026Jonathan Nöther, Adish Singla, Goran RadanovicRed-TeamingMalicious Code

  4. RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems

    Jun 22, 2026Yarin Yerushalmi Levi, Roy Betser, Amit Giloni +5Red-TeamingMultiple Attack Surfaces

  5. Diffuse AI Control on Fuzzy Tasks

    Jun 8, 2026Mikhail Terekhov, Caglar Gulcehre, Vivek Hebbar +1SabotageRed-Teaming

  6. MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring

    May 10, 2026Monika Jotautaitė, Maria Angelica Martinez, Ollie Matthews +1Red-TeamingCoding Agents

  7. Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours

    May 5, 2026Raja Sekhar Rao Dheekonda, Will Pearce, Nick LandersRed-TeamingLogical Connectives

  8. Training a General Purpose Automated Red Teaming Model

    Apr 24, 2026Aishwarya Padmakumar, Leon Derczynski, Traian Rebedea +1Red-TeamingLanguage Model Evasion Attacks