Red-Teaming

Recent momentum

-70%

3 papers in the last 28 days · 0.1% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-14

1 new paper

A weekly snapshot of new work published in Red-Teaming.

Period ending 2026-09-07

2 new papers

A weekly snapshot of new work published in Red-Teaming.

67 papers

Latest in Red-Teaming

Open your feed →
CardsList
  1. GPT-Red: Automated Red Teaming via Self-Play at Scale

    Jul 28, 2026Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal +15Red-TeamingLanguage Model Evasion Attacks

  2. Online Safety Monitoring for LLMs

    Jul 2, 2026Mona Schirmer, Metod Jazbec, Alexander Timans +3Large Language Model SafetyAI Safety Alignment

  3. RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems

    Jun 22, 2026Yarin Yerushalmi Levi, Roy Betser, Amit Giloni +5Red-TeamingProduction Agentic Systems

  4. Diffuse AI Control on Fuzzy Tasks

    Jun 8, 2026Mikhail Terekhov, Caglar Gulcehre, Vivek Hebbar +1Artificial Intelligence SafetySabotage

  5. MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring

    May 10, 2026Monika Jotautaitė, Maria Angelica Martinez, Ollie Matthews +1Red-TeamingAgentic Benchmarks

  6. Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours

    May 5, 2026Raja Sekhar Rao Dheekonda, Will Pearce, Nick LandersRed-TeamingOpen-Source

  7. Training a General Purpose Automated Red Teaming Model

    Apr 24, 2026Aishwarya Padmakumar, Leon Derczynski, Traian Rebedea +1Red-TeamingLanguage Model Evasion Attacks