cs.CLOct 4, 2026

Red-TTT: Test-Time Training for Automated Jailbreaking Large Language Models

Authors: Tongyan Hu, Hao Li, Xiaogeng Liu, Ruida Wang, Zhengyu Liu, Shuyao Xu, Ning Zhang, Ziyang Li, +3 more

Organizations: Johns Hopkins University · National University of Singapore · Washington University in St. Louis · University of Illinois Urbana-Champaign · Stanford University

Abstract

Large language models remain vulnerable to jailbreaks, and automated red teaming is the standard way to find jailbreaks in large language models at scale. Current methods either draw more samples at test time through search, rewriting, and tree expansion, or train a stronger attacker offline with reinforcement learning. Both share a limitation: once an attack on a specific target behavior begins, the attacker's weights are frozen. Any signal it gathers about the behavior stays in its context window and is discarded afterward. The attacker never adapts its proposal distribution mid-attack, so success depends almost entirely on the sampling budget, and under a budget affordable at scale, many behaviors remain unbroken. We propose Red-TTT, which updates the attacker's parameters during the attack on each behavior. At each round, the attacker samples a group of candidates, scores them against the victim's replies, and takes a policy-gradient step before drawing the next group, so what it discovers about the current victim is consolidated into weights rather than accumulated as context. We also adapt the training objective to red teaming, where success is judged by the single best sample rather than the average. Red-TTT requires only sampling access to the victim and integrates into existing attack pipelines with no other changes. Against the Best-of-N baseline, Red-TTT raises attack success rate from 55.9% to 72.4% on average at a budget of 120 samples, improving over the baseline in every configuration and cracking many behaviors previous method cannot. The code is available at https://github.com/SaFo-Lab/Red-TTT

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AutoRISE: Agent-Driven Strategy Evolution for Red-Teaming Large Language Models

    Apr 23, 2026Tanmay Gautam, Alireza Bahramali, Sandeep AtluriRed-TeamingAttacker Large Language Model

  2. Training a General Purpose Automated Red Teaming Model

    Apr 24, 2026Aishwarya Padmakumar, Leon Derczynski, Traian Rebedea +1Red-TeamingAttacker Large Language Model

  3. LASH: Adaptive Semantic Hybridization for Black-Box Jailbreaking of Large Language Models

    May 20, 2026Abdullah Al Nomaan Nafi, Fnu Suya, Swarup Bhunia +1Large Language Model JailbreaksJailbreak Attacks