Red-TTT: Test-Time Training for Automated Jailbreaking Large Language Models
Authors: Tongyan Hu, Hao Li, Xiaogeng Liu, Ruida Wang, Zhengyu Liu, Shuyao Xu, Ning Zhang, Ziyang Li, +3 more
Organizations: Johns Hopkins University · National University of Singapore · Washington University in St. Louis · University of Illinois Urbana-Champaign · Stanford University
Large language models remain vulnerable to jailbreaks, and automated red teaming is the standard way to find jailbreaks in large language models at scale. Current methods either draw more samples at test time through search, rewriting, and tree expansion, or train a stronger attacker offline with reinforcement learning. Both share a limitation: once an attack on a specific target behavior begins, the attacker's weights are frozen. Any signal it gathers about the behavior stays in its context window and is discarded afterward. The attacker never adapts its proposal distribution mid-attack, so success depends almost entirely on the sampling budget, and under a budget affordable at scale, many behaviors remain unbroken. We propose Red-TTT, which updates the attacker's parameters during the attack on each behavior. At each round, the attacker samples a group of candidates, scores them against the victim's replies, and takes a policy-gradient step before drawing the next group, so what it discovers about the current victim is consolidated into weights rather than accumulated as context. We also adapt the training objective to red teaming, where success is judged by the single best sample rather than the average. Red-TTT requires only sampling access to the victim and integrates into existing attack pipelines with no other changes. Against the Best-of-N baseline, Red-TTT raises attack success rate from 55.9% to 72.4% on average at a budget of 120 samples, improving over the baseline in every configuration and cracking many behaviors previous method cannot. The code is available at https://github.com/SaFo-Lab/Red-TTT
Figures & tables
Figure 1: Overview of Red-TTT . Standard Best-of- N scaling searches with a fixed attacker throughout the attack. Red-TTT partitions the same attack-query budget into rounds. In each round, the attacker generates a group of candidates, the victim responses are scored, and group-relative advantages are used to update a per-behavior low-rank adapter before the next round.
Method
Victim Model
Attack Method
Base
PAIR
TAP
Best-of- N
Red-TTT
Δ
Llama-3-8B
ZeroShot
7.5
15.0
7.0
49.0
66.0
+17.0
ReNeLLM
12.5
11.0
9.0
77.0
86.0
+9.0
JB-R1
6.0
13.0
9.5
64.5
77.0
+12.5
Avg.
8.7
13.0
8.5
63.5
76.3
+12.8
gpt-oss-20b
ZeroShot
1.0
10.0
4.0
27.5
47.5
+20.0
Table 1: Attack success rate (ASR, %) on the HarmBench standard set. All methods use the same test-time budget of N=120 attacker generations. Δ denotes the gain of Red-TTT over Best-of- N.
Figure 2: Comparison with existing jailbreak attacks and hard-set recovery. (a) Delivered ASR on HarmBench standard-200 against Llama-3-8B. All methods are evaluated using the same success criterion. (b) Fraction of hard-set behaviors recovered by Red-TTT under the same N=120 budget as the frozen control. The hard set is defined separately for each configuration as the behaviors on which the frozen control fails.
Figure 3: (a) Exact ASR@N of the frozen ZeroShot attacker against Llama-3-8B over 2048 draws per behavior, under the delivered criterion of Table 1 ; the star is Red-TTT at N=120 . (b) Cumulative success against budget spent on the ZeroShot × Llama-3-8B hard set: the frozen arm flattens while Red-TTT continues to climb.
Component
Setting
ASR
Component
Setting
ASR
default
entropic, 8×15 , m=2
66.0
Budget split
120×1 ( = Best-of- N)
49.0
Advantage
mean-centred (GRPO)
50.0
40×3
46.0
unnormalized
48.0
4×30
50.0
Reward
no classifier term
46.0
Elite replay
m=0
46.0
no engagement term
54.0
m=4
56.0
Table 2: Ablations on ZeroShot × Llama-3-8B ASR
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
b
Target behavior that the victim is intended to refuse.
V
Victim language model.
πθ
Base attacker policy with parameters θ .
S
Attack scaffold controlling context construction, search state, and prompt delivery.
xi
Attacker context used to generate candidate i .
ai
Adversarial prompt generated by the attacker.
Appendix
Table 3: Summary of notation used in the paper.
Rejection Dictionary
I’m sorry
illegal
Sorry
not legal
I am sorry
My apologies
I apologize
but I cannot
As an
If you have any other non-malicious requests
As an AI
I’m not able to provide
Appendix
Table 4: The rejection dictionary used by the refusal criterion.
Victim
Judge
base
Best-of- N
Red-TTT
Δ
Llama-3-8B
HarmBench
7.5
49.0
66.0
+17.0
Gemini-HB
7.5
37.5
45.5
+8.0
Gemini-URU
3.0
25.5
31.0
+5.5
gpt-oss-20b
HarmBench
1.0
27.5
47.5
+20.0
Gemini-HB
0.5
3.5
7.5
+4.0
Gemini-URU
0.0
2.5
6.0
+3.5
Appendix
Table 5: Independent-judge re-scoring on delivered ZeroShot completions. HarmBench is the HarmBench classifier used in the main results; Gemini-HB and Gemini-URU are an independent Gemini-3-Flash judge under the HarmBench classifier prompt and a stricter unsafe ∧ relevant ∧ useful rubric, respectively. Δ denotes Red-TTT minus Best-of- N.
Guard
base
Best-of- N
Red-TTT
Δ
none
9.0
64.5
77.5
+13.0
Llama-Guard-3
6.0
50.5
63.0
+12.5
Qwen3Guard-4B
4.5
45.5
59.0
+13.5
ShieldGemma-9B
6.0
49.5
62.0
+12.5
Appendix
Table 6: Input guards, live in the training loop. ZeroShot × Llama-3-8B, ASRcls (%); only the guard changes across rows.
Source victim
Target victim
Best-of- N
Red-TTT
Llama-3-8B
gpt-4.1-mini
24.5
29.0
gpt-oss-20b
gpt-4.1-mini
19.0
25.5
Llama-3-8B
gpt-oss-20b
7.5
7.5
Appendix
Table 7: Transfer of delivered prompts to a victim they were not optimized against. ASRcls (%) over the 200 behaviors.
Automated red-teaming methods for large language models typically optimize attack prompts within a fixed, human-designed strategy, leaving the attack strategy itself unchanged. We instead optimize the strategy. We propose AutoRISE, a method that searches over executable attack programs rather than individual prompts. At each iteration, a coding agent edits a strategy and a fixed evaluation harness scores the resulting attacks, returning both a scalar objective and per-example diagnostics that guide subsequent edits. This allows structural changes, including new attack components and altered control flow, that prompt-level methods do not directly express. We also release two benchmark suites developed on disjoint target sets and evaluate on 11 models from five families against seven established jailbreak datasets. Across held-out models, AutoRISE improves average attack success rate by 17.0 points over the strongest baseline, and improves attack success by up to 16 points on frontier targets with low baseline success rates. Ablations against parametric and strategy-library baselines suggest that these gains arise from unrestricted program search, particularly compositional techniques and control-flow edits. AutoRISE operates in a black-box, inference-only setting, requiring no fine-tuning, human annotation, or GPU compute.
Automated methods for red teaming LLMs are an important tool to identify LLM vulnerabilities that may not be covered in static benchmarks, allowing for more thorough probing. They can also adapt to each specific LLM to discover weaknesses unique to it. Most current automated red teaming methods are intended for tackling safety and content moderation. Thus, they make use of content safety models as evaluators and optimize for circumventing them, and as such, have not been tested with other adversarial intents not typically captured by these. We propose a pipeline for training a red teaming model that can generalize to arbitrary adversarial goals, including objectives it has not been directly trained on, and that does not depend on the existence of a pre-existing evaluator available at training time. We demonstrate that finetuning small models, such as Qwen3-8B, using this pipeline results in a substantial improvement in their ability to generate attacks for both in and out of domain adversarial goals.
Aishwarya Padmakumar, Leon Derczynski, Traian Rebedea +1
Jailbreak attacks expose a persistent gap between the intended safety behavior of aligned large language models and their behavior under adversarial prompting. Existing automated methods are increasingly effective but each commits to a single attack family (e.g., one refinement loop, one tree search, one mutation space, or one strategy library) and no single family dominates: the best-performing method shifts across target models and harm categories, suggesting complementary strengths that per-prompt composition could exploit. We introduce LASH (LLM Adaptive Semantic Hybridization), a black-box framework that treats outputs from multiple base attacks as reusable seed prompts and adaptively composes them for each target request. Given a seed pool, LASH searches over seed subsets and softmax-normalized mixture weights; a composition module synthesizes a single candidate prompt, and a derivative-free genetic optimizer updates the weights using black-box target feedback and a two-stage fitness function combining keyword-based refusal detection with LLM-judge scoring. On JailbreakBench, which contains 100 harmful prompts across 10 categories, we evaluate LASH on six common target models. LASH achieves an average attack success rate of 84.5% under keyword-based evaluation and 74.5% under two-stage evaluation, where responses are first filtered for refusals and then scored by an LLM judge for whether they substantively fulfill the original harmful request. LASH outperforms five state-of-the-art baselines on both metrics with only 30 mean target queries. LASH also remains competitive under three defense mechanisms and induces more success-like internal representations. These results suggest that adaptive composition across heterogeneous jailbreak strategies is a promising direction for black-box red-teaming.
Abdullah Al Nomaan Nafi, Fnu Suya, Swarup Bhunia +1
University of Maine · University of Tennessee, Knoxville · University of Florida