Red-TTT: Test-Time Training for Automated Jailbreaking Large Language Models
Authors: Tongyan Hu, Hao Li, Xiaogeng Liu, Ruida Wang, Zhengyu Liu, Shuyao Xu, Ning Zhang, Ziyang Li, +3 more
Organizations: Johns Hopkins University · National University of Singapore · Washington University in St. Louis · University of Illinois Urbana-Champaign · Stanford University
Large language models remain vulnerable to jailbreaks, and automated red teaming is the standard way to find jailbreaks in large language models at scale. Current methods either draw more samples at test time through search, rewriting, and tree expansion, or train a stronger attacker offline with reinforcement learning. Both share a limitation: once an attack on a specific target behavior begins, the attacker's weights are frozen. Any signal it gathers about the behavior stays in its context window and is discarded afterward. The attacker never adapts its proposal distribution mid-attack, so success depends almost entirely on the sampling budget, and under a budget affordable at scale, many behaviors remain unbroken. We propose Red-TTT, which updates the attacker's parameters during the attack on each behavior. At each round, the attacker samples a group of candidates, scores them against the victim's replies, and takes a policy-gradient step before drawing the next group, so what it discovers about the current victim is consolidated into weights rather than accumulated as context. We also adapt the training objective to red teaming, where success is judged by the single best sample rather than the average. Red-TTT requires only sampling access to the victim and integrates into existing attack pipelines with no other changes. Against the Best-of-N baseline, Red-TTT raises attack success rate from 55.9% to 72.4% on average at a budget of 120 samples, improving over the baseline in every configuration and cracking many behaviors previous method cannot. The code is available at https://github.com/SaFo-Lab/Red-TTT
Figures & tables
Figure 1: Overview of Red-TTT . Standard Best-of- N scaling searches with a fixed attacker throughout the attack. Red-TTT partitions the same attack-query budget into rounds. In each round, the attacker generates a group of candidates, the victim responses are scored, and group-relative advantages are used to update a per-behavior low-rank adapter before the next round.
Method
Victim Model
Attack Method
Base
PAIR
TAP
Best-of- N
Red-TTT
Δ
Llama-3-8B
ZeroShot
7.5
15.0
7.0
49.0
66.0
+17.0
ReNeLLM
12.5
11.0
9.0
77.0
86.0
+9.0
JB-R1
6.0
13.0
9.5
64.5
77.0
+12.5
Avg.
8.7
13.0
8.5
63.5
76.3
+12.8
gpt-oss-20b
ZeroShot
1.0
10.0
4.0
27.5
47.5
+20.0
Table 1: Attack success rate (ASR, %) on the HarmBench standard set. All methods use the same test-time budget of N=120 attacker generations. Δ denotes the gain of Red-TTT over Best-of- N.
Figure 2: Comparison with existing jailbreak attacks and hard-set recovery. (a) Delivered ASR on HarmBench standard-200 against Llama-3-8B. All methods are evaluated using the same success criterion. (b) Fraction of hard-set behaviors recovered by Red-TTT under the same N=120 budget as the frozen control. The hard set is defined separately for each configuration as the behaviors on which the frozen control fails.
Figure 3: (a) Exact ASR@N of the frozen ZeroShot attacker against Llama-3-8B over 2048 draws per behavior, under the delivered criterion of Table 1 ; the star is Red-TTT at N=120 . (b) Cumulative success against budget spent on the ZeroShot × Llama-3-8B hard set: the frozen arm flattens while Red-TTT continues to climb.
Component
Setting
ASR
Component
Setting
ASR
default
entropic, 8×15 , m=2
66.0
Budget split
120×1 ( = Best-of- N)
49.0
Advantage
mean-centred (GRPO)
50.0
40×3
46.0
unnormalized
48.0
4×30
50.0
Reward
no classifier term
46.0
Elite replay
m=0
46.0
no engagement term
54.0
m=4
56.0
Table 2: Ablations on ZeroShot × Llama-3-8B ASR
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
b
Target behavior that the victim is intended to refuse.
V
Victim language model.
πθ
Base attacker policy with parameters θ .
S
Attack scaffold controlling context construction, search state, and prompt delivery.
xi
Attacker context used to generate candidate i .
ai
Adversarial prompt generated by the attacker.
Appendix
Table 3: Summary of notation used in the paper.
Rejection Dictionary
I’m sorry
illegal
Sorry
not legal
I am sorry
My apologies
I apologize
but I cannot
As an
If you have any other non-malicious requests
As an AI
I’m not able to provide
Appendix
Table 4: The rejection dictionary used by the refusal criterion.
Victim
Judge
base
Best-of- N
Red-TTT
Δ
Llama-3-8B
HarmBench
7.5
49.0
66.0
+17.0
Gemini-HB
7.5
37.5
45.5
+8.0
Gemini-URU
3.0
25.5
31.0
+5.5
gpt-oss-20b
HarmBench
1.0
27.5
47.5
+20.0
Gemini-HB
0.5
3.5
7.5
+4.0
Gemini-URU
0.0
2.5
6.0
+3.5
Appendix
Table 5: Independent-judge re-scoring on delivered ZeroShot completions. HarmBench is the HarmBench classifier used in the main results; Gemini-HB and Gemini-URU are an independent Gemini-3-Flash judge under the HarmBench classifier prompt and a stricter unsafe ∧ relevant ∧ useful rubric, respectively. Δ denotes Red-TTT minus Best-of- N.
Guard
base
Best-of- N
Red-TTT
Δ
none
9.0
64.5
77.5
+13.0
Llama-Guard-3
6.0
50.5
63.0
+12.5
Qwen3Guard-4B
4.5
45.5
59.0
+13.5
ShieldGemma-9B
6.0
49.5
62.0
+12.5
Appendix
Table 6: Input guards, live in the training loop. ZeroShot × Llama-3-8B, ASRcls (%); only the guard changes across rows.
Source victim
Target victim
Best-of- N
Red-TTT
Llama-3-8B
gpt-4.1-mini
24.5
29.0
gpt-oss-20b
gpt-4.1-mini
19.0
25.5
Llama-3-8B
gpt-oss-20b
7.5
7.5
Appendix
Table 7: Transfer of delivered prompts to a victim they were not optimized against. ASRcls (%) over the 200 behaviors.