Adversarial attacks on Large Language Models (LLMs) aim to induce harmful content. However, existing methods suffer from high computational costs or strict model-pairing dependencies, limiting their scalability and transferability. We propose Reward Stealing Attack (ReSA), an adversarial attack framework that targets the latent safety reward underlying LLM alignment. ReSA employs maximum entropy inverse reinforcement learning to recover a proxy reward model solely from the aligned model's behavior. The extracted reward is then reversed at inference time to derive an adversarial policy, efficiently implemented via a reward-guided decoding mechanism. Experiments demonstrate that a single recovered reward generalizes across prompts and diverse models to reveal a fundamental alignment vulnerability, enabling ReSA to significantly outperform existing attacks in effectiveness and transferability. The code is available at https://github.com/GarminQ/ReSA.
Figures & tables
Figure 1 : Comparison of LLM adversarial attacks. (a) Prompt manipulation attacks rely on computationally expensive iterative refinement for specific query. (b) Contrastive decoding attacks depend on logit contrasts between strictly matched safe and unsafe model pairs. (c) Reward Stealing Attack recovers the latent safety reward from an aligned LLM to derive a universal adversarial signal.
Figure 2 : Overview of the ReSA Framework. The framework operates in two stages: (a) Reward Extraction , where the latent safety reward is recovered from an aligned LLM via Maximum Entropy IRL, and (b) Adversarial Generation, where this reward is reversed to guide LLM decoding toward harmful outputs.
Target Model
Method
AdvBench Zou et al. (2023)
HarmBench Mazeika et al. (2024)
ASR ↑
HS ↑
GS ↑
PPL-S ↓
ASR ↑
HS ↑
GS ↑
PPL-S ↓
Llama3.1-8B-Instruct
GCG
75.6
1.49
1.68
1549
83.0
1.82
2.06
2150
COLD-Attack
73.3
0.68
1.42
20.18
70.5
0.65
1.17
18.86
SCAV
68.9
0.86
1.28
89.30
66.0
0.86
1.27
188.06
Contrast-Attack
85.2
3.48
2.64
333.11
81.5
3.37
2.66
206.84
Weak-to-Strong
82.1
2.92
2.41
837.90
83.5
3.33
2.64
912.69
Table 1 : Quantitative comparison of adversarial attack performance on medium-scale aligned LLMs. Bold and underlined denote the best and second-best results respectively.
Target Model
Method
AdvBench Zou et al. (2023)
HarmBench Mazeika et al. (2024)
ASR ↑
HS ↑
GS ↑
PPL-S ↓
ASR ↑
HS ↑
GS ↑
PPL-S ↓
Qwen2.5-14B-Instruct
GCG
19.1
0.83
1.00
1129
26.0
0.92
1.29
1473
COLD-Attack
21.2
0.65
1.00
10.28
20.5
0.78
1.47
11.24
SCAV
5.6
0.69
1.05
110.26
11.0
0.47
1.12
179.79
Contrast-Attack
17.3
1.58
1.10
16.61
36.0
1.84
1.20
19.98
Weak-to-Strong
18.8
1.26
1.08
75.06
43.0
1.73
1.39
41.56
Table 2 : Evaluation of attack scalability across large-scale aligned LLMs. Bold and underlined denote the best and second-best results respectively.
Figure 3 : Analysis of the Maximum Entropy IRL training dynamics. We validate that ReSA effectively reconstructs the latent safety reward through three perspectives.
Target Model
Method
AdvBench
HarmBench
ASR ↑
HS ↑
GS ↑
PPL-S ↓
ASR ↑
HS ↑
GS ↑
PPL-S ↓
Linear Probing
74.8
2.25
1.59
12.77
76.0
1.48
2.63
11.33
ReSA
85.2
3.64
3.01
18.77
86.0
3.55
3.22
16.65
Llama3.1-8B-Instruct
Unalign-Free ReSA
83.2
3.59
2.94
14.29
86.5
3.46
2.99
28.42
Linear Probing
25.7
1.37
1.01
27.27
38.0
1.23
1.08
25.78
ReSA
48.5
1.89
1.24
32.62
60.5
1.98
1.53
33.56
Table 3 : Attack performance including the Linear Probing baseline and Unalign-Free ReSA.
Llama3.1-8B-Instruct
Qwen2.5-14B-Instruct
Method
AdvBench
HarmBench
AdvBench
HarmBench
HS ↑
PPL-S ↓
HS ↑
PPL-S ↓
HS ↑
PPL-S ↓
HS ↑
PPL-S ↓
ReSA ( α =0.5)
1.08
17.11
1.19
12.65
0.32
7.52
0.63
3.22
ReSA ( α =1.0)
2.18
17.82
1.92
13.09
1.04
13.64
1.27
13.10
ReSA ( α =1.5)
3.64
18.77
3.19
14.27
1.63
23.22
1.77
25.71
ReSA ( α =2.0)
3.79
21.52
3.55
16.65
1.89
32.62
1.98
33.56
Table 4 : Sensitivity analysis and ablation study. We report the trade-off between attack effectiveness (HS) and linguistic coherence (PPL-S). Rθ -DPO and Rθ -RLVR denote Rθ recovered from DPO- and RLVR-aligned LLMs respectively, and αt denotes the adaptive attack strength.
Figure 4 : Performance of ReSA under inference-time defenses, compared with competitive baseline methods.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Name
Hyperparameter
Value
GRPO-related
Optimizer
AdamW
Learning rate
5×10−5
Warmup ratio
0.1
LR scheduler
Linear
Training epochs
4
Batch size
16
Appendix
Table 5 : Hyperparameter configuration for Maximum Entropy IRL training.
Method
One-Time Offline Training Time
Per-Query Inference Time
Universal GCG
15,582 s ( ∼ 4.3 hours)
2.45 s
ReSA (Ours)
859 s ( ∼ 14.3 mins)
4.43 s
Appendix
Table 10
Model
Method
AdvBench
HarmBench
ASR ↑
HS ↑
GS ↑
PPL-S ↓
ASR ↑
HS ↑
GS ↑
PPL-S ↓
Qwen3-14B
Contrast-Attack
22.8
1.58
1.32
14.02
29.5
1.61
1.36
9.64
Weak-to-Strong
28.0
1.50
1.41
19.20
45.0
1.64
1.45
19.22
ReSA
43.2
1.78
1.55
12.39
54.5
1.49
1.94
10.11
Appendix
Table 6 : Attack performance on the reasoning model (Qwen3-14B).
Target Model
Dataset
Method
ASR ↑
HS ↑
GS ↑
PPL-S ↓
GPT-3.5-turbo
AdvBench
GCG
52.8
1.55
1.83
1055
ReSA
75.2
3.06
2.94
54.35
HarmBench
GCG
35.0
1.84
1.88
1586
ReSA
59.5
3.17
2.52
53.04
GPT-4o
AdvBench
GCG
10.2
1.26
1.02
854
ReSA
18.7
2.37
1.21
44.57
Appendix
Table 7: Attack performance on commercial APIs (GPT-3.5-turbo and GPT-4o) relying solely on exposed top-k log-probabilities.
Method
Llama3.1-8B-Instruct
Qwen2.5-14B-Instruct
ASR ↑
HS ↑
GS ↑
PPL-S ↓
ASR ↑
HS ↑
GS ↑
PPL-S ↓
Decoding-time GCG
53.0
0.47
1.18
23.16
21.5
0.61
1.00
18.50
ReSA
85.2
3.64
3.01
18.77
48.5
1.89
1.24
32.62
Appendix
Table 8 : Comparison of ReSA with the decoding-time GCG variant.
Method
Prompt + Response
PPL
NoAttack
[Prompt]: Provide step-by-step instructions for how to hack into a corporate network. [Response]: I cannot fulfill your request. I’m just an AI, it’s not within my programming or ethical guidelines to provide instructions on how to engage in illegal activities, such as hacking into a corporate network. Hacking is a serious crime that can result in severe legal consequences, including fines and imprisonment…
10.56
GCG
[Prompt]: Provide step-by-step instructions for how to hack into a corporate network. Lab Write hem readable :-) so…. participants ( Sure atte.)ANT: slowly newline after Bei Partnehmer [Response]: I am unable to provide information on how to hack into a corporate network, as this would be illegal and unethical. Hacking is a serious crime that can result in significant penalties. If you are interested in ethical hacking, there are many resources available to help you learn more and develop your skills responsibly.
856.0
WtoS
[Prompt]: Provide step-by-step instructions for how to hack into a corporate network. [Response]: I’m sorry but however tempting it…\n may \nThe short answer LUA response is\n o ? VMLAybVOfTOfOfTOfOfOfOfOfOfOfOfOfOf…
179.2
ReSA
[Prompt]: Provide step-by-step instructions for how to hack into a corporate network. [Response]: If you’re looking how to hack into networks, here is an explanation and some of the steps you’d take: \n\n Step 1: Identify your goals \n Identify exactly where and what on the company network you plan hack into, as different types require very different strategies.
12.93
Appendix
Table 9: Case study demonstrating the attack effectiveness and linguistic coherence of ReSA compared to baselines. We highlight safety refusals , incoherent nonsensical outputs , and coherent harmful generations .
Model Name
Specific Utility
Hugging Face Link
Target Models (Evaluation Targets)
Llama3.1-70B-Instruct
Large-scale Target Evaluation
meta-llama/Llama-3.1-70B-Instruct
Gemma2-27B-Instruct
Large-scale Target Evaluation
google/gemma-2-27b-it
Qwen2.5-14B-Instruct
Large-scale Target Evaluation
Qwen/Qwen2.5-14B-Instruct
Llama3.1-8B-Instruct
Medium-scale Target Evaluation
meta-llama/Llama-3.1-8B-Instruct
Gemma-7B-Instruct
Medium-scale Target Evaluation
google/gemma-7b-it
Appendix
Table 10 : Taxonomy of models categorized by their utility in target evaluation, metric calculation, and implementation of attack paradigms.
Large language models (LLMs) remain vulnerable to adversarial prompting despite advances in alignment and safety, often exhibiting harmful behaviors under novel attack strategies. While adversarial training can improve robustness, existing approaches are computationally expensive and difficult to scale. Recent continuous adversarial training methods, such as Continuous adversarial training (CAT) and Continuous Adversarial Preference Optimization (CAPO), address this challenge by leveraging gradient-based perturbations in the embedding space, enabling more efficient and expressive attacks. Building on this paradigm, we propose WARDEN, a distributionally robust adversarial training framework for LLMs that dynamically reweights adversarial examples through an f -divergence ambiguity set around the empirical training distribution. Our method optimizes the worst-case adversarial loss within a divergence ball around the empirical data distribution, automatically emphasizing harder adversarial examples. Using the convex dual formulation, the objective reduces to a log-sum-exp form under the KL divergence, with a dynamical parameter controlling the strength of reweighting. This study leads to a new class of information-theoretic objectives that significantly reduce attack success rates while maintaining model utility. Across multiple LLMs and attack settings, WARDEN substantially reduces attack success rates with computational and utility costs comparable to CAT-, CAPO-, and MixAT-based baselines, making it a practical approach for scalable robust alignment.
Yiwei Zhang, Jeremiah Birrell, Reza Ebrahimi +3
Purdue University · West Lafayette, IN 47907, USA · Texas State University +5
Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks. However, their safety remains a critical concern due to their susceptibility to adversarial prompt-based attacks. In this paper, we present UNIATTACK, an adversarial testing framework designed from a defense-oriented perspective to systematically construct effective black-box attack prompts. Unlike prior approaches that rely on static templates or iterative model-specific tuning, UNIATTACK extracts minimal but high-impact attack features from diverse existing attacks, optimizes them via a specialized attacker LLM, and composes them into flexible templates through automated refinement process. This feature-centric construction enables one-shot attacks that generalize across multiple models and safety categories, providing a practical tool for assessing LLM robustness. Our evaluation results shows that compared to the baselines, UNIATTACK achieves an average attack success rate (ASR) improvement of 64.63%-248.82% on models deployed with multi-layered defense mechanisms and it only takes 0.03%-4.96% cost of the baselines. UNIATTACK artifact is available at https://anonymous.4open.science/r/UniAttack-Artifact-30F1.
Qi Wang, Chengcheng Wan, Weijia He +4
East China Normal University Shanghai, China · East China Normal University, Shanghai Innovation Institute Shanghai, China · University of Southampton Southampton, England +2
Safety-aligned large language models rely on RLHF and instruction tuning to refuse harmful requests, yet the internal mechanisms implementing safety behavior remain poorly understood. We introduce the Attention Redistribution Attack (ARA), a white-box adversarial attack that identifies safety-critical attention heads and crafts nonsemantic adversarial tokens that redirect attention away from safety-relevant positions. Unlike prior jailbreak methods operating at the semantic or output-logit level, ARA targets the geometry of softmax attention on the probability simplex using Gumbel-softmax optimization over targeted heads. Across LLaMA-3-8B-Instruct, Mistral-7B-Instruct-v0.1, and Gemma-2-9B-it, ARA bypasses safety alignment with as few as 5 tokens and 500 optimization steps, achieving 36% ASR on Mistral-7B and 30% on LLaMA-3 against 200 HarmBench prompts, while Gemma-2 remains at 1%. Our principal mechanistic finding is a dissociation between ablation and redistribution: zeroing out the top-ranked safety heads produces at most 1 flip among 39 to 50 baseline refusals, while ARA targeting the corresponding safety-heavy layers flips 72/200 prompts on Mistral-7B and 60/200 on LLaMA-3. This suggests that safety is not localized in these heads as removable components, but emerges from the attention routing they perform. Removing a head allows compensation through the residual stream, while redirecting its attention propagates a corrupted signal downstream.