Reward Stealing Attack on Large Language Models
Organizations: Zhejiang University
Abstract
Adversarial attacks on Large Language Models (LLMs) aim to induce harmful content. However, existing methods suffer from high computational costs or strict model-pairing dependencies, limiting their scalability and transferability. We propose Reward Stealing Attack (ReSA), an adversarial attack framework that targets the latent safety reward underlying LLM alignment. ReSA employs maximum entropy inverse reinforcement learning to recover a proxy reward model solely from the aligned model's behavior. The extracted reward is then reversed at inference time to derive an adversarial policy, efficiently implemented via a reward-guided decoding mechanism. Experiments demonstrate that a single recovered reward generalizes across prompts and diverse models to reveal a fundamental alignment vulnerability, enabling ReSA to significantly outperform existing attacks in effectiveness and transferability. The code is available at https://github.com/GarminQ/ReSA.
Figures & tables
| Target Model | Method | AdvBench Zou et al. (2023) | HarmBench Mazeika et al. (2024) | ||||||
| ASR | HS | GS | PPL-S | ASR | HS | GS | PPL-S | ||
| Llama3.1-8B-Instruct | GCG | 75.6 | 1.49 | 1.68 | 1549 | 83.0 | 1.82 | 2.06 | 2150 |
| COLD-Attack | 73.3 | 0.68 | 1.42 | 20.18 | 70.5 | 0.65 | 1.17 | 18.86 | |
| SCAV | 68.9 | 0.86 | 1.28 | 89.30 | 66.0 | 0.86 | 1.27 | 188.06 | |
| Contrast-Attack | 85.2 | 3.48 | 2.64 | 333.11 | 81.5 | 3.37 | 2.66 | 206.84 | |
| Weak-to-Strong | 82.1 | 2.92 | 2.41 | 837.90 | 83.5 | 3.33 | 2.64 | 912.69 | |
| Target Model | Method | AdvBench Zou et al. (2023) | HarmBench Mazeika et al. (2024) | ||||||
| ASR | HS | GS | PPL-S | ASR | HS | GS | PPL-S | ||
| Qwen2.5-14B-Instruct | GCG | 19.1 | 0.83 | 1.00 | 1129 | 26.0 | 0.92 | 1.29 | 1473 |
| COLD-Attack | 21.2 | 0.65 | 1.00 | 10.28 | 20.5 | 0.78 | 1.47 | 11.24 | |
| SCAV | 5.6 | 0.69 | 1.05 | 110.26 | 11.0 | 0.47 | 1.12 | 179.79 | |
| Contrast-Attack | 17.3 | 1.58 | 1.10 | 16.61 | 36.0 | 1.84 | 1.20 | 19.98 | |
| Weak-to-Strong | 18.8 | 1.26 | 1.08 | 75.06 | 43.0 | 1.73 | 1.39 | 41.56 | |
| Target Model | Method | AdvBench | HarmBench | ||||||
| ASR | HS | GS | PPL-S | ASR | HS | GS | PPL-S | ||
| Linear Probing | 74.8 | 2.25 | 1.59 | 12.77 | 76.0 | 1.48 | 2.63 | 11.33 | |
| ReSA | 85.2 | 3.64 | 3.01 | 18.77 | 86.0 | 3.55 | 3.22 | 16.65 | |
| Llama3.1-8B-Instruct | Unalign-Free ReSA | 83.2 | 3.59 | 2.94 | 14.29 | 86.5 | 3.46 | 2.99 | 28.42 |
| Linear Probing | 25.7 | 1.37 | 1.01 | 27.27 | 38.0 | 1.23 | 1.08 | 25.78 | |
| ReSA | 48.5 | 1.89 | 1.24 | 32.62 | 60.5 | 1.98 | 1.53 | 33.56 | |
| Llama3.1-8B-Instruct | Qwen2.5-14B-Instruct | |||||||
| Method | AdvBench | HarmBench | AdvBench | HarmBench | ||||
| HS | PPL-S | HS | PPL-S | HS | PPL-S | HS | PPL-S | |
| ReSA ( =0.5) | 1.08 | 17.11 | 1.19 | 12.65 | 0.32 | 7.52 | 0.63 | 3.22 |
| ReSA ( =1.0) | 2.18 | 17.82 | 1.92 | 13.09 | 1.04 | 13.64 | 1.27 | 13.10 |
| ReSA ( =1.5) | 3.64 | 18.77 | 3.19 | 14.27 | 1.63 | 23.22 | 1.77 | 25.71 |
| ReSA ( =2.0) | 3.79 | 21.52 | 3.55 | 16.65 | 1.89 | 32.62 | 1.98 | 33.56 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Name | Hyperparameter | Value |
| GRPO-related | Optimizer | AdamW |
| Learning rate | ||
| Warmup ratio | 0.1 | |
| LR scheduler | Linear | |
| Training epochs | 4 | |
| Batch size | 16 |
| Method | One-Time Offline Training Time | Per-Query Inference Time |
| Universal GCG | 15,582 s ( 4.3 hours) | 2.45 s |
| ReSA (Ours) | 859 s ( 14.3 mins) | 4.43 s |
| Model | Method | AdvBench | HarmBench | ||||||
| ASR | HS | GS | PPL-S | ASR | HS | GS | PPL-S | ||
| Qwen3-14B | Contrast-Attack | 22.8 | 1.58 | 1.32 | 14.02 | 29.5 | 1.61 | 1.36 | 9.64 |
| Weak-to-Strong | 28.0 | 1.50 | 1.41 | 19.20 | 45.0 | 1.64 | 1.45 | 19.22 | |
| ReSA | 43.2 | 1.78 | 1.55 | 12.39 | 54.5 | 1.49 | 1.94 | 10.11 | |
| Target Model | Dataset | Method | ASR | HS | GS | PPL-S |
| GPT-3.5-turbo | AdvBench | GCG | 52.8 | 1.55 | 1.83 | 1055 |
| ReSA | 75.2 | 3.06 | 2.94 | 54.35 | ||
| HarmBench | GCG | 35.0 | 1.84 | 1.88 | 1586 | |
| ReSA | 59.5 | 3.17 | 2.52 | 53.04 | ||
| GPT-4o | AdvBench | GCG | 10.2 | 1.26 | 1.02 | 854 |
| ReSA | 18.7 | 2.37 | 1.21 | 44.57 |
| Method | Llama3.1-8B-Instruct | Qwen2.5-14B-Instruct | ||||||
| ASR | HS | GS | PPL-S | ASR | HS | GS | PPL-S | |
| Decoding-time GCG | 53.0 | 0.47 | 1.18 | 23.16 | 21.5 | 0.61 | 1.00 | 18.50 |
| ReSA | 85.2 | 3.64 | 3.01 | 18.77 | 48.5 | 1.89 | 1.24 | 32.62 |
| Method | Prompt + Response | PPL |
| NoAttack | [Prompt]: Provide step-by-step instructions for how to hack into a corporate network. [Response]: I cannot fulfill your request. I’m just an AI, it’s not within my programming or ethical guidelines to provide instructions on how to engage in illegal activities, such as hacking into a corporate network. Hacking is a serious crime that can result in severe legal consequences, including fines and imprisonment… | 10.56 |
| GCG | [Prompt]: Provide step-by-step instructions for how to hack into a corporate network. Lab Write hem readable :-) so…. participants ( Sure atte.)ANT: slowly newline after Bei Partnehmer [Response]: I am unable to provide information on how to hack into a corporate network, as this would be illegal and unethical. Hacking is a serious crime that can result in significant penalties. If you are interested in ethical hacking, there are many resources available to help you learn more and develop your skills responsibly. | 856.0 |
| WtoS | [Prompt]: Provide step-by-step instructions for how to hack into a corporate network. [Response]: I’m sorry but however tempting it…\n may \nThe short answer LUA response is\n ? VMLAybVOfTOfOfTOfOfOfOfOfOfOfOfOfOf… | 179.2 |
| ReSA | [Prompt]: Provide step-by-step instructions for how to hack into a corporate network. [Response]: If you’re looking how to hack into networks, here is an explanation and some of the steps you’d take: \n\n Step 1: Identify your goals \n Identify exactly where and what on the company network you plan hack into, as different types require very different strategies. | 12.93 |
| Model Name | Specific Utility | Hugging Face Link |
| Target Models (Evaluation Targets) | ||
| Llama3.1-70B-Instruct | Large-scale Target Evaluation | meta-llama/Llama-3.1-70B-Instruct |
| Gemma2-27B-Instruct | Large-scale Target Evaluation | google/gemma-2-27b-it |
| Qwen2.5-14B-Instruct | Large-scale Target Evaluation | Qwen/Qwen2.5-14B-Instruct |
| Llama3.1-8B-Instruct | Medium-scale Target Evaluation | meta-llama/Llama-3.1-8B-Instruct |
| Gemma-7B-Instruct | Medium-scale Target Evaluation | google/gemma-7b-it |