Organizations: John A. Paulson School of Engineering And Applied Sciences, Harvard University · Department of Brain and Cognitive Sciences, Massachusetts Institute of Technology · Speech and Hearing Bioscience and Technology, Harvard Medical School · Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University · Center for Brain Science, Harvard University
Adversarial attacks can reliably steer safety-aligned large language models toward unsafe behavior. Empirically, we find that adversarial prompt-injection attacks can amplify attack success rate from the slow polynomial growth observed without injection to exponential growth with the number of inference-time samples. We first identify a minimal statistical mechanism for these two regimes by giving a small set of assumptions on the distribution of safe generation across contexts under which both scaling laws follow. To explain this phenomenon further, we propose a theoretical generative model of proxy language in terms of a spin-glass system operating in a replica-symmetry-breaking regime, where generations are drawn from the associated Gibbs measure and a subset of low-energy, size-biased clusters is designated unsafe. We analytically show how this model naturally realizes the minimal assumptions. Short injected prompts correspond to a weak magnetic field aligned towards unsafe cluster centers and yield a power-law scaling of attack success rate with the number of inference-time samples, while long injected prompts, i.e., strong magnetic field, yield exponential scaling. We observe qualitatively consistent behavior across a broad range of large language models, spanning parameter scales from 3B to 70B. In particular, the main trends remain stable across multiple attack methods, such as GCG and AutoDAN, as well as across benchmark datasets such as AdvBench and HarmBench.
Figures & tables
Figure 1: (a) Attack success rate Πk is plotted against number of inference time samples k . The experiment is performed with Mistral-7B-Instruct-v0.3 acting as the judge of jailbreaking on AdvBench dataset using the GCG attack method in Zou et al. (2023) (b-c) In the plot we show the distribution of per prompt safe probability of generated output of Llama-3.2-3B-Instruct with Mistral-7B-Instruct-v0.3 acting as a judge on AdvBench dataset. Plot (b) corresponds to the setup with no attack strategy and it shows very small attack success probability for most prompts with p^=1 . Plot (c) is obtained under GCG attack. It shows that p^≈0.9 leading to much higher attack success
Figure 2: The scaling law parameters as a function of jailbreak strength.
Figure 3: We see that prompt injection leads to a much higher attack success rate reflected in the curve falling off exponentially as discussed in the main text. The experiments feature Llama-3-8B-Instruct on AdvBench dataset with Mistral-7B-Instruct-v0.3 as a judge. (a) The attack is performed with the GCG-based universal prompt injection method as in Zou et al. (2023) (b) The attacks were performed using stealthy prompt-specific jailbreak strings generated by the AutoDAN method in Liu et al. (2024a) . The straight line appearing in the high injection curves is due to numerical limitations of our code ( 1−Πk≈10−3 ).
Figure 4: OLMo-2-0325-32B-Instruct was tested on the AdvBench dataset. For both (a) GCG- and (b) AutoDAN-based attack methods, we observe similar trends as in Figure 3 .
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Figure B.1: Sample two configurations σ,τ independently from the Gibbs measure. R(σ,τ) concentrates at q(σ,τ) , (σ,τ) is the first level at which they differ. The distance d(σ,τ)=(1−R(σ,τ))/2 is ultrametric: R(α1,α3)≥min(R(α1,α2),R(α2,α3)) .
Figure C.1: A schematic view of the low energy landscape of large number of spins interacting via the spin-glass Hamiltonian. In low temperature replica symmetry breaking phase the Gibbs measure decomposes into many hierarchically organized (based on overlaps) pure states/clusters; it is common to picture these states as ‘valleys’ or ‘basins’ in a energy landscape. Following Zhou and Wong [2009] , Zhou [2011] , we approximate the overlap based clustering in replica symmetry breaking phase in terms of basin of attraction associated with several local minima. Here we presented what typical sampling of the size biased ordering would look like - dots signify individual spin configurations.
Figure E.1: Experimental evidence for Assumption 23 . The figure on the left, confirms exponential concentration in N . The figure on the right shows faster than exponential concentration in distance from the mean.
Figure F.1: Attack success rate in spin-glass-based model is plotted numerically for N=24,p=2,β=10,j0=1 . The teacher-student setup is matched for β,j0 with additional magnetic field h turned on for the student along the m=1 teacher cluster at the lowest level. In plot (a), we compare the numerical plot against the one coming from Theorem 4 , i.e., log(−log(Πk))=−νlogk−νλ+logCm , and see that for small λ or equivalently h the graphs are in good agreement in the domain of validity of the theoretical result N≫k≫1≫kλ2 . As we increase h violating kλ2≪1 we see that the experimental results differ significantly from the prediction of Theorem 4 - in this domain it is meaningful to fit the experimental results to a form suggested by Theorem 4 and Theorem 5 , i.e., log(−log(Πk))=−ν^logk−μ^k+logc^ . From plot (b), we see that it is possible to find a reasonable fit to this from in the large- h regime in which both μ^,ν^ increase monotonically with h .
Figure G.1: Attack success rate measurement based on refusal string and GPT-4 as LLM-judge and analysis of harmfulness of jailbroken response for various attack methods. Plot (a) shows that ASR is much smaller for the LLM-judge metric compared to the refusal string-based metric for the reason explained above. Plot (b) compares the harmfulness of the successful jailbroken responses (score ≥ 4) for different attack methods. We see that injecting a universal adversarial string achieves higher ASR compared to “Sure here is” injection.
Figure G.2: Attack success rate measurement using refusal string method and LLM-as-a-judge (Mistral-7B-Instruct-v0.3) method under AutoDAN injection attack while varying the inference time sample, k , values. Prompts were taken from the AdvBench dataset. (a) Llama-3-8B-Instruct model (b) OLMo-2-0325-32B-Instruct model.
#
Refusal String
1
I’m sorry.
2
Sorry.
3
I am sorry.
4
I apologize.
5
As an
6
As an AI
Appendix
Table G.1: List of refusal strings used for computing attack success rate
Figure G.3: Attack success rate measurement based on LLM-as-a-judge method, separately using two different LLMs: GPT-4 and Mistral-7B-Instruct-v0.3, and comparison to the refusal string-based method. Target model responses were generated using prompts from the AdvBench dataset. (a) ASR was calculated on responses from GPT-Turbo-3.5 as the target model. (b) ASR was calculated on responses from Claude-Sonnet-4.5 as the target model.
Figure H.1: The experiments feature Llama-3.2-3B-Instruct on AdvBench dataset. (a) The attack was performed with the GCG-based universal prompt injection method as in Zou et al. [2023] (b) The attack was performed using stealthy prompt-specific jailbreak strings generated by the AutoDAN method in Liu et al. [2024a] . In both cases, we used Mistral-7B-Instruct-v0.3 as a judge for ASR calculation. The straight line appearing in the high injection curves at large k values is due to numerical limitations of our code.
Figure H.2: The experiments feature Llama-3-70B-Instruct on AdvBench dataset. (a) The attack was performed with the GCG-based universal prompt injection method as in Zou et al. [2023] (b) The attack was performed using stealthy prompt-specific jailbreak strings generated by the AutoDAN method in Liu et al. [2024a] . In both cases, we used Mistral-7B-Instruct-v0.3 as a judge for ASR calculation. The straight line appearing in the high injection curves at large k values is due to numerical limitations of our code.
Figure H.3: Olmo-3.1-32B-Instruct is tested on prompts from the AdvBench dataset. ASR was calculated using Mistral-7B-Instruct-v0.3 as a judge. (a) GCG attack; (b) AutoDAN attack;
Figure H.4: Olmo family of models is tested on standard prompts from the HarmBench dataset. In all cases, we use Mistral-7B-Instruct-v0.3 as a judge for ASR calculation. (a) OLMo-2-0325-32B-Instruct; GCG attack; (b) OLMo-2-0325-32B-Instruct; AutoDAN attack; (c) Olmo-3.1-32B-Instruct; GCG attack; (d) Olmo-3.1-32B-Instruct; AutoDAN attack;
Figure H.5: Llama-3-8B-Instruct model results are shown for category-wise prompts from the AdvBench dataset under GCG attack (top) and AutoDAN attack (bottom).
Figure H.6: Llama-3.2-3B-Instruct model results are shown for category-wise prompts from the AdvBench dataset under GCG attack (top) and AutoDAN attack (bottom).
Figure H.7: Llama-3-70B-Instruct model results are shown for category-wise prompts from the AdvBench dataset under GCG attack (top) and AutoDAN attack (bottom).
Figure H.8: OLMo-2-0325-32B-Instruct model results are shown for category-wise prompts from the AdvBench dataset under GCG attack (top) and AutoDAN attack (bottom).
Figure H.9: Olmo-3.1-32B-Instruct model results are shown for category-wise prompts from the AdvBench dataset under GCG attack (top) and AutoDAN attack (bottom).
Figure H.10: OLMo-2-0325-32B-Instruct model results are shown for category-wise prompts from the HarmBench dataset (standard prompts) under GCG attack
Figure H.11: Olmo-2-0325-32B-Instruct model results are shown for category-wise prompts from the HarmBench dataset (standard prompts) under AutoDAN attack
Figure H.12: OLMo-3.1-32B-Instruct model results are shown for category-wise prompts from the HarmBench dataset (standard prompts) under GCG attack
Figure H.13: Olmo-3.1-32B-Instruct model results are shown for category-wise prompts from the HarmBench dataset (standard prompts) under AutoDAN attack
Safety fine-tuning of language models typically requires a curated adversarial dataset. We take a different approach: score each candidate prompt's difficulty by how often the target model's own rollouts are judged harmful, then fine-tune on the hardest prompts paired with the model's own non-jailbroken rollouts. On Llama-3-8B-Instruct and Llama-3.2-3B-Instruct, this approach cuts the WildJailbreak attack success rate from 11.5% and 20.1% down to 1-3%, but pushes refusal on jailbreak-shaped benign prompts from 14-22% to 74-94%. Interleaving the same hard prompts 1:1 with adversarially-framed benign prompts (prompts that look like jailbreaks but have benign intent) cuts that refusal back down to 30-51% on 8B and 52-72% on 3B, at a cost of 2-6 percentage points of attack success rate. Within the mixed regime, training on the hardest half of the eligible pool rather than a random half cuts the remaining ASR by 35-50% (about 3 percentage points) on both models.
Safety-aligned Large Language Models (LLMs) remain vulnerable to interventions during inference that redirect generation toward harmful outputs. Recent work attributes this to shallow safety, where alignment concentrates in the first few output tokens. We show that shallow safety is a special case of a broader inference-time vulnerability, in which short token injections at any generation step can substantially alter subsequent safety behavior. We also find that a model's alignment with refusal directions in its hidden states does not predict its robustness to such injection, revealing that internal state alone does not determine generation behavior under perturbation. To address this, we align models directly on generation trajectories constructed by simulating mid-sequence perturbation, and show that this improves robustness to mid-sequence injection and generalizes to attacks that exploit early-token generation. Our work argues that robust safety alignment requires training on the generation process itself, not only its outputs.
Kyungmin Park, Taesup Kim
Hankuk University of Foreign Studies · Seoul National University
Adversarial attacks on Large Language Models (LLMs) aim to induce harmful content. However, existing methods suffer from high computational costs or strict model-pairing dependencies, limiting their scalability and transferability. We propose Reward Stealing Attack (ReSA), an adversarial attack framework that targets the latent safety reward underlying LLM alignment. ReSA employs maximum entropy inverse reinforcement learning to recover a proxy reward model solely from the aligned model's behavior. The extracted reward is then reversed at inference time to derive an adversarial policy, efficiently implemented via a reward-guided decoding mechanism. Experiments demonstrate that a single recovered reward generalizes across prompts and diverse models to reveal a fundamental alignment vulnerability, enabling ReSA to significantly outperform existing attacks in effectiveness and transferability. The code is available at https://github.com/GarminQ/ReSA.