Large Language Models (LLMs) are vulnerable to attacks like prompt injection, backdoor attacks, and adversarial attacks, which manipulate prompts or models to generate harmful outputs. In this paper, departing from traditional deep learning attack paradigms, we explore their intrinsic relationship and collectively term them Prompt Trigger Attacks (PTA). This raises a key question: Given a prompt, can we tell whether a hidden trigger is steering the model's behavior? We propose UniGuardian, to the best of our knowledge the first training-free LLM detector to jointly detect successfully activated prompt injection, backdoor, and adversarial attacks without knowing the attack type. Its shared mechanism measures how structured prompt perturbations shift the model's output distribution. Additionally, we introduce a single-forward strategy to optimize the detection pipeline, enabling simultaneous attack detection and text generation within a shared batched forward pass at each decoding step. Our experiments confirm that UniGuardian accurately and efficiently identifies trigger-activated prompts in LLMs.
Figures & tables
Figure 1: Overview of three types of attack on LLMs: (a) Prompt Injection manipulate prompts to inject specific outputs. (b) Backdoor Attacks embeds backdoor in the model and activated when a prompt contains triggers. (c) Adversarial Attacks introduce perturbations in the input text to manipulate the model to mislead LLMs.
Figure 2: Overview of UniGuardian. (a) Given a prompt, the LLM generates a base output generation. (b) A random masking strategy creates prompt variations by masking different word subsets. The LLM processes these masked prompts, measuring differences between the logits Li and Lb . (c) The single-forward strategy is introduced to accelerate trigger detection, allowing triggers to be identified simultaneously with text generation.
Model
Method
Prompt Injections
Jailbreak
SST2
Open Question
SMS Spam
auROC
auPRC
auROC
auPRC
auROC
auPRC
auROC
auPRC
auROC
auPRC
-
Prompt-Guard-86M
0.5732
0.5567
0.5000
0.5305
0.5000
0.4997
0.5000
0.5000
0.5538
0.5284
PPL Detection
0.3336
0.4193
0.1932
0.3676
0.2342
0.3531
0.2822
0.3679
0.2051
0.3784
3B
Llama-Guard-3-1B
0.5839
0.5651
0.5628
0.5652
0.4987
0.4991
0.4727
0.4870
0.4803
0.4905
Llama-Guard-3-8B
0.5000
0.5172
0.5530
0.5751
0.5132
0.5101
0.5015
0.5010
0.5000
0.5000
Granite-Guardian-3.1-8B
0.6339
0.7302
0.7382
0.7820
0.5978
0.5531
0.4216
0.4365
0.6322
0.5681
Table 1: Detection performance on prompt injection. RAP encounters OOM on 32B and 70B models and is omitted.
Model
Method
SST2
Open Question
SMS Spam
auROC
auPRC
auROC
auPRC
auROC
auPRC
-
Prompt-Guard-86M
0.5000
0.4997
0.5000
0.5000
0.5000
0.4910
PPL Detection
0.6043
0.6136
0.7138
0.7096
0.5818
0.5823
8B (LoRA)
Llama-Guard-3-1B
0.4866
0.4932
0.4985
0.4992
0.5294
0.5065
Llama-Guard-3-8B
0.4978
0.4997
0.4955
0.5000
0.4877
0.4910
Granite-Guardian-3.1-8B
0.1290
0.3281
0.1853
0.3420
0.2049
0.3395
Table 2: Comparison of detection performance on backdoor attacks (Trigger: cf ).
Model
Method
SST2
Open Question
SMS Spam
auROC
auPRC
auROC
auPRC
auROC
auPRC
-
Prompt-Guard-86M
0.5000
0.4997
0.5000
0.5000
0.5000
0.4910
PPL Detection
0.3228
0.3807
0.4081
0.4209
0.2866
0.3608
8B (LoRA)
Llama-Guard-3-1B
0.5162
0.5081
0.4742
0.4877
0.5216
0.5022
Llama-Guard-3-8B
0.4967
0.4997
0.4970
0.5000
0.4930
0.4910
Granite-Guardian-3.1-8B
0.2135
0.3484
0.1342
0.3310
0.2850
0.3618
Table 3: Comparison of detection performance on backdoor attacks (Trigger: I watched 3D movies ).
Model
Method
SST2
Emotion
auROC
auPRC
auROC
auPRC
-
Prompt-Guard-86M
0.5024
0.5012
0.5000
0.5000
PPL Detection
0.6266
0.6177
0.5348
0.5330
32B
Llama-Guard-3-1B
0.4844
0.4924
0.5082
0.5041
Llama-Guard-3-8B
0.5000
0.5000
0.4984
0.5000
Granite-Guardian-3.1-8B
0.7840
0.7739
0.7303
0.6428
Table 4: Detection performance on adversarial attacks. RAP encounters OOM and is omitted.
Model
Method
auROC
auPRC
8B
ONION
0.9340
0.9278
RAP
0.8493
0.8060
UniGuardian
0.9338
0.9355
70B
ONION
0.8322
0.8369
RAP
OOM
OOM
UniGuardian
0.9592
0.9530
Table 5: Detection on the GSM8K controlled prompt-injection setting.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
# Test
Prompt Injections
116
Jailbreak
262
SST2
1821
Open Question
660
SMS Spam
558
Emotion
612 13 13 13 Only joy and sadness classes.
Appendix
Table 6: Number of test samples of datasets.
Figure 4: Distribution of suspicion scores for poisoned and clean input on prompt injection (32B).
Figure 5: Distribution of suspicion scores for poisoned and clean input on backdoor attacks.
Figure 6: Distribution of suspicion scores for poisoned and clean input on adversarial attacks (70B).
Figure 7: Impact of n on detection performance. The x-axis represents n , which is defined as n=x× (length of prompt).
Figure 8: Impact of m on detection performance. The x-axis represents m , which is defined as m= (length of prompt) x .
Figure 9: Threshold selection for prompt injection detection (Jailbreak Dataset on 32B and 70B models).
Figure 10: Threshold selection for backdoor attack detection (SST2 Dataset on two types of backdoor attacks).
Figure 11: Threshold selection for adversarial attack detection (SST2 Dataset on 32B and 70B models).
Method
auROC
auPRC
PPL-based Detection
0.4746
0.4699
Llama-Guard-3-1B
0.5064
0.5032
Llama-Guard-3-8B
0.5457
0.5397
Granite-Guardian-3.1-8B
0.5034
0.6035
LLM-based Detection
0.8117
0.7283
UniGuardian (Ours)
0.6692
0.6575
Appendix
Table 7: Detection performance under adaptive adversarial rephrasing attacks generated by AdvPrompter on Llama 3.1 8B.
Table 9: Detection performance on HarmBench HumanJailBreaks.
Prompt Injection ( L=118 )
Jailbreak ( L=1235 )
k
m
n=0.2L
n=0.5L
n=L
n=2L
n=0.2L
n=0.5L
n=L
n=2L
1
L0.1
0.1778
0.3948
0.6337
0.8658
0.3299
0.6321
0.8649
0.9817
L0.3
0.5476
0.8693
0.9829
0.9997
0.7992
0.9819
0.9997
1.0000
L0.5
0.8695
0.9946
1.0000
1.0000
0.9992
1.0000
1.0000
1.0000
L0.7
0.9980
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
2
L0.1
0.3240
0.6337
0.8658
0.9820
0.5510
0.8647
0.9817
0.9997
Appendix
Table 10: Effect of hyperparameter scaling on trigger coverage. We show the probability P=1−(1−Lm)nk of masking at least one of the k trigger tokens when generating n masked variants of a prompt of length L , with m tokens masked per variant. The results illustrate that choosing n proportional to L and a sublinear mask size m=L0.3 achieves near-saturated trigger coverage.
Model Size
Method
Prompt Injections
Jailbreak Classification
-
Prompt-Guard-86M (Local)
1.94
0.26
PPL Detection
7.17
3.28
32B
Llama-Guard-3-1B (Local)
1.98
0.34
Llama-Guard-3-8B (Local)
3.86
0.83
Granite-Guardian-3.1-8B (Local)
4.46
1.16
LLM-based detection (Local)
7.48
3.49
Appendix
Table 11: Per-sample runtime (seconds per sample) of various detection methods for prompt injection detection on the Prompt Injection and Jailbreak datasets.
Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks. However, their safety remains a critical concern due to their susceptibility to adversarial prompt-based attacks. In this paper, we present UNIATTACK, an adversarial testing framework designed from a defense-oriented perspective to systematically construct effective black-box attack prompts. Unlike prior approaches that rely on static templates or iterative model-specific tuning, UNIATTACK extracts minimal but high-impact attack features from diverse existing attacks, optimizes them via a specialized attacker LLM, and composes them into flexible templates through automated refinement process. This feature-centric construction enables one-shot attacks that generalize across multiple models and safety categories, providing a practical tool for assessing LLM robustness. Our evaluation results shows that compared to the baselines, UNIATTACK achieves an average attack success rate (ASR) improvement of 64.63%-248.82% on models deployed with multi-layered defense mechanisms and it only takes 0.03%-4.96% cost of the baselines. UNIATTACK artifact is available at https://anonymous.4open.science/r/UniAttack-Artifact-30F1.
Qi Wang, Chengcheng Wan, Weijia He +4
East China Normal University Shanghai, China · East China Normal University, Shanghai Innovation Institute Shanghai, China · University of Southampton Southampton, England +2
Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks. Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space. We challenge this assumption with a novel answer-side backdoor for multi-turn dialogue. Instead of inserting the trigger into the input, the adversary uses a benign first-turn prompt to naturally induce the model to generate a specific, seemingly innocuous word. Once merged into the dialogue history, this self-generated word becomes the trigger. When a later harmful query arrives, the model detects its own trigger and bypasses its safety refusal, while the user input stays perfectly clean. Across four LLMs, our attack reaches near-perfect Attack Success Rates, approaching 100% at only a 5% poisoning rate, while preserving general utility and clean-input safety, and it evades mainstream input-centric defenses. Representation-level analysis shows that the self-generated trigger consistently suppresses the model's refusal signal, exposing a critical blind spot in current LLM defenses.
Yibo Zhang, Tianrong Guan, Liang Lin +3
Queen Mary University of London, UK · Squirrel AI Learning, USA
Prompt injection is a critical security threat in large language model (LLM) applications, where attackers hijack model behavior by embedding malicious instructions in user or external data. Existing detection methods only detect the presence of injection and refuse to respond upon detection, overlooking the fact that for many modern aligned models, well-crafted instructions can resist most injection attacks. This means that the injection robustness varies significantly across instructions and models. This leads to widespread unnecessary over-refusal: inputs containing injections that the model could have handled correctly are rejected incorrectly. To deal with this over-refusal issue, we propose BASIS (Robustness-Aware Prompt Injection Defense). This defense method uses the Attention Competition Ratio (ρ) as features to train two sparse linear probes: an existence probe and a breach probe. Both probes make defense decisions through cascaded gating, which does not require additional LLM inference. BASIS comprises three stages: injection existence detection, per-sample breach prediction, and instruction robustness assessment; the online cascade refuses only when the model would actually be compromised and thus avoids over-refusal on robust instructions. Experiments across four tasks and six open-source LLMs show that BASIS maintains near-perfect injection detection while substantially reducing over-refusal on safe attack samples, especially under robust instruction templates.
Laiqiao Qin, Tianqing Zhu, Longxiang Gao +1
City University of Macau, Macau · Qilu University of Technology