Large Language Models (LLMs) are vulnerable to attacks like prompt injection, backdoor attacks, and adversarial attacks, which manipulate prompts or models to generate harmful outputs. In this paper, departing from traditional deep learning attack paradigms, we explore their intrinsic relationship and collectively term them Prompt Trigger Attacks (PTA). This raises a key question: Given a prompt, can we tell whether a hidden trigger is steering the model's behavior? We propose UniGuardian, to the best of our knowledge the first training-free LLM detector to jointly detect successfully activated prompt injection, backdoor, and adversarial attacks without knowing the attack type. Its shared mechanism measures how structured prompt perturbations shift the model's output distribution. Additionally, we introduce a single-forward strategy to optimize the detection pipeline, enabling simultaneous attack detection and text generation within a shared batched forward pass at each decoding step. Our experiments confirm that UniGuardian accurately and efficiently identifies trigger-activated prompts in LLMs.
Figures & tables
Figure 1: Overview of three types of attack on LLMs: (a) Prompt Injection manipulate prompts to inject specific outputs. (b) Backdoor Attacks embeds backdoor in the model and activated when a prompt contains triggers. (c) Adversarial Attacks introduce perturbations in the input text to manipulate the model to mislead LLMs.
Figure 2: Overview of UniGuardian. (a) Given a prompt, the LLM generates a base output generation. (b) A random masking strategy creates prompt variations by masking different word subsets. The LLM processes these masked prompts, measuring differences between the logits Li and Lb . (c) The single-forward strategy is introduced to accelerate trigger detection, allowing triggers to be identified simultaneously with text generation.
Model
Method
Prompt Injections
Jailbreak
SST2
Open Question
SMS Spam
auROC
auPRC
auROC
auPRC
auROC
auPRC
auROC
auPRC
auROC
auPRC
-
Prompt-Guard-86M
0.5732
0.5567
0.5000
0.5305
0.5000
0.4997
0.5000
0.5000
0.5538
0.5284
PPL Detection
0.3336
0.4193
0.1932
0.3676
0.2342
0.3531
0.2822
0.3679
0.2051
0.3784
3B
Llama-Guard-3-1B
0.5839
0.5651
0.5628
0.5652
0.4987
0.4991
0.4727
0.4870
0.4803
0.4905
Llama-Guard-3-8B
0.5000
0.5172
0.5530
0.5751
0.5132
0.5101
0.5015
0.5010
0.5000
0.5000
Granite-Guardian-3.1-8B
0.6339
0.7302
0.7382
0.7820
0.5978
0.5531
0.4216
0.4365
0.6322
0.5681
Table 1: Detection performance on prompt injection. RAP encounters OOM on 32B and 70B models and is omitted.
Model
Method
SST2
Open Question
SMS Spam
auROC
auPRC
auROC
auPRC
auROC
auPRC
-
Prompt-Guard-86M
0.5000
0.4997
0.5000
0.5000
0.5000
0.4910
PPL Detection
0.6043
0.6136
0.7138
0.7096
0.5818
0.5823
8B (LoRA)
Llama-Guard-3-1B
0.4866
0.4932
0.4985
0.4992
0.5294
0.5065
Llama-Guard-3-8B
0.4978
0.4997
0.4955
0.5000
0.4877
0.4910
Granite-Guardian-3.1-8B
0.1290
0.3281
0.1853
0.3420
0.2049
0.3395
Table 2: Comparison of detection performance on backdoor attacks (Trigger: cf ).
Model
Method
SST2
Open Question
SMS Spam
auROC
auPRC
auROC
auPRC
auROC
auPRC
-
Prompt-Guard-86M
0.5000
0.4997
0.5000
0.5000
0.5000
0.4910
PPL Detection
0.3228
0.3807
0.4081
0.4209
0.2866
0.3608
8B (LoRA)
Llama-Guard-3-1B
0.5162
0.5081
0.4742
0.4877
0.5216
0.5022
Llama-Guard-3-8B
0.4967
0.4997
0.4970
0.5000
0.4930
0.4910
Granite-Guardian-3.1-8B
0.2135
0.3484
0.1342
0.3310
0.2850
0.3618
Table 3: Comparison of detection performance on backdoor attacks (Trigger: I watched 3D movies ).
Model
Method
SST2
Emotion
auROC
auPRC
auROC
auPRC
-
Prompt-Guard-86M
0.5024
0.5012
0.5000
0.5000
PPL Detection
0.6266
0.6177
0.5348
0.5330
32B
Llama-Guard-3-1B
0.4844
0.4924
0.5082
0.5041
Llama-Guard-3-8B
0.5000
0.5000
0.4984
0.5000
Granite-Guardian-3.1-8B
0.7840
0.7739
0.7303
0.6428
Table 4: Detection performance on adversarial attacks. RAP encounters OOM and is omitted.
Model
Method
auROC
auPRC
8B
ONION
0.9340
0.9278
RAP
0.8493
0.8060
UniGuardian
0.9338
0.9355
70B
ONION
0.8322
0.8369
RAP
OOM
OOM
UniGuardian
0.9592
0.9530
Table 5: Detection on the GSM8K controlled prompt-injection setting.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
# Test
Prompt Injections
116
Jailbreak
262
SST2
1821
Open Question
660
SMS Spam
558
Emotion
612 13 13 13 Only joy and sadness classes.
Appendix
Table 6: Number of test samples of datasets.
Figure 4: Distribution of suspicion scores for poisoned and clean input on prompt injection (32B).
Figure 5: Distribution of suspicion scores for poisoned and clean input on backdoor attacks.
Figure 6: Distribution of suspicion scores for poisoned and clean input on adversarial attacks (70B).
Figure 7: Impact of n on detection performance. The x-axis represents n , which is defined as n=x× (length of prompt).
Figure 8: Impact of m on detection performance. The x-axis represents m , which is defined as m= (length of prompt) x .
Figure 9: Threshold selection for prompt injection detection (Jailbreak Dataset on 32B and 70B models).
Figure 10: Threshold selection for backdoor attack detection (SST2 Dataset on two types of backdoor attacks).
Figure 11: Threshold selection for adversarial attack detection (SST2 Dataset on 32B and 70B models).
Method
auROC
auPRC
PPL-based Detection
0.4746
0.4699
Llama-Guard-3-1B
0.5064
0.5032
Llama-Guard-3-8B
0.5457
0.5397
Granite-Guardian-3.1-8B
0.5034
0.6035
LLM-based Detection
0.8117
0.7283
UniGuardian (Ours)
0.6692
0.6575
Appendix
Table 7: Detection performance under adaptive adversarial rephrasing attacks generated by AdvPrompter on Llama 3.1 8B.
Table 9: Detection performance on HarmBench HumanJailBreaks.
Prompt Injection ( L=118 )
Jailbreak ( L=1235 )
k
m
n=0.2L
n=0.5L
n=L
n=2L
n=0.2L
n=0.5L
n=L
n=2L
1
L0.1
0.1778
0.3948
0.6337
0.8658
0.3299
0.6321
0.8649
0.9817
L0.3
0.5476
0.8693
0.9829
0.9997
0.7992
0.9819
0.9997
1.0000
L0.5
0.8695
0.9946
1.0000
1.0000
0.9992
1.0000
1.0000
1.0000
L0.7
0.9980
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
2
L0.1
0.3240
0.6337
0.8658
0.9820
0.5510
0.8647
0.9817
0.9997
Appendix
Table 10: Effect of hyperparameter scaling on trigger coverage. We show the probability P=1−(1−Lm)nk of masking at least one of the k trigger tokens when generating n masked variants of a prompt of length L , with m tokens masked per variant. The results illustrate that choosing n proportional to L and a sublinear mask size m=L0.3 achieves near-saturated trigger coverage.
Model Size
Method
Prompt Injections
Jailbreak Classification
-
Prompt-Guard-86M (Local)
1.94
0.26
PPL Detection
7.17
3.28
32B
Llama-Guard-3-1B (Local)
1.98
0.34
Llama-Guard-3-8B (Local)
3.86
0.83
Granite-Guardian-3.1-8B (Local)
4.46
1.16
LLM-based detection (Local)
7.48
3.49
Appendix
Table 11: Per-sample runtime (seconds per sample) of various detection methods for prompt injection detection on the Prompt Injection and Jailbreak datasets.
East China Normal University Shanghai, China · East China Normal University, Shanghai Innovation Institute Shanghai, China · University of Southampton Southampton, England +2