Large language models (LLMs) perform remarkably well on complex tasks, yet remain highly vulnerable to prompt injection attacks, where malicious instructions embedded in external data can override user intent. Existing defenses remain limited by model fine-tuning requirements, vulnerability to adaptive attacks, or reliance on brittle handcrafted prompts. We argue that a fundamental source of this vulnerability is the lack of an explicit representation of trust provenance. To address this, we introduce Learnable Trust-Boundary Delimiters (LTBD), a lightweight defense that explicitly encodes trust boundaries in the input while keeping the LLM parameters unchanged. LTBD uses a small number of learnable delimiters to distinguish trusted user instructions from untrusted external data, enabling the model to better respect the intended trust hierarchy. Experimental results show that LTBD substantially outperforms inference-time defenses and performs competitively with training-based approaches, while preserving benign-task utility and introducing negligible inference overhead. In particular, LTBD achieves 0.00% ASR on AlpacaFarm and only 0.11-0.19% ASR on TaskTracker. LTBD also remains effective under adaptive attacks, where adversaries have full knowledge of the defense and explicitly attempt to bypass it.
Figures & tables
Figure 1: Overview of Learnable Trust-Boundary Delimiters (LTBD) . (a) Prompt injections embedded in external data steer the LLM away from the original task. (b) LTBD insert four learnable delimiters to explicitly mark the trusted instruction region and the untrusted data region. The delimiter delimiter embeddings are the only trainable parameters, while the base LLM fθ remains frozen. During inference the input is wrapped with the learned delimiters, enabling the model to follow the trusted instruction and treat the (potentially attacked) external data as data rather than a new instruction.
Model
Defense Type
Method
AlpacaFarm [ 15 ]
SEP [ 18 ]
TaskTracker [ 17 ]
CyberSecEval2 [ 19 ]
WinRate {\color[rgb]{1,0,0}\uparrow}
ASR {\color[rgb]{0,0.5469,0.2695}\downarrow}
WinRate {\color[rgb]{1,0,0}\uparrow}
ASR {\color[rgb]{0,0.5469,0.2695}\downarrow}
ASR {\color[rgb]{0,0.5469,0.2695}\downarrow}
ASR {\color[rgb]{0,0.5469,0.2695}\downarrow}
Llama3-8B- Instruct
No Defense
None
33.82
55.29
59.72
74.47
16.4
45.45
Prompting-based
TextGrad [ 20 ]
22.9
0.00
1.1
3.5
0.25
1.8
Reminder [ 13 ]
24.4
34.6
48.3
75.2
19.8
43.6
Sandwich [ 14 ]
26.8
56.7
46.9
63.4
5.5
41.8
DefensiveToken [ 16 ]
33.46
0.48
58.60
3.20
0.27
3.64
Table 1: Comparison of utility (WinRate % {\color[rgb]{1,0,0}\uparrow} ) and security (ASR % {\color[rgb]{0,0.5469,0.2695}\downarrow} ) across prompt injection defenses.
Model
Method
AlpacaFarm [ 15 ]
SEP [ 18 ]
WinRate {\color[rgb]{1,0,0}\uparrow}
ASR {\color[rgb]{0,0.5469,0.2695}\downarrow}
ASR {\color[rgb]{0,0.5469,0.2695}\downarrow}
Llama3-8B- Instruct
Prefix + CE
27.60
6.73
15.43
Boundary + CE
34.51
0.00
2.34
Boundary + CE + Pref (ours)
35.73
0.00
2.04
Llama3.1-8B- Instruct
Prefix + CE
27.50
23.08
33.81
Boundary + CE
30.98
0.48
3.43
Table 2: Ablation study on trust boundary placement and loss type.
Model
Plain Injection
Adaptive Attack
Llama3-8B-Instruct
0.00
0.00
Llama3.1-8B-Instruct
0.48
0.48
Falcon3-7B-Instruct
0.00
0.00
Qwen2.5-7B-Instruct
0.00
0.48
Table 3: ASR (% {\color[rgb]{0,0.5469,0.2695}\downarrow} ) under Adaptive Attacks.
Prompt injection is a critical security threat in large language model (LLM) applications, where attackers hijack model behavior by embedding malicious instructions in user or external data. Existing detection methods only detect the presence of injection and refuse to respond upon detection, overlooking the fact that for many modern aligned models, well-crafted instructions can resist most injection attacks. This means that the injection robustness varies significantly across instructions and models. This leads to widespread unnecessary over-refusal: inputs containing injections that the model could have handled correctly are rejected incorrectly. To deal with this over-refusal issue, we propose BASIS (Robustness-Aware Prompt Injection Defense). This defense method uses the Attention Competition Ratio (ρ) as features to train two sparse linear probes: an existence probe and a breach probe. Both probes make defense decisions through cascaded gating, which does not require additional LLM inference. BASIS comprises three stages: injection existence detection, per-sample breach prediction, and instruction robustness assessment; the online cascade refuses only when the model would actually be compromised and thus avoids over-refusal on robust instructions. Experiments across four tasks and six open-source LLMs show that BASIS maintains near-perfect injection detection while substantially reducing over-refusal on safe attack samples, especially under robust instruction templates.
Laiqiao Qin, Tianqing Zhu, Longxiang Gao +1
City University of Macau, Macau · Qilu University of Technology
We identify a security-fidelity tradeoff in defending LLMs against indirect prompt injection: defenses resist injected instructions largely by suppressing untrusted text, which corrupts tasks that must preserve it, such as translation and document editing. Attack-success metrics cannot see this, because a model that ignores an injection and one that faithfully processes it as data score identically. We introduce SecFid, a benchmark built so that executing an injection, processing it as data, and ignoring it produce distinguishable outputs. This makes fidelity measurable and exposes a frontier: across 1,168 examples and 48 configurations, no model or defense achieves both objectives. The highest-fidelity model reaches 96.5% fidelity at 47.8% security, while the most secure defenses invert this, at 99.3% security but only 71.0%-73.9% fidelity. Even defenses with identical security differ in how they earn it: some repair hijacks into faithful processing, others simply suppress benign content. A decision-theoretic analysis shows why no fixed choice can be right everywhere: the correct behavior is not a property of the defense but of the deployment, set by its relative cost of a hijack versus a dropped span. Security alone therefore measures only half of robustness, and reporting it without fidelity hides the price at which it was bought.
Large Language Models (LLMs) are vulnerable to attacks like prompt injection, backdoor attacks, and adversarial attacks, which manipulate prompts or models to generate harmful outputs. In this paper, departing from traditional deep learning attack paradigms, we explore their intrinsic relationship and collectively term them Prompt Trigger Attacks (PTA). This raises a key question: Given a prompt, can we tell whether a hidden trigger is steering the model's behavior? We propose UniGuardian, to the best of our knowledge the first training-free LLM detector to jointly detect successfully activated prompt injection, backdoor, and adversarial attacks without knowing the attack type. Its shared mechanism measures how structured prompt perturbations shift the model's output distribution. Additionally, we introduce a single-forward strategy to optimize the detection pipeline, enabling simultaneous attack detection and text generation within a shared batched forward pass at each decoding step. Our experiments confirm that UniGuardian accurately and efficiently identifies trigger-activated prompts in LLMs.
Huawei Lin, Yingjie Lao, Tony Geng +2
Rochester Institute of Technology · Tufts University · University of Rochester +1