cs.CROct 8, 2026

LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense

Authors: Luman Zhao, Minghui Xu, Yue Zhang, Yijun Yang

Organizations: Shandong University, Shandong Academy of Artificial Intelligence

Abstract

Large language models (LLMs) perform remarkably well on complex tasks, yet remain highly vulnerable to prompt injection attacks, where malicious instructions embedded in external data can override user intent. Existing defenses remain limited by model fine-tuning requirements, vulnerability to adaptive attacks, or reliance on brittle handcrafted prompts. We argue that a fundamental source of this vulnerability is the lack of an explicit representation of trust provenance. To address this, we introduce Learnable Trust-Boundary Delimiters (LTBD), a lightweight defense that explicitly encodes trust boundaries in the input while keeping the LLM parameters unchanged. LTBD uses a small number of learnable delimiters to distinguish trusted user instructions from untrusted external data, enabling the model to better respect the intended trust hierarchy. Experimental results show that LTBD substantially outperforms inference-time defenses and performs competitively with training-based approaches, while preserving benign-task utility and introducing negligible inference overhead. In particular, LTBD achieves 0.00% ASR on AlpacaFarm and only 0.11-0.19% ASR on TaskTracker. LTBD also remains effective under adaptive attacks, where adversaries have full knowledge of the defense and explicitly attempt to bypass it.

Figures & tables

Explore similar work

CardsList
  1. BASIS: Breach-Aware Selective Prompt Injection Shielding with Prefill Attention Probes

    Aug 8, 2026Laiqiao Qin, Tianqing Zhu, Longxiang Gao +1Prompt Injection DefensePrompt Injection Attacks on LLMs

  2. Security--Fidelity Tradeoffs: The Hidden Cost of Prompt Injection Defense

    Jun 29, 2026Mitchell Hermon, Rahul Gupta, Weitong Ruan +2LLM Safety BenchmarksPrompt Injection Defense

  3. UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models

    Feb 18, 2025Huawei Lin, Yingjie Lao, Tony Geng +2LLM Backdoor AttacksAdversarial Attacks on LLMs