cs.CLFeb 18, 2025

UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models

Authors: Huawei Lin, Yingjie Lao, Tony Geng, Tan Yu, Weijie Zhao

Organizations: Rochester Institute of Technology · Tufts University · University of Rochester · NVIDIA

Abstract

Large Language Models (LLMs) are vulnerable to attacks like prompt injection, backdoor attacks, and adversarial attacks, which manipulate prompts or models to generate harmful outputs. In this paper, departing from traditional deep learning attack paradigms, we explore their intrinsic relationship and collectively term them Prompt Trigger Attacks (PTA). This raises a key question: Given a prompt, can we tell whether a hidden trigger is steering the model's behavior? We propose UniGuardian, to the best of our knowledge the first training-free LLM detector to jointly detect successfully activated prompt injection, backdoor, and adversarial attacks without knowing the attack type. Its shared mechanism measures how structured prompt perturbations shift the model's output distribution. Additionally, we introduce a single-forward strategy to optimize the detection pipeline, enabling simultaneous attack detection and text generation within a shared batched forward pass at each decoding step. Our experiments confirm that UniGuardian accurately and efficiently identifies trigger-activated prompts in LLMs.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Automated jailbreak attack targeting multiple defense strategies

    Jun 15, 2026Qi Wang, Chengcheng Wan, Weijia He +4Attacker Large Language ModelJailbreak Attacks

  2. The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models

    Oct 6, 2026Yibo Zhang, Tianrong Guan, Liang Lin +3

  3. BASIS: Breach-Aware Selective Prompt Injection Shielding with Prefill Attention Probes

    Aug 8, 2026Laiqiao Qin, Tianqing Zhu, Longxiang Gao +1Prompt-Injection DetectorsPrompt Engineering