Prompt Injection Attacks on LLMs

LLM: Large Language Model

Latest papers 95

All topics
CardsList
  1. Secure Speculative Decoding for Large Language Models

    Oct 6, 2026Yichi Zhang, Zhiqi Wang, Neil Gong +1Adversarial Attacks on LLMsSpeculative Decoding

  2. RAG-PIBench: A Leakage-Aware Benchmark for Prompt-Injection Detection in Trustworthy RAG Systems

    Oct 6, 2026Niveen O. Jaffal, Ahmet Yuksel, David MohaisenRAG SecurityPrompt Injection Attacks on LLMs

  3. RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents

    Oct 5, 2026Mohamed Dhouib, Clement Elliker, Alexi Canesse +5Prompt Injection DefenseLLM Agents

  4. Can CaMeLs Talk? Securing Multi-Agent Systems Against Indirect Prompt Injection Attacks

    Oct 5, 2026James Peters-Gill, Avi Semler, Henning Bartsch +2AI Agent Security BenchmarksInformation Flow Control

  5. Readable Before Actionable: Causal Tracing of Indirect Prompt Injection

    Oct 4, 2026Zhe Yu, Wenpeng Xing, Xingxing Yang +1Causal Reasoning in Language ModelsIndirect Prompt Injection

  6. Blocking at the Boundary: Auditing Long-Horizon Agents against Staged Prompt Injection

    Oct 4, 2026Jingkai Liu, Yufei Han, Xiaoting Lyu +2Prompt InjectionAI Agent Auditing

  7. Divide and Inject: Can Agents Reconstruct an Indirect Prompt Injection from Fragments?

    Sep 29, 2026Michael Lee, Zhipeng Wei, Yue Dong +1Indirect Prompt InjectionPrompt Injection Attacks

  8. Render Before Reading: Visual Rendering as a Prompt Injection Defense

    Sep 28, 2026Jie Zhang, Andrei Baroian, Jan N. van Rijn +2Multimodal Large Language ModelsModality Gap in VLMs

  9. Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection

    Sep 28, 2026Yan Zhan, Yunze Song, Mengkai Hou +3Prompt InjectionPrompt Injection Defense

  10. CoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-based Agents

    Sep 28, 2026Xiao Yang, Yangchen Ou, Yuhan Gao +3Adversarial TrainingLLM Agent Security

  11. Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning

    Sep 27, 2026Chenlong Yin, Xiaolong Jin, Wei Zou +2Adversarial Attacks on LLMsLLM Red Teaming

  12. ENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation

    Sep 24, 2026Qingyu Wu, Zeyu Feng, Yongda Yu +2Adversarial Prompt GenerationPrompt Injection Attacks

  13. Prefilling the Reasoning Channel: Output-Prefix Attacks on Reasoning LLMs

    Sep 24, 2026Lukáš Brůna, Robert Bridges, Adam EkAdversarial Attacks on LLMsPrompt Injection Attacks

  14. ActGuard: Pre-execution Action Auditing against Indirect Prompt Injection in LLM Agents

    Sep 14, 2026Bingzheng Wang, Xiaoyan Gu, Wentao Wang +3Runtime Enforcement for AI AgentsPrompt Injection Defense

  15. DriftNet: A Dual-Head Trajectory Transformer for Detecting and Localizing Prompt Injection in LLM Agents

    Sep 12, 2026Asif Pinjari, Mithun Paul Saint-GermainAI Agent Security BenchmarksAI Agent Monitoring

  16. An Experimental Evaluation of Multimodal Prompt Injection Attacks on Agentic AI Frameworks

    Sep 10, 2026Viet K. Nguyen, Mohammad I. HusainAI Agent SecurityAI Agent Security Benchmarks

  17. CoER: Defending against Adaptive Indirect Prompt Injection via Adversarial Co-Evolution and Refinement

    Sep 7, 2026Boyang Zhang, Qingxin Xiao, Lingwei Dang +1Reinforcement LearningAI Agent Security

  18. Stealing Reasoning Traces from Proprietary LLM APIs

    Aug 10, 2026Alexander Panfilov, David Schmotz, Ilia Shumailov +5Model Extraction AttacksAdversarial Attacks

  19. Toward Metacognitive One-Shot Indirect Prompt Injection: Strategy Abstraction Via Outcome-Conditioned Reflection

    Aug 9, 2026Sihan Hou, Xinmeng Hou, Zhijun Zhang +5AI Agent Security BenchmarksIndirect Prompt Injection

  20. BASIS: Breach-Aware Selective Prompt Injection Shielding with Prefill Attention Probes

    Aug 8, 2026Laiqiao Qin, Tianqing Zhu, Longxiang Gao +1Prompt Injection DefensePrompt Injection Attacks on LLMs

  21. NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs

    Aug 7, 2026Aditya Katkar, Om Karkele, Kartik Mandhane +2Tool-Augmented Language Model AgentsLLM Guardrails

  22. StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection

    Aug 6, 2026Zhuoxin Zhan, Akbar Rafiey, Avery Ma +2Computer-Use AgentsAI Agent Safety

  23. Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots

    Aug 6, 2026S. M . Bhagya P. Samarakoon, M. A. Viraj J. Muthugala, W. K. R. Sachinthana +1VLM RobustnessAdversarial Attacks on VLMs

  24. Robust Context-Aware Detection of Malicious Instructions in Text

    Aug 5, 2026Buzhao Liu, Xinhang Ma, Yevgeniy VorobeychikAdversarial TrainingPrompt Injection Defense

  25. AgentAntibody: An Adaptive Immune System for Defending LLM Agents against Prompt Injection

    Aug 4, 2026Shihao Weng, Yang Feng, Xiaofei Xie +1Continual Learning for LLM AgentsLLM Agent Security

  26. Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure

    Aug 1, 2026Jianshuo Dong, Yiming Liu, Maosen Zhang +6Indirect Prompt InjectionPrompt Injection Defense