Adversarial Attacks on LLMs

LLM: Large Language Model

Latest papers 238

All topics
CardsList
  1. One Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent Guardrails

    Oct 8, 2026Seyedarmin Azizi, Erfan Baghaei Potraghloo, Massoud PedramLLM GuardrailsLLM Agent Security

  2. Secure Speculative Decoding for Large Language Models

    Oct 6, 2026Yichi Zhang, Zhiqi Wang, Neil Gong +1Adversarial Attacks on LLMsSpeculative Decoding

  3. Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge

    Oct 5, 2026Wonjun Lee, Kyungsik Yang, Gaeun Ji +5Adversarial Attacks on LLMsLLM Security

  4. Jailbreaking Open-Weight LLMs via Random Embedding Perturbations

    Oct 5, 2026Abhinav Sudhakar Dubey, Scott Sirri, Vaggos Chatziafratis +1Adversarial AttacksAdversarial Attacks on LLMs

  5. Reward Stealing Attack on Large Language Models

    Oct 5, 2026Jiaming Qian, Pengyang Zhou, Jiahe Xu +1LLM AlignmentModel Extraction Attacks

  6. Reflections and Fragments: Securing LLMs Against Sequential Mosaic Attacks

    Oct 4, 2026Emanuele La Malfa, Saar Cohen, Gabriele La Malfa +4Adversarial Attacks on LLMsLLM Safety

  7. TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety

    Oct 1, 2026Fengpeng Li, Kemou Li, Qizhou Wang +3LLM Fine-TuningAdversarial Attacks on LLMs

  8. CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion

    Sep 30, 2026Zhen Liang, Hai Huang, Wentao ChenAdversarial Attacks on LLMsJailbreak Attacks

  9. Hiding in Plain Sight: Decoupling Pretext from Actuation for Skill Poisoning in LLM Agents

    Sep 30, 2026Wenxin Wu, Lingyong Yan, Lei Sha +2Adversarial Attacks on LLMsLLM Agent Security

  10. MADBench: Benchmarking the Security of Multi-Agent Debate

    Sep 30, 2026Yuwan Liu, Jiaming Zhang, Yue Huang +1Adversarial Attacks on LLMsMulti-Agent Debate

  11. ACTR: Aligning Thoughts and Responses for Multilingual Safety in Reasoning LLMs

    Sep 29, 2026Xianhui Zhang, Jian Yu, Chengyu Xie +6Adversarial Attacks on LLMsLLM Safety Alignment

  12. Controlled Decoding Attacks on Black-Box LLMs

    Sep 29, 2026Jesson Wang, Shawn Li, Wei Yang +4Constrained DecodingAdversarial Attacks on LLMs

  13. LLMs Learn to Evade Latent Monitors from Prior Feedback Alone

    Sep 29, 2026Hugo Lyons Keenan, Christopher Leckie, Sarah ErfaniLLM AuditingAdversarial Attacks on LLMs

  14. Epistemic Policy Divergence in Multi-Turn LLM Contamination: A Protocol-Gradient Investigation

    Sep 28, 2026Fahrell Giovanny, Geby Bayuningtyas, Sahrul Mukharom +1Knowledge Conflicts in Language ModelsAdversarial Attacks on LLMs

  15. One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs

    Sep 27, 2026Sen Nie, Jie Zhang, Zhongqi Wang +2Adversarial Attacks on VLMsAdversarial Attacks on LLMs

  16. SafeMol: Dual-Modality Safety Alignment for Molecular Multimodal Models

    Sep 27, 2026Xinmiao Wang, Ruijie Wang, Menghui Wang +5Adversarial Attacks on LLMsLLM Safety Alignment

  17. Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning

    Sep 27, 2026Chenlong Yin, Xiaolong Jin, Wei Zou +2Adversarial Attacks on LLMsLLM Red Teaming

  18. JevOut: Natural Context Can Flip Decision Models

    Sep 24, 2026Zixiang Xu, Zirui Song, Chiyu Zhang +4Adversarial Attacks on LLMsLLM Decision-Making

  19. Prefilling the Reasoning Channel: Output-Prefix Attacks on Reasoning LLMs

    Sep 24, 2026Lukáš Brůna, Robert Bridges, Adam EkAdversarial Attacks on LLMsPrompt Injection Attacks

  20. The Tokens Remember: When Tokenization Bypasses Knowledge Editing and Unlearning

    Sep 24, 2026Manit Baser, Aditya Nawal, Dinil Mon Divakaran +1Adversarial Attacks on LLMsPrivacy Leakage in Language Models

  21. ALIBI: Adversarial Legitimacy Injection in Binary Input against LLM Malware Analyzers

    Sep 17, 2026Hyeongjun Choi, Wonyoung Jung, Haehoon Seo +1Adversarial Attacks on LLMsLLM Security