LLM Defense Mechanisms

Momentum

32 papers in the last four weeks, up 256% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 258

All topics
CardsList
  1. AI Agents May Always Fall for Prompt Injections

    May 17, 2026Sahar Abdelnabi, Eugene BagdasarianArtificial Intelligence SafetyArtificial Intelligence Agents

  2. Asking Back: Interaction-Layer Antidistillation Watermarks

    May 15, 2026Guang Yang, Amir Ghasemian, Fengchen Liu +3WatermarksLLM Defense Mechanisms

  3. EntropyScan: Towards Model-level Backdoor Detection in LVLMs via Visual Attention Entropy

    May 15, 2026Xuanyu Ge, Zhongqi Wang, Jie Zhang +2Clean Label Backdoor AttackLarge Language Model Backbones

  4. A Few Steps Further: Why Defenses Against Malicious Finetuning Erode Under Continued Training

    May 14, 2026Itay Zloczower, Eyal Lenga, Gilad Gressel +1Model-Agnostic DefenseLLM Defense Mechanisms

  5. MemLineage: Lineage-Guided Enforcement for LLM Agent Memory

    May 14, 2026Ciyan Ouyang, Rui HouUntrusted ContentLLM Defense Mechanisms

  6. Mitigating Data Scarcity in Psychological Defense Classification with Context-Aware Synthetic Augmentation

    May 14, 2026Hoang-Thuy-Duong Vu, Quoc-Cuong Pham, Huy-Hieu PhamAugmentationLLM Defense Mechanisms

  7. Inducing Overthink: Hierarchical Genetic Algorithm-based DoS Attack on Black-Box Large Language Reasoning Models

    May 13, 2026Shuqiang Wang, Wei Cao, Jiaqi Weng +4LLM Reasoning StrategiesLarge Reasoning Models

  8. Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy

    May 13, 2026Adarsh Kumarappan, Ananya MujooMulti-Agent Large Language Model SystemsReinforcement Learning From Human Feedback

  9. CoT-Guard: Small Models for Strong Monitoring

    May 12, 2026Nirav Diwan, Han Wang, Berkcan Kapusuzoglu +6Realistic Threat ModelLLM Defense Mechanisms

  10. Re-Triggering Safeguards within LLMs for Jailbreak Detection

    May 11, 2026Zheng Lin, Zhenxing Niu, Haoxuan Ji +2Large Language Model JailbreaksJailbreak Attacks

  11. Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing

    May 11, 2026Zheng Lin, Zhenxing Niu, Haoxuan Ji +1Large Language Model JailbreaksJailbreak Attacks

  12. Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims

    May 11, 2026Phongsakon Mark Konrad, Toygar Tanyel, Serkan AyvazLarge Language Model SafetyModel Fine-Tuning

  13. Strategic commitments shape collective cybersecurity under AI inequality

    May 10, 2026Adeela Bashir, Zia Ush Shamszaman, Zhao Song +2CybersecurityLLM Defense Mechanisms

  14. NEXUS: Continual Learning of Symbolic Constraints for Safe and Robust Embodied Planning

    May 10, 2026Tiehan Cui, Peipei Liu, Yanxu Mao +3Embodied AgentsEmbodied Artificial Intelligence

  15. Not All Turns Matter: Credit Assignment for Multi-Turn Jailbreaking

    May 9, 2026Zhida He, Xiaoyu Wen, Han Qi +7Jailbreak AttacksJailbreaks

  16. Robust Server Defense Against Unreliable Clients in One-Shot Fair Collaborative Machine Learning

    May 9, 2026Chia-Yuan Wu, Frank E. Curtis, Daniel P. RobinsonFederated LearningAlgorithmic Fairness

  17. The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play

    May 8, 2026Gabriele La Malfa, Emanuele La Malfa, Saar Cohen +4Self-PlayArtificial Intelligence Safety

  18. Seed Hijacking of LLM Sampling and Quantum Random Number Defense

    May 8, 2026Ziyang You, Xiaoke Yang, Zhanling Fan +3Attacker Large Language ModelRandom

  19. WebTrap: Stealthy Mid-Task Hijacking of Browser Agents During Navigation

    May 8, 2026Zhichao Liu, Wenbo Pan, Haining Yu +3HijackingBrowser Agent

  20. Fortifying Time Series: DTW-Certified Robust Anomaly Detection

    May 8, 2026Shijie Liu, Tansu Alpcan, Christopher Leckie +1Time-Series Anomaly DetectionTime Series

  21. Nürnberg NLP at PsyDefDetect: Multi-Axis Voter Ensembles for Psychological Defence Mechanism Classification

    May 8, 2026Philipp Steigerwald, Eric Rudolph, Jens AlbrechtPsychologyNatural Language Processing

  22. Beyond Defenses: Manifold-Aligned Regularization for Intrinsic 3D Point Cloud Robustness

    May 8, 2026Pedro Alonso, Chongshou Li, Tianrui LiPoint CloudsRegularization

  23. PropGuard: Safeguarding LLM-MAS via Propagation-Aware Exploration and Remediation

    May 8, 2026Bingyu Yan, Xiaoming Zhang, Jinyu Hou +4LLM Defense MechanismsRemediation

  24. Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks

    May 7, 2026Guoxin Lu, Letian Sha, Qing Wang +4Large Language Model SafetyLarge Language Model Fine-Tuning

  25. SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety

    May 7, 2026Zhe Liu, Zonghao Ying, Wenxin Zhang +5Large Language Model SafetyLarge Language Model Agents

  26. One Turn Too Late: Learning When to Intervene Against Multi-Turn Malicious Intent

    May 7, 2026Xinjie Shen, Rongzhe Wei, Peizhi Niu +6AttackerMalicious Agents

  27. On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference

    May 6, 2026Zhengyi Li, Yakai Wang, Kang Yang +6Transformer ArchitecturesLLM Defense Mechanisms

  28. MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents

    May 5, 2026Ishrith GowdaPoisoningLLM Defense Mechanisms

  29. Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory

    May 4, 2026Yuhui Wang, Tanqiu Jiang, Jiacheng Liang +2Multi-Llm AgentsLarge Language Model Agents

  30. PIIGuard: Mitigating PII Harvesting under Adversarial Sanitization

    May 4, 2026Mingshuo Liu, Yiwei Zha, Min ChenAttacker Large Language ModelData Leakage

  31. Tenability and Weak Semantics: Modeling Non-uniform Defense -- Extended Version

    May 3, 2026Uri Andrews, Luca San Mauro, John SpoerlArgumentation FrameworksLLM Defense Mechanisms

  32. Catching the Infection Before It Spreads: Foresight-Guided Defense in Multi-Agent Systems

    May 3, 2026Yue Ma, Ziyuan Yang, Yi ZhangModel-Based Multi-Agent SystemsLLM Defense Mechanisms

  33. A Theoretical Game of Attacks via Compositional Skills

    May 1, 2026Xinbo Wu, Huan Zhang, Abhishek Umrawal +1Attacker Large Language ModelAdversarial Prompts

  34. Controlled Steering-Based State Preparation for Adversarial-Robust Quantum Machine Learning

    Apr 30, 2026Sahan Sanjaya, Hari Krishna Parvatham, Emma Andrews +1Quantum Machine LearningQuantum Error Correction

  35. TwinGate: Stateful Defense against Decompositional Jailbreaks in Untraceable Traffic via Asymmetric Contrastive Learning

    Apr 30, 2026Bowen Sun, Chaozhuo Li, Yaodong Yang +2Large Language Model JailbreaksJailbreak Success Rates

  36. Security Attack and Defense Strategies for Autonomous Agent Frameworks: A Layered Review with OpenClaw as a Case Study

    Apr 30, 2026Luyao Xu, Xiang ChenOpenclawSecurity

  37. AdaBFL: Multi-Layer Defensive Adaptive Aggregation for Bzantine-Robust Federated Learning

    Apr 30, 2026Zehui Tang, Yuchen Liu, Feihu HuangFederated LearningByzantine Attacks

  38. Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control

    Apr 29, 2026Mahiro Nakao, Kazuhiro TakemotoLarge Language Model SafetyLLM Defense Mechanisms

  39. The Unseen Adversaries: Robust and Generalized Defense Against Adversarial Patches

    Apr 29, 2026Vishesh Kumar, Akshay AgarwalAdversarial TrainingOut-Of-Distribution Detection

  40. SnapGuard: Lightweight Prompt Injection Detection for Screenshot-Based Web Agents

    Apr 28, 2026Mengyao Du, Han Fang, Haokai Ma +4Prompt-Injection DetectorsIndirect Prompt Injection

  41. GAMMAF: A Common Framework for Graph-Based Anomaly Monitoring Benchmarking in LLM Multi-Agent Systems

    Apr 27, 2026Pablo Mateo-Torrejón, Alfonso Sánchez-MaciánGeneralist Graph Anomaly DetectionInterpretable Anomaly Detection

  42. Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing

    Apr 27, 2026Kaisheng Fan, Weizhe Zhang, Yishu Gao +2Inference-Time DefenseLLM Defense Mechanisms

  43. Evaluation of Prompt Injection Defenses in Large Language Models

    Apr 26, 2026Priyal Deep, Shane Emmons, Amy Fox +4Attacker Large Language ModelLLM Defense Mechanisms

  44. Breaking the Secret: Economic Interventions for Combating Collusion in Embodied Multi-Agent Systems

    Apr 26, 2026Qi Liu, Xiaohui Chen, Zhihui Zhao +5Model-Based Multi-Agent SystemsCollusion

  45. UNSEEN: A Cross-Stack LLM Unlearning Defense against AR-LLM Social Engineering Attacks

    Apr 25, 2026Tianlong Yu, Yang Yang, Xiao Luo +6Attacker Large Language ModelLLM Defense Mechanisms

  46. Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers

    Apr 23, 2026Jiali Wei, Ming Fan, Guoheng Sun +3Attacker Large Language ModelClean Label Backdoor Attack

  47. Adaptive Defense Orchestration for RAG: A Sentinel-Strategist Architecture against Multi-Vector Attacks

    Apr 22, 2026Pranav Pallerla, Wilson Naik Bhukya, Bharath Vemula +1Model-Agnostic DefenseLLM Defense Mechanisms

  48. SafeRedirect: Defeating Internal Safety Collapse via Task-Completion Redirection in Frontier LLMs

    Apr 22, 2026Chao Pan, Yu Wu, Xin YaoLarge Language Model SafetyLLM Defense Mechanisms

  49. Adversarial Arena: Crowdsourcing Data Generation through Interactive Competition

    Apr 20, 2026Prasoon Goyal, Sattvik Sahai, Michael Johnston +14Large Language Model GenerationAdversarial Robustness

  50. CASCADE: A Component Ablation and Corpus Audit of a Layered Local Defense for MCP-Based Systems

    Apr 18, 2026İpek Abasıkeleş Turgut, Edip GümüşLLM Defense MechanismsModel Context Protocol

  51. When Safe Concepts Become Unsafe: Multi-Concept Compositional Vulnerabilities in Text-to-Image Models

    Apr 17, 2026Chaoshuo Zhang, Yibo Liang, Mengke Tian +7Large Language Model SafetyVisual Representations