LLM Defense Mechanisms

Momentum

32 papers in the last four weeks, up 256% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 258

All topics
CardsList
  1. DefenseSplat: Enhancing the Robustness of 3D Gaussian Splatting via Frequency-Aware Filtering

    Feb 22, 2026Yiran Qiao, Yiren Lu, Yunlai Zhou +43D Gaussian3D Reconstruction

  2. Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent Defense

    Feb 9, 2026Jiacheng Liu, Yaxin Luo, Jiacheng Cui +3CaptchaAgentic Benchmarks

  3. Noise-Aware and Dynamically Adaptive Federated Defense Framework for SAR Image Target Recognition

    Dec 31, 2025Yuchao Hou, Zixuan Zhang, Jie Wang +9Synthetic Aperture RadarFederated Learning

  4. MAS-Shield: A Defense Framework for Secure and Efficient LLM MAS

    Nov 28, 2025Kaixiang Wang, Zhaojiacheng Zhou, Bunyod Suvonov +8LLM Defense MechanismsLinguistics

  5. BackdoorVLM: A Benchmark for Backdoor Attacks and Defenses on Vision-Language Models

    Nov 24, 2025Juncheng Li, Yige Li, Hanxun Huang +5LLM Defense Mechanisms

  6. GraphToxin: Reconstructing Full Unlearned Graphs from Graph Unlearning

    Nov 14, 2025Ying Song, Balaji PalanisamyMembership Inference AttacksUnlearnable Examples

  7. "Give a Positive Review Only": An Early Investigation Into In-Paper Prompt Injection Attacks and Defenses for AI Reviewers

    Nov 3, 2025Qin Zhou, Zhexin Zhang, Zhi Li +1Peer ReviewLLM Defense Mechanisms

  8. Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety

    Oct 18, 2025Vamshi Krishna Bonagiri, Ponnurangam Kumaragurum, Khanh Nguyen +1Large Language Model AgentsLarge Language Model Safety

  9. Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks

    Oct 1, 2025Shoumik Saha, Jifan Chen, Sam Mayers +3Jailbreak AttacksSecurity Evaluation

  10. Adversarial Defense in Cybersecurity: A Systematic Review of GANs for Threat Detection and Mitigation

    Sep 24, 2025Tharcisse Ndayipfukamiye, Jianguo Ding, Doreen Sebastian Sarwatt +2Threat DetectionCybersecurity

  11. Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs

    Aug 9, 2025Jinhwa Kim, Ian G. HarrisJailbreak AttacksLarge Language Models Fail

  12. CO-DEFEND: Continuous Decentralized Federated Learning for Secure DoH-Based Threat Detection

    Apr 2, 2025Diego Cajaraville-Aboy, Marta Moure-Garrido, Carlos Beis-Penedo +5Threat DetectionFederated Learning

  13. Eraser: Jailbreaking Defense in Large Language Models via Unlearning Harmful Knowledge

    Apr 8, 2024Weikai Lu, Ziqian Zeng, Jianwei Wang +5Large Language Model JailbreaksJailbreaks

  14. Builder, Defender, Breaker: Measurable Independence and Bounded Autonomy When Generative Models Build, Defend and Test Software

    Date pendingMohamed Chahine GhanemLLM Defense MechanismsCode Generation

  15. A False Average: Pooled CoT-Monitor Accuracy Conceals a Reasoning-Dependent Fragility

    Date pendingShikhar Shiromani, Leo RichterEvasionChain-of-Thought Reasoning