LLM Safety

LLM: Large Language Model

Momentum

36 papers in the last four weeks, up 200% on the four weeks before. 0.4% of all new papers.

Jul 13Week of Sep 28

Latest papers 272

All topics
CardsList
  1. The Role of Fine-grained Harm Signals in LLM Safety

    Sep 16, 2026Soyeon Park, Seogyeong Jeong, Sunwoo Kim +1LLM InterpretabilityLLM Safety

  2. Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration

    Sep 14, 2026Aashiq Muhamed, Mona T. Diab, Virginia SmithAdversarial Attacks on LLMsLLM Safety

  3. Inoculation Midtraining with Learned Neologisms

    Sep 14, 2026Kyle O'Brien, Edward James Young, Puria Radmard +4LLM AlignmentLLM Safety

  4. K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

    Sep 14, 2026Laura M. Vowels, Matthew J. Vowels, Shivali Sharma +9LLM Safety BenchmarksMental Health Counseling

  5. Beyond Safe Answers: Segment-Aware Listwise Alignment for Reasoning Safety in Large Reasoning Models

    Sep 14, 2026JungMin Yun, Junehyoung Kwon, Hayeong Ryu +3LLM AlignmentDirect Preference Optimization

  6. Semantic Fibers and Cross-Gram Interference: A Calculus of Safety Drift in Overcomplete Representations

    Sep 14, 2026Mohammed Ahnouch, Lotfi ElaachackLanguage Model Safety EvaluationLLM Safety

  7. SIRF: A Spec-Internalized Risk Foundation Model for Industrial Content Risk Control

    Sep 12, 2026Suwan Wu, Yumeng Lin, Pengcheng Yuan +1Content ModerationLLM Safety

  8. An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS

    Sep 12, 2026Roberto Campbell, Momin Abbass, Muneeza Azmat +5LLM Safety AlignmentLow-Rank Adaptation

  9. RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety

    Sep 11, 2026Adithiyan Rajan Indira Saravanan, Kathleen C. FraserLLM Safety BenchmarksLLM Safety

  10. Combating Instruction Conflict via Energy-Driven Latent Conflict Detection

    Sep 8, 2026Mingyu Ma, Yuxin Wu, Jingbo Wang +3LLM Safety

  11. Risk-Conditioned Fine-Tuning of Large Language Models

    Sep 8, 2026Zixuan Liu, Fangzheng Wu, Brian Summa +1Risk-Sensitive RLLLM Safety

  12. IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks

    Sep 3, 2026Saikat Mondal, Mamta, Deeksha Varshney +2LLM Safety BenchmarksMultilingual Language Model Evaluation

  13. When Retrieval Helps: Selective Retrieval for Single-Turn Mental-Health QA

    Sep 3, 2026Hyunseo Oh, Chong-Kwon Kim, Yoonhyuk ChoiRetrieval-Augmented GenerationMental Health Counseling

  14. Selective Knowledge Edit Reversal via Gated Singular Vector Shrinkage

    Sep 2, 2026Weifeng Jiang, Ruirui Chen, Qianren Mao +3Knowledge EditingLLM Safety

  15. GAPS: Dimension-Level Gates for Conditional Activation Steering

    Sep 1, 2026Moghis Fereidouni, Muhammad Umair Haider, Hassan Sajjad +1Language Model SteeringLLM Safety

  16. The Safeguard Worked. Is the LLM System Safer?

    Sep 1, 2026Pingyu Wu, Weiming Zhang, Nenghai YuLLM GuardrailsLLM Safety

  17. Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance

    Aug 24, 2026Or Biton, Tomer Krichli, Itai Allouche +1Language Model DecodingLanguage Model Safety Evaluation

  18. Pragmatic Attack Surface: Vulnerabilities of Implicit Context in Large Language Models

    Aug 10, 2026Bocheng Chen, Han Zi, Roucheng Ou +5Pragmatic Reasoning in Language ModelsAdversarial Attacks on LLMs

  19. Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways

    Aug 10, 2026Shuyi Miao, Wangjie Qiu, Pengyang Shao +4LLM Safety AlignmentLLM Safety

  20. HoloAegis: Frozen Representation, Topological Inference --- Minimally Parametric Safety Manifolds and Their Capability Boundaries for LLM Guardrails

    Aug 9, 2026Tak Ho Alex Li, Kaijie Liu, Lik-Hang Lee +3Language Model Safety EvaluationLLM Guardrails

  21. Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production

    Aug 9, 2026Cong Ming, Jingyi Chen, Bin Liu +4Continual Learning for LLMsLLM Guardrails

  22. Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning

    Aug 5, 2026Yuxuan Huang, Xingyu Zeng, Tianhang Zheng +1LLM Safety AlignmentLLM Security

  23. DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots

    Aug 5, 2026Jared Moore, Andrea Mock, Yifan Mai +9LLM Safety BenchmarksLLM Safety