LLM Safety

LLM: Large Language Model

Momentum

36 papers in the last four weeks, up 200% on the four weeks before. 0.4% of all new papers.

Jul 13Week of Sep 28

Latest papers 272

All topics
CardsList
  1. TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts

    Apr 17, 2026Hua-Rong Chu, Kuan-Chun Wang, Yao-Te HuangLanguage Model Safety EvaluationLLM Guardrails

  2. FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models

    Apr 16, 2026Zixuan Weng, Jinghuai Zhang, Kunlin Cai +3Language Model SteeringLLM Safety

  3. Segment-Level Coherence for Robust Harmful Intent Probing in LLMs

    Apr 16, 2026Xuanli He, Bilgehan Sel, Faizan Ali +3LLM SafetyLLM Jailbreak Attacks

  4. CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification

    Apr 16, 2026Yian Wang, Yuen Chen, Agam Goyal +1Causal Reasoning in Language ModelsLLM Safety

  5. HazardArena: Evaluating Semantic Safety in Vision-Language-Action Models

    Apr 14, 2026Zixing Chen, Yifeng Gao, Li Wang +8RoboticsAI Safety Evaluation

  6. Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types

    Apr 10, 2026Hadas Orgad, Boyi Wei, Kaden Zheng +4LLM AlignmentEmergent Misalignment in Language Models

  7. Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services

    Mar 11, 2026Fabrizio Dimino, Bhaskarjit Sarmah, Stefano PasqualiLanguage Model Safety EvaluationLLM Safety

  8. In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement

    Jan 19, 2026Anudeex Shetty, Aditya Joshi, Salil S. KanhereLanguage Model Safety EvaluationAdversarial Attacks on LLMs

  9. Test-Time Detoxification without Training or Learning Anything

    Jan 14, 2026Baturay Saglam, Dionysis KalogeriasLarge Language Model-Guided OptimizationLLM Safety

  10. Large language models can effectively convince people to believe conspiracies

    Jan 8, 2026Thomas H. Costello, Kellin Pelrine, Matthew Kowal +6Deception in Language ModelsLLM Safety

  11. Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control

    Jan 6, 2026Harshvardhan Saini, Yiming Tang, Dianbo LiuLanguage Model SteeringLLM Interpretability

  12. The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety

    Jan 3, 2026Ian Rios-SialerLLM Safety

  13. Safety boundary maintenance in consumer AI systems responding to pediatric health queries: a cross-platform benchmark evaluation under naturalistic and adversarially pressured conditions

    Dec 26, 2025Vahideh Zolfaghari, Leila Mashhadi, Mitra Ahadi +2HealthcareLLM Safety

  14. SGM: Safety Glasses for Multimodal Large Language Models via Neuron-Level Detoxification

    Dec 17, 2025Hongbo Wang, AprilPyone MaungMaung, Isao EchizenMultimodal RobustnessLLM Safety

  15. Training-Free Policy Violation Detection via Activation-Space Whitening in LLMs

    Dec 3, 2025Oren Rachmil, Avishag Shapira, Roy Betser +5Language Model Safety EvaluationLLM Auditing

  16. A Neurosymbolic Approach to Natural Language Formalization and Verification

    Nov 12, 2025Chenyang An, Sam Bayless, Stefano Buliani +27Formal VerificationLLM Safety

  17. Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs

    Aug 9, 2025Jinhwa Kim, Ian G. HarrisLLM Safety AlignmentLLM Safety

  18. LLMs Encode Harmfulness and Refusal Separately

    Jul 16, 2025Jiachen Zhao, Jing Huang, Zhengxuan Wu +2LLM InterpretabilityLLM Refusal Behavior

  19. PRISON: Unmasking the Criminal Potential of Large Language Models

    Jun 19, 2025Xinyi Wu, Geng Hong, Pei Chen +3Language Model Safety EvaluationLLM Auditing

  20. HauntAttack: When Attack Follows Reasoning as a Shadow

    Jun 8, 2025Jingyuan Ma, Rui Li, Zheng Li +4Adversarial Attacks on LLMsJailbreak Attacks

  21. Popular but Wrong: Understanding and Mitigating LLM Overconfidence through Knowledge Popularity

    May 23, 2025Shiyu Ni, Keping Bi, Jiafeng Guo +1LLM Hallucination MitigationConfidence Estimation in Language Models

  22. Decoding One Safety Trigger Token for Balancing Safety and Usability in Large Language Models

    May 12, 2025Haoran Gu, Handing Wang, Yi Mei +2LLM Safety AlignmentLLM Safety

  23. Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models

    Mar 8, 2025Thomas Winninger, Boussad Addad, Katarzyna KapustaAdversarial Prompt GenerationAdversarial Attacks on LLMs

  24. Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models

    Dec 21, 2024Yanxu Mao, Peipei Liu, Tiehan Cui +3Adversarial Attacks on LLMsMultimodal Large Language Models