LLM Safety Evaluation

LLM: Large Language Model

Momentum

16 papers in the last four weeks, level with the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 177

All topics
CardsList
  1. AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety

    Jun 3, 2026Yanjing Ren, Reza Ebrahimi, TengTeng MaLLM Safety BenchmarksLLM-as-a-Judge

  2. PreAct-Bench: Benchmarking Predictive Monitoring in LLMs

    Jun 3, 2026Hainiu Xu, Italo Luis da Silva, Jiangnan Ye +7LLM Safety BenchmarksMoral Reasoning in Language Models

  3. Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You Need to Break LLMs

    Jun 2, 2026Vincent Limbach, Jonas Dornbusch, David Lüdke +2LLM Jailbreak AttacksLLM Safety Evaluation

  4. DDOR: Delta Debugging for Explainable Overrefusal Testing and Repair

    Jun 2, 2026Qinyan Zhou, Peixin Zhang, Jun Sun +2LLM GuardrailsLLM Refusal Behavior

  5. Food Noise & False Safety: A Systematic Evaluation of How LLMs Fail to Adapt to Eating Disorder Queries with Clinician Feedback

    Jun 1, 2026Giulia Pucci, Emily Hemendinger, Ruizhe Li +3Mental HealthLLM Safety

  6. Lost in Delusion: Examining LLM Safety Under User Delusions and Distress

    May 31, 2026Andrew Aquilina, Chetna Nihalani, Vasudha Varadarajan +3Mental HealthLLM Safety Alignment

  7. TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety

    May 30, 2026Zhepei Hong, Lin Wang, Liting Li +5AI Agent SafetyLLM Safety Evaluation

  8. Position: Anthropomorphic Misalignment Research Needs Stronger Evidence

    May 29, 2026Vansh Gupta, Peter Nutter, Samuel Stante +5LLM Safety Evaluation

  9. Triaging Threats to Specialized Guardrails

    May 29, 2026Wenjie Jacky Mo, Xiaofei Wen, Rui Cai +6LLM Safety BenchmarksLLM Guardrails

  10. Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles

    May 28, 2026Drishti Goel, Agam Goyal, Veda Duddu +8HealthcareLLM Auditing

  11. SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing

    May 28, 2026Almene De Meran Meguimtsop, Maria Leonor Pacheco, Daniel E. AcunaAI for ScienceAdversarial Attacks on LLMs

  12. A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities

    May 28, 2026Kenji Imamura, Masao Ideuchi, Atsushi FujitaAI Safety EvaluationLLM Safety Evaluation

  13. Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

    May 28, 2026Aditya Nawal, Manit Baser, Mohan GurusamyAI Agent Security BenchmarksLLM Safety Evaluation

  14. Code as a Weapon: A Consensus-Labeled Prompt Bank for Measuring Coding-Model Compliance with Malicious-Code Requests

    May 27, 2026Richard J. Young, Gregory D. MoodyClassificationLLM Safety Evaluation

  15. Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems

    May 26, 2026Aman Priyanshu, Supriti Vijay, Esha PahwaData LeakageMulti-Agent LLM Systems

  16. When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries

    May 26, 2026Yige Li, Jun Sun, Wei Zhao +5LLM Safety BenchmarksHealthcare

  17. KZ-SafetyPrompts: A Kazakh Safety Evaluation Prompt Dataset for Large Language Models

    May 26, 2026Wajdi Zaghouani, Shimaa Amer Ibrahim, Aruzhan Muratbek +2LLM Safety BenchmarksLanguage Model Safety Evaluation

  18. Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety

    May 25, 2026Fay Elhassan, David Sasu, Alexandra Kulinkina +2LLM Safety BenchmarksHuman Preference Evaluation

  19. SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks

    May 25, 2026Yanhang Li, Zhichao Fan, Zexin ZhuangPairwise ComparisonPairwise Preference Evaluation

  20. From Automation to Collaboration: Human-in-the-Loop Methods for Safe and Trustworthy NLP

    May 24, 2026Most. Sharmin Sultana Samu, MD. Tanvir Ahmed Seum, Md. Rakibul IslamLLM AuditingHuman-in-the-Loop Annotation

  21. Furina: Fragmented Uncertainty-Driven Refusal Instability Attack

    May 24, 2026Tongxi Wu, Jian Zhang, Yang GaoAdversarial Attacks on VLMsLLM Safety

  22. Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment

    May 20, 2026Roland Pihlakas, Jan Llenzl DagohoyLLM Refusal BehaviorLLM Safety Evaluation

  23. RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts

    May 20, 2026Lukas Weidener, Marko Brkić, Mihailo Jovanović +2LLM Safety BenchmarksLanguage Model Safety Evaluation

  24. Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents

    May 18, 2026Ahmad Al-Tawaha, Shangding Gu, Peizhi Niu +2Long-Horizon Agent EvaluationLLM Safety Evaluation