LLM Refusal Behavior

LLM: Large Language Model

Momentum

1 paper in the last four weeks, down 86% on the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 54

All topics
CardsList
  1. SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing

    May 28, 2026Almene De Meran Meguimtsop, Maria Leonor Pacheco, Daniel E. AcunaAI for ScienceAdversarial Attacks on LLMs

  2. The Ethics of LLM Sandbox and Persona Dynamics

    May 27, 2026Tim Gebbie, Stewart GebbieLLM GuardrailsLLM Safety Alignment

  3. Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations

    May 27, 2026Matteo Gioele Collu, Riccardo Conte, Alberto Giaretta +4LLM InterpretabilityAdversarial Attacks on LLMs

  4. Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal

    May 26, 2026Kia-Jüng Yang, Dominik Meier, Jiachen Zhao +2CoT ReasoningLLM Refusal Behavior

  5. Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment

    May 20, 2026Roland Pihlakas, Jan Llenzl DagohoyLLM Refusal BehaviorLLM Safety Evaluation

  6. RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts

    May 20, 2026Lukas Weidener, Marko Brkić, Mihailo Jovanović +2LLM Safety BenchmarksLanguage Model Safety Evaluation

  7. Residual Paving: Diagnosing the Routing Bottleneck in Selective Refusal Editing

    May 18, 2026Bryce Hinkley, Peyman NajafiradResidual LearningLLM Refusal Behavior

  8. LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs

    May 13, 2026Rodrigo Nogueira, Thales Sales Almeida, Giovana Kerche Bonás +7Adversarial Prompt GenerationLLM Guardrails

  9. The Refusal--Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models

    May 6, 2026Alif Al Hasan, Sumon BiswasLLM Safety BenchmarksLanguage Model Safety Evaluation

  10. Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection

    May 2, 2026Xulin Hu, Che Wang, Wei Yang Bryan Lim +2Adversarial Attacks on LLMsLLM Refusal Behavior

  11. Please refuse to answer me! Mitigating Over-Refusal in Large Language Models via Adaptive Contrastive Decoding

    Apr 18, 2026Yupeng Qi, Ziyu Lyu, Lixin Cui +2Language Model Safety EvaluationLLM Refusal Behavior

  12. From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions

    Mar 9, 2026Rishab Alagharu, Ishneet Sukhvinder Singh, Shaibi Shamsudeen +2Language Model SteeringLLM Refusal Behavior

  13. There Is More to Refusal in Large Language Models than a Single Direction

    Feb 2, 2026Faaiz Joad, Majd Hawasly, Sabri Boughorbel +2LLM InterpretabilityLLM Refusal Behavior

  14. LLMs Encode Harmfulness and Refusal Separately

    Jul 16, 2025Jiachen Zhao, Jing Huang, Zhengxuan Wu +2LLM InterpretabilityLLM Refusal Behavior

  15. Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP

    Date pendingElisabetta Rocchetti, Alfio FerraraLLM Refusal BehaviorCausal Interventions in Language Models