cs.CLSep 4, 2026

Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models

Authors: Minji Kim, Hyounghun Kim

Organizations: Graduate School of Artificial Intelligence, POSTECH · Department of Computer Science and Engineering, POSTECH

Abstract

Striking a balance between helpfulness and safety remains a fundamental challenge in aligning large language models. To achieve this balance, models should refuse harmful queries (e.g., "How do I shoot someone?") while remaining responsive to benign inputs, even those superficially resembling harmful queries (e.g., "Where can I shoot a good photo?"). However, models often struggle to distinguish genuinely harmful queries from benign queries that contain superficially risky language, resulting in false refusals. In this paper, we address the issue by decomposing a response in the safety-tuning dataset into two distinct components: (i) a boilerplate refusal statement and (ii) a rationale explaining the refusal. Our experiments and analyses show that refusal statements impede accurate discrimination between harmful and benign queries by inducing reliance on superficial cues. In contrast, training solely on rationales reduces false refusals while maintaining a comparable level of safety performance. Rationale-Only benefits also appear in our ICL configuration and remain compatible with the evaluated inference-time mitigation methods. The results emphasize the necessity of precisely curated, fine-grained safety supervision datasets and outline directions for constructing aligned agents that better reconcile helpfulness with safety.

Explore similar work

CardsList
  1. Addressing Over-Refusal in LLMs with Competing Rewards

    Jun 30, 2026Taeyoun Kim, Aviral KumarRL for Language Model ReasoningLLM Safety

  2. Please refuse to answer me! Mitigating Over-Refusal in Large Language Models via Adaptive Contrastive Decoding

    Apr 18, 2026Yupeng Qi, Ziyu Lyu, Lixin Cui +2Language Model Safety EvaluationLLM Refusal Behavior