Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibration
Authors: Zixuan Wang, Bingjie Zhang, He Zhao, Dandan Guo
Organizations: School of Artificial Intelligence, Jilin University · CSIRO, Australia · Center of Excellence for Generative AI, King Abdullah University of Science and Technology
Large language models (LLMs) aligned for safety often suffer from over-refusal, incorrectly rejecting benign yet safety-related instructions. Prior studies primarily attribute this to static representation overlap, largely overlooking the underlying dynamic mechanisms. In this paper, we present the mechanistic analysis of over-refusal through the lens of internal routing conflicts within transformer attention. We discover that a sparse subset of Hypersensitive Safety Heads misfires on Hard-Safe prompts, exhibiting abnormal attention entanglement that forcefully binds harmless target entities to refusal semantics. This triggers a severe, high-entropy routing conflict that deprives target entities of necessary attention. To counteract this, we propose Semantic Routing Calibration (SRC), a lightweight, training-free inference framework. SRC precisely localizes and dynamically suppresses these hypersensitive safety heads at the inference stage. Coupled with a dual-branch logits fusion that acts as a safety regularizer during subsequent decoding, SRC seamlessly restores trustworthy reasoning. Extensive experiments demonstrate that SRC alleviates over-refusal, with intrinsic safety performance preserved as much as feasible.
Figures & tables
Figure 1 : Overview of our Semantic Routing Calibration (SRC) framework. (1) Hypersensitive safety head localization: We identify hypersensitive safety heads by comparing noun-centric attention discrepancies between Hard-Safe and Unsafe queries for LLMs. (2) Calibration and generation: We dynamically trigger intervention based on the first token’s refusal tendency for test input. If activated, SRC calibrates these hypersensitive safety heads and applies dual-branch logits fusion to ensure safety-aligned generation.
Figure 2 : Layer-wise semantic routing dynamics on Qwen2.5-7B. (a)–(d) in the first row compares attention allocation ratios across functional token groups, while (e)–(h) in the second row shows attention entropy to reflect routing dispersion. See Figure 12 for Llama-3-8B and Figure 13 for Qwen2.5-1.5B.
Model / Method
Training-Free
Safety & Over-refusal
General Capability
XS
COCO
OR
OK
PHtest
Safety width 0.7pt
MMLU
ARC-e
ARC-c
OBQA
PIQA
Qwen-2.5-1.5B
STL Bianchi et al. [2024]
×
0.73
0.88
0.72
0.75
0.75
0.72 width 0.7pt
0.59
0.77
0.48
0.41
0.76
STL-aug Bianchi et al. [2024]
×
0.75
0.90
0.76
0.76
0.75
0.77 width 0.7pt
0.59
0.77
0.48
0.41
0.76
DCR Lu et al. [2026]
×
0.98
0.98
0.83
0.86
0.86
0.81 width 0.7pt
0.58
0.75
0.47
0.38
0.76
Surgical Wang et al. [2025]
✓
0.81
0.84
0.54
0.84
0.54
0.78 width 0.7pt
0.59
0.76
0.48
0.40
0.76
Table 1 : Overall evaluation results on Qwen-2.5-1.5B, Qwen-2.5-7B and Llama-3-8B. ✓ denotes training-free methods. Left: safety and over-refusal benchmarks. Right: general capability benchmarks.
Figure 3 : Comparison of inference throughput and peak VRAM usage on LLaMA-3-8B. Baseline denotes the safety-aligned model. Evaluated on four NVIDIA A40 GPUs, Float16, batch size 1, max 200 tokens.
Figure 4 : Case study before and after intervention using Llama-3-8B. Attention scores indicate how strongly the generated token attends to each input token.
Target Token
TopK
Strategies
Oversafety ↑
Safety ↑
General ↑
N
L
V
O
Large
Small
Δsim
Fusion
OKtest
Xstest
Xstest(U)
MMLU
✓
-
-
-
-
✓
✓
✓
0.71
0.95
0.485
0.49
✓
-
-
-
✓
-
-
✓
0.96
0.99
0.83
0.48
✓
-
-
-
✓
-
✓
-
0.70
0.93
0.87
0.51
✓
-
-
-
✓
-
✓
✓
0.96
0.99
0.84
0.50
-
-
-
✓
✓
-
✓
✓
0.95
0.97
0.85
0.50
Table 2 : Ablation of SRC components on LLaMA-3-8B. N/L/V/O denote Noun/Last/Verb/Other tokens. “Largest/Smallest” select Top- K heads with highest/lowest Sl,h scores. The highlighted row indicates our final configuration.
Method
Safety Prompt
XSTest ↑
OKTest ↑
CoCoNot ↑
Safety ↑
SRC
×
0.86
0.94
0.97
0.82
✓
0.99
0.96
0.99
0.90
SCANS
×
0.84
0.90
0.97
0.90
✓
0.92
0.95
0.94
0.93
Table 3 : Ablation study on the shared safety prompt for SRC and SCANS on LLaMA-3-8B.
Method
Training Data
OR-Bench ↑
Group 1: Default STL Training
STL
Default
0.34
STL + SRC
Default
0.87
Group 2: OR-Bench-Augmented Training
STL
Default + OR-Bench
0.76
STL + SRC
Default + OR-Bench
0.84
Table 4 : Comparison between OR-Bench-augmented fine-tuning and training-free SRC on Qwen-7B. Higher OR-Bench scores indicate better over-refusal mitigation.
Figure 5 : Heatmap of hypersensitive scores Sl,h on LLaMA-3-8B. Higher values indicate stronger Hard-Safe vs. Unsafe attention discrepancies.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Impact of enhancing different token groups under varying intervention strengths α . Left: over-refusal score on Xstest. Right: safety rate on Xstestunsafe benchmarks.
Token
Top-16
Top-32
Top-64
Jaccard ↑
Spearman ρ↑
Jaccard ↑
Spearman ρ↑
Jaccard ↑
Spearman ρ↑
Noun (Ours)
0.778
0.927
0.699
0.902
0.670
0.857
Verb
0.662
0.554
0.779
0.783
0.707
0.830
Last Token
0.608
0.708
0.537
0.791
0.518
0.755
Appendix
Table 5 : Reproducibility of localized attention heads across independent synthetic datasets. Higher Jaccard similarity and Spearman correlation indicate greater reproducibility.
Method
Comp.
XSTest
OKTest
CoCoNot
Safety
Fusion + Prompt
F+P
81
77
92
86
Head Only
H
92
70
98
86
SRC
H+F+P
99
96
99
90
Appendix
Table 6 : Component ablation on LLaMA-3-8B. H: head intervention; F: logits fusion; P: safety prompt.
Dataset
# Samples
Category
Construction / Source
Usage
Synthetic Dataset Dsyn
M
Hard-Safe / Unsafe
Paired instructions sharing identical templates and verbs, differing only in target noun entities.
Hypersensitive head localization
Alpaca [ Dubois et al., 2024 ]
30
Safe
Randomly sampled benign instruction-following data from Alpaca.
Attention analysis ( Danalyze )
OR-Bench [ Cui et al., 2024 ]
30
Hard-Safe
Seemingly-toxic benign prompts generated from toxic-word seeds and verified by multiple LLMs.
Expert-written and manually verified seemingly-toxic benign prompts./Expert-written and manually verified toxic prompts.
Over-refusal/Safety evaluation
CoCoNot [ Brahman et al., 2024 ]
379
Hard-Safe
Seed prompts expanded by GPT-4 and verified by both humans and LLMs.
Over-refusal evaluation
Appendix
Table 7 : Statistics and construction details of all datasets used in this work. Danalyze is used for semantic routing analysis, while the synthetic paired dataset Dsyn is used for hypersensitive safety head localization. The remaining benchmarks are used for over-refusal, safety, and general capability evaluation.
Figure 7 : Hyperparameter analysis on Llama3-8B. Red line: over-refusal mitigation. Blue line: safety preservation. Results are evaluated on downstream test benchmarks, while the gray dashed line denotes the hyperparameter selected on Danalyze .
Hyperparameter
Qwen2.5-1.5B
Qwen2.5-7B
Llama3-8B
τ
0.10
0.00
0.02
α
0.3
-0.3
-0.2
β
0.7
0.5
0.7
K
32
64
64
Appendix
Table 8 : Hyperparameter settings used in our experiments.
Figure 8 : Δsim distributions of refusal and non-refusal responses across different models. From top to bottom: Llama-3-8B, Qwen2.5-1.5B, and Qwen2.5-7B. The dashed line denotes the model-specific threshold τ used to separate the two distributions for intervention triggering.
Figure 9 : Comparison of different token-group based localization strategies. Noun-based localization achieves the best balance between over-refusal reduction and safety preservation, indicating that hypersensitive safety activation mainly originates from noun-centric semantic routing.
Figure 10 : Trade-off dynamics across different attention head intervention sets on LLaMA-3-8B.
Delay t
OR ↑
Safety ↑
Baseline
0.76
0.98
1
0.86
0.90
2
0.90
0.87
3
0.96
0.85
Appendix
Table 9: Ablation study on logits fusion delay token t . Larger delay values improve over-refusal mitigation but gradually reduce safety preservation.
Setting
XSTest
OKTest
CoCoNot
Safety
Greedy + STL
66
87
87
95
Greedy + SRC
92
91
97
92
Top- p + STL
65
80
85
96
Top- p + SRC
86
88
97
94
Appendix
Table 10 : Performance comparison under greedy and Top- p decoding.
Figure 11 : Effect of constructed dataset size on over-refusal and safety performance on Llama-3-8B. Increasing the dataset size gradually degrades both over-refusal mitigation and safety performance.
Figure 12 : Layer-wise semantic-routing analysis on Llama3-8B across different token categories and safety conditions. The first row illustrates the average attention allocation ratio of generated tokens to different input token groups. The second row displays the attention entropy that quantifies the dispersion of attention distributions.
Figure 13 : Layer-wise semantic-routing analysis on Qwen2.5-1.5B across different token categories and safety conditions. The first row illustrates the average attention allocation ratio of generated tokens to different input token groups. The second row displays the attention entropy that quantifies the dispersion of attention distributions.
Head Category
α
Over-refusal Rate ↓
Safety Rate ↑
Baseline
1.0
52
99
Ours (Top-64)
−0.2
13
88
−0.5
11
85
−0.6
10
75
Lower-ranked (64–128)
−0.5
34
96
−1.3
27
87
Appendix
Table 11 : Ablation study on intervention head selection. Suppressing generic safety heads or lower-ranked secondary heads (64–128) leads to severe safety degradation, whereas targeting our primary hypersensitive safety heads (Top-64) yields an optimal safety-utility balance.
Representation
Best Layer
Accuracy
AUROC
Performance Gap
Clean
Jailbreak
Clean
Jailbreak
Acc
AUROC
Last-token Hidden State
13
92.5
62.9
98.2
43.1
-29.6 ↓
-55.1 ↓
Noun-centric Attention
8
83.3
63.4
91.5
68.8
-19.9 ↓
-22.7 ↓
Appendix
Table 12: Generalization of linear probes trained on AdvBench etc. (Unsafe) and OR-Bench/XSTest (Hard-Safe) under jailbreak distribution shift on Llama-3-8B. Results are reported for the best-performing layer of each representation.
Evaluation Metric
Clean
Jailbreak
Performance Gap
Retention (%)
Safety ↑
90
86
-4 ↓
95.6
XSTest ↑
99
97
-2 ↓
98.0
OR ↑
86
93
+7 ↑
108.1
Appendix
Table 13 : SRC remains robust under the same jailbreak-induced distribution shift without learning a task-specific classifier.
Safety training on language models often induces over-refusal: improved safety on harmful prompts at the cost of increased refusal on harmless ones. Though this trade-off can be mitigated by training models with reinforcement learning (RL) to reason before answering, it does not remove the underlying problem that reasoning can often be a "rubber stamp" for a predetermined response. In this paper, we address the safety-refusal trade-off by rethinking how models are trained to reason about safety. Our key insight is that unsafe reasoning can itself serve as a useful exploratory signal. Rather than preemptively blocking harmful thoughts, we encourage the model to sufficiently explore unsafe reasoning but produce a safe response. The harmful exploration improves the model's ability to distinguish harmful from harmless prompts by resolving ambiguity, allowing it to remain safe while complying only when appropriate. We cast this as an adversarial optimization problem in which a reasoning player explores strategies for producing an unsafe response and an answer player ensures that the final output is safe. We train a single model with dense rewards to play both roles within one chain-of-thought, across different segments. To achieve this, we find that process rewards are crucial for stable optimization of competing objectives. Our resulting model SEAR deliberately engages in harmful reasoning as exploration while reliably flipping back to a safe answer. We demonstrate that this behavior helps mitigate over-refusal and defend against attacks that directly manipulate the reasoning to be harmful.
Safety-aligned large language models (LLMs) often generate refusal responses to harmless queries due to the over-refusal problem. However, existing methods for mitigating over-refusal cannot maintain a low refusal ratio for harmless queries while keeping a high refusal ratio for malicious ones. In this paper, we analyze how system prompts with varying safety levels affect LLM refusal behaviors when facing over-refusal queries. A key observation is that, when LLMs suffer from the over-refusal issue, non-refusal tokens remain present in the next-token candidate list, but the model systematically fails to select them, despite the generation of refusal tokens. Based on this observation, we propose a training-free and model-agnostic approach, Adaptive Contrastive Decoding (AdaCD), to mitigate over-refusal while maintaining LLM safety. First, AdaCD compares the output distributions of the LLM with or without an extreme safety system prompt to refine the refusal token distribution. Second, we introduce an adaptive contrastive decoding strategy that dynamically incorporates or removes the refusal token distribution, adaptively boosting the probability of selecting refusal or non-refusal tokens. Experimental results on five benchmark datasets show that, on average, AdaCD reduces the refusal ratio for over-refusal queries by 10.35%, yet still increases the refusal ratio for malicious queries by 0.13%. Code is available at https://github.com/OutdoorManofML/AdaCD.
Yupeng Qi, Ziyu Lyu, Lixin Cui +2
School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University · School of Information, Central University of Finance and Economics · School of Artificial Intelligence, Beijing Normal University +1
While safety alignment and guardrails help large language models (LLMs) avoid harmful outputs, they can also induce overrefusal, i.e., unwarranted rejection of benign queries that merely appear risky. We present DDOR (Delta Debugging for OverRefusal), a fully automated and explainable framework for overrefusal testing and repair in a black-box setting, where only model inputs and outputs are accessible and internal safety mechanisms remain opaque. DDOR applies delta debugging to localize minimal refusal-triggering fragments (mRTFs) that provide phrase-level, explainable evidence for why a refusal occurs. Conditioned on these mRTFs, DDOR generates diverse, context-rich prompts and performs multi-oracle validation to filter intrinsically unsafe or ambiguous cases, producing scalable and model-specific overrefusal test suites (approximately 1K cases per model). Beyond evaluation, we further leverage localized mRTFs to perform targeted prompt repair, substantially reducing overrefusal while preserving the original intent and maintaining safety on genuinely harmful inputs. Overall, DDOR offers a practical end-to-end solution to both evaluate and mitigate overrefusal, improving LLM usability without sacrificing safety.
Qinyan Zhou, Peixin Zhang, Jun Sun +2
Southeast University, China · School of Computing and Information Systems, Singapore Management University, Singapore · Zhejiang University, China +1