Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibration
Authors: Zixuan Wang, Bingjie Zhang, He Zhao, Dandan Guo
Organizations: School of Artificial Intelligence, Jilin University · CSIRO, Australia · Center of Excellence for Generative AI, King Abdullah University of Science and Technology
Large language models (LLMs) aligned for safety often suffer from over-refusal, incorrectly rejecting benign yet safety-related instructions. Prior studies primarily attribute this to static representation overlap, largely overlooking the underlying dynamic mechanisms. In this paper, we present the mechanistic analysis of over-refusal through the lens of internal routing conflicts within transformer attention. We discover that a sparse subset of Hypersensitive Safety Heads misfires on Hard-Safe prompts, exhibiting abnormal attention entanglement that forcefully binds harmless target entities to refusal semantics. This triggers a severe, high-entropy routing conflict that deprives target entities of necessary attention. To counteract this, we propose Semantic Routing Calibration (SRC), a lightweight, training-free inference framework. SRC precisely localizes and dynamically suppresses these hypersensitive safety heads at the inference stage. Coupled with a dual-branch logits fusion that acts as a safety regularizer during subsequent decoding, SRC seamlessly restores trustworthy reasoning. Extensive experiments demonstrate that SRC alleviates over-refusal, with intrinsic safety performance preserved as much as feasible.
Figures & tables
Figure 1 : Overview of our Semantic Routing Calibration (SRC) framework. (1) Hypersensitive safety head localization: We identify hypersensitive safety heads by comparing noun-centric attention discrepancies between Hard-Safe and Unsafe queries for LLMs. (2) Calibration and generation: We dynamically trigger intervention based on the first token’s refusal tendency for test input. If activated, SRC calibrates these hypersensitive safety heads and applies dual-branch logits fusion to ensure safety-aligned generation.
Figure 2 : Layer-wise semantic routing dynamics on Qwen2.5-7B. (a)–(d) in the first row compares attention allocation ratios across functional token groups, while (e)–(h) in the second row shows attention entropy to reflect routing dispersion. See Figure 12 for Llama-3-8B and Figure 13 for Qwen2.5-1.5B.
Model / Method
Training-Free
Safety & Over-refusal
General Capability
XS
COCO
OR
OK
PHtest
Safety width 0.7pt
MMLU
ARC-e
ARC-c
OBQA
PIQA
Qwen-2.5-1.5B
STL Bianchi et al. [2024]
×
0.73
0.88
0.72
0.75
0.75
0.72 width 0.7pt
0.59
0.77
0.48
0.41
0.76
STL-aug Bianchi et al. [2024]
×
0.75
0.90
0.76
0.76
0.75
0.77 width 0.7pt
0.59
0.77
0.48
0.41
0.76
DCR Lu et al. [2026]
×
0.98
0.98
0.83
0.86
0.86
0.81 width 0.7pt
0.58
0.75
0.47
0.38
0.76
Surgical Wang et al. [2025]
✓
0.81
0.84
0.54
0.84
0.54
0.78 width 0.7pt
0.59
0.76
0.48
0.40
0.76
Table 1 : Overall evaluation results on Qwen-2.5-1.5B, Qwen-2.5-7B and Llama-3-8B. ✓ denotes training-free methods. Left: safety and over-refusal benchmarks. Right: general capability benchmarks.
Figure 3 : Comparison of inference throughput and peak VRAM usage on LLaMA-3-8B. Baseline denotes the safety-aligned model. Evaluated on four NVIDIA A40 GPUs, Float16, batch size 1, max 200 tokens.
Figure 4 : Case study before and after intervention using Llama-3-8B. Attention scores indicate how strongly the generated token attends to each input token.
Target Token
TopK
Strategies
Oversafety ↑
Safety ↑
General ↑
N
L
V
O
Large
Small
Δsim
Fusion
OKtest
Xstest
Xstest(U)
MMLU
✓
-
-
-
-
✓
✓
✓
0.71
0.95
0.485
0.49
✓
-
-
-
✓
-
-
✓
0.96
0.99
0.83
0.48
✓
-
-
-
✓
-
✓
-
0.70
0.93
0.87
0.51
✓
-
-
-
✓
-
✓
✓
0.96
0.99
0.84
0.50
-
-
-
✓
✓
-
✓
✓
0.95
0.97
0.85
0.50
Table 2 : Ablation of SRC components on LLaMA-3-8B. N/L/V/O denote Noun/Last/Verb/Other tokens. “Largest/Smallest” select Top- K heads with highest/lowest Sl,h scores. The highlighted row indicates our final configuration.
Method
Safety Prompt
XSTest ↑
OKTest ↑
CoCoNot ↑
Safety ↑
SRC
×
0.86
0.94
0.97
0.82
✓
0.99
0.96
0.99
0.90
SCANS
×
0.84
0.90
0.97
0.90
✓
0.92
0.95
0.94
0.93
Table 3 : Ablation study on the shared safety prompt for SRC and SCANS on LLaMA-3-8B.
Method
Training Data
OR-Bench ↑
Group 1: Default STL Training
STL
Default
0.34
STL + SRC
Default
0.87
Group 2: OR-Bench-Augmented Training
STL
Default + OR-Bench
0.76
STL + SRC
Default + OR-Bench
0.84
Table 4 : Comparison between OR-Bench-augmented fine-tuning and training-free SRC on Qwen-7B. Higher OR-Bench scores indicate better over-refusal mitigation.
Figure 5 : Heatmap of hypersensitive scores Sl,h on LLaMA-3-8B. Higher values indicate stronger Hard-Safe vs. Unsafe attention discrepancies.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Impact of enhancing different token groups under varying intervention strengths α . Left: over-refusal score on Xstest. Right: safety rate on Xstestunsafe benchmarks.
Token
Top-16
Top-32
Top-64
Jaccard ↑
Spearman ρ↑
Jaccard ↑
Spearman ρ↑
Jaccard ↑
Spearman ρ↑
Noun (Ours)
0.778
0.927
0.699
0.902
0.670
0.857
Verb
0.662
0.554
0.779
0.783
0.707
0.830
Last Token
0.608
0.708
0.537
0.791
0.518
0.755
Appendix
Table 5 : Reproducibility of localized attention heads across independent synthetic datasets. Higher Jaccard similarity and Spearman correlation indicate greater reproducibility.
Method
Comp.
XSTest
OKTest
CoCoNot
Safety
Fusion + Prompt
F+P
81
77
92
86
Head Only
H
92
70
98
86
SRC
H+F+P
99
96
99
90
Appendix
Table 6 : Component ablation on LLaMA-3-8B. H: head intervention; F: logits fusion; P: safety prompt.
Dataset
# Samples
Category
Construction / Source
Usage
Synthetic Dataset Dsyn
M
Hard-Safe / Unsafe
Paired instructions sharing identical templates and verbs, differing only in target noun entities.
Hypersensitive head localization
Alpaca [ Dubois et al., 2024 ]
30
Safe
Randomly sampled benign instruction-following data from Alpaca.
Attention analysis ( Danalyze )
OR-Bench [ Cui et al., 2024 ]
30
Hard-Safe
Seemingly-toxic benign prompts generated from toxic-word seeds and verified by multiple LLMs.
Expert-written and manually verified seemingly-toxic benign prompts./Expert-written and manually verified toxic prompts.
Over-refusal/Safety evaluation
CoCoNot [ Brahman et al., 2024 ]
379
Hard-Safe
Seed prompts expanded by GPT-4 and verified by both humans and LLMs.
Over-refusal evaluation
Appendix
Table 7 : Statistics and construction details of all datasets used in this work. Danalyze is used for semantic routing analysis, while the synthetic paired dataset Dsyn is used for hypersensitive safety head localization. The remaining benchmarks are used for over-refusal, safety, and general capability evaluation.
Figure 7 : Hyperparameter analysis on Llama3-8B. Red line: over-refusal mitigation. Blue line: safety preservation. Results are evaluated on downstream test benchmarks, while the gray dashed line denotes the hyperparameter selected on Danalyze .
Hyperparameter
Qwen2.5-1.5B
Qwen2.5-7B
Llama3-8B
τ
0.10
0.00
0.02
α
0.3
-0.3
-0.2
β
0.7
0.5
0.7
K
32
64
64
Appendix
Table 8 : Hyperparameter settings used in our experiments.
Figure 8 : Δsim distributions of refusal and non-refusal responses across different models. From top to bottom: Llama-3-8B, Qwen2.5-1.5B, and Qwen2.5-7B. The dashed line denotes the model-specific threshold τ used to separate the two distributions for intervention triggering.
Figure 9 : Comparison of different token-group based localization strategies. Noun-based localization achieves the best balance between over-refusal reduction and safety preservation, indicating that hypersensitive safety activation mainly originates from noun-centric semantic routing.
Figure 10 : Trade-off dynamics across different attention head intervention sets on LLaMA-3-8B.
Delay t
OR ↑
Safety ↑
Baseline
0.76
0.98
1
0.86
0.90
2
0.90
0.87
3
0.96
0.85
Appendix
Table 9: Ablation study on logits fusion delay token t . Larger delay values improve over-refusal mitigation but gradually reduce safety preservation.
Setting
XSTest
OKTest
CoCoNot
Safety
Greedy + STL
66
87
87
95
Greedy + SRC
92
91
97
92
Top- p + STL
65
80
85
96
Top- p + SRC
86
88
97
94
Appendix
Table 10 : Performance comparison under greedy and Top- p decoding.
Figure 11 : Effect of constructed dataset size on over-refusal and safety performance on Llama-3-8B. Increasing the dataset size gradually degrades both over-refusal mitigation and safety performance.
Figure 12 : Layer-wise semantic-routing analysis on Llama3-8B across different token categories and safety conditions. The first row illustrates the average attention allocation ratio of generated tokens to different input token groups. The second row displays the attention entropy that quantifies the dispersion of attention distributions.
Figure 13 : Layer-wise semantic-routing analysis on Qwen2.5-1.5B across different token categories and safety conditions. The first row illustrates the average attention allocation ratio of generated tokens to different input token groups. The second row displays the attention entropy that quantifies the dispersion of attention distributions.
Head Category
α
Over-refusal Rate ↓
Safety Rate ↑
Baseline
1.0
52
99
Ours (Top-64)
−0.2
13
88
−0.5
11
85
−0.6
10
75
Lower-ranked (64–128)
−0.5
34
96
−1.3
27
87
Appendix
Table 11 : Ablation study on intervention head selection. Suppressing generic safety heads or lower-ranked secondary heads (64–128) leads to severe safety degradation, whereas targeting our primary hypersensitive safety heads (Top-64) yields an optimal safety-utility balance.
Representation
Best Layer
Accuracy
AUROC
Performance Gap
Clean
Jailbreak
Clean
Jailbreak
Acc
AUROC
Last-token Hidden State
13
92.5
62.9
98.2
43.1
-29.6 ↓
-55.1 ↓
Noun-centric Attention
8
83.3
63.4
91.5
68.8
-19.9 ↓
-22.7 ↓
Appendix
Table 12: Generalization of linear probes trained on AdvBench etc. (Unsafe) and OR-Bench/XSTest (Hard-Safe) under jailbreak distribution shift on Llama-3-8B. Results are reported for the best-performing layer of each representation.
Evaluation Metric
Clean
Jailbreak
Performance Gap
Retention (%)
Safety ↑
90
86
-4 ↓
95.6
XSTest ↑
99
97
-2 ↓
98.0
OR ↑
86
93
+7 ↑
108.1
Appendix
Table 13 : SRC remains robust under the same jailbreak-induced distribution shift without learning a task-specific classifier.
School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University · School of Information, Central University of Finance and Economics · School of Artificial Intelligence, Beijing Normal University +1