Beyond Attention Imbalance: Mitigating Hallucinations via Spectral Surgery
Authors: Siqi Lu, Suo Wei, Yongbin Zheng, Jianhang Yao, Wanying Xu, Peng Wang
Organizations: College of Intelligence Science and Technology, National University of Defense Technology, China · School of Computer Science, Northwestern Polytechnical University, China · Ningbo Institute, Northwestern Polytechnical University, China · National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean, China
While Large Vision-Language Models (LVLMs) achieve remarkable success, hallucinations remain a significant barrier to their reliable deployment. Recent studies primarily attribute these issues to cross-modal attention imbalances; most solutions therefore focus on reweighting visual tokens or suppressing language priors. However, such approaches often overlook the spectral characteristics of the visual information flow and frequently rely on Contrastive Decoding (CD), which doubles inference time. Instead of following conventional approaches, we identify two distinct hallucination patterns-Perceptual-Semantic Dissociation and Localized Fixation-and propose FLASH (Frequency-Localized Attention SHaping), a training-free and CD-free framework. FLASH utilizes a Spectral Vortex Score to detect vision heads within multi-head attention layers and applies adaptive spectral modulation to rectify the visual information flow during decoding. Empirical results demonstrate that FLASH achieves a superior balance between performance and efficiency compared to SOTA methods.
Figures & tables
Figure 1 : Two distinct hallucination patterns. (a) seeing but not understanding and (b) biased attention aggregation. Hallucinated responses are highlighted in red.
Figure 2 : Comparison of different methods for mitigating hallucinations. Results are reported relative to Vanilla LLaVA-1.5 on the POPE (COCO-R) dataset. (a) For fixed-weight configurations, a gain coefficient of 1.5 is applied to visual tokens. (b) Circle size indicates the inference time.
Figure 3 : Impact of low-pass filtering on performance and spectral energy distribution. (a) Impact of different frequency band information on model performance. (b) High-frequency energy ratio of the Value Matrix and Attention Weights across layers, with the high-frequency cutoff fixed at 0.5. We employ random sampling in this experiment.
Figure 4 : High-frequency (HF) energy analysis of queries and keys. (a) Comparison of HF energy distributions between hallucinatory and non-hallucinatory samples. (b) HF energy match rate, defined as the ratio HF( qtext ) / HF( kvision ).
Figure 5 : Overview of the FLASH framework. FLASH rectifies hallucinatory signals via spectrum modulation of attention scores and value matrices within the MHA layers. Specifically, spectral shaping of scores is designed to broaden contextual coverage ( seeing comprehensively ), while modulation of value matrices aims to enhance perceptual fidelity ( seeing clearly ).
Figure 6 : Spatial distributions and log-magnitude spectra of vision heads compared with those of other heads.
LVLMs
POPE-MS-COCO
MME
CHAIR
AMBER
Random
Popular
Adversarial
Acc. ↑
F1 ↑
Acc.
F1
Acc.
F1
Score ↑
CHAIR S ↓
CHAIR I ↓
Cover ↑
Hal ↓
Cog ↓
LLaVA-1.5 7B
88.97
88.90
85.63
86.03
79.23
80.99
621.67
48.10
12.75
51.00
30.50
3.20
+VCD
89.07
88.99
85.60
85.99
79.27
81.00
636.67
48.80
12.85
51.55
27.00
2.85
+PAI
89.30
89.27
86.07
86.45
79.23
81.06
636.67
47.80
12.35
47.10
24.25
1.45
+SID
89.40
89.04
85.93
85.93
80.33
81.38
606.67
48.10
12.40
52.40
30.50
2.15
Table 1 : Comparison with SOTA methods. Rows shaded in green denote the performance of vanilla LVLMs, while subsequent rows report results after integrating various hallucination mitigation methods into these models. The evaluation spans both discriminative benchmarks (POPE and MME) and generative benchmarks (CHAIR and AMBER). For each LVLM, the best and second-best results among the mitigation methods are highlighted in pink and purple , respectively. All experiments employed greedy decoding.
Methods
POPE
AMBER
Acc.↑
F1↑
Cover↑
Hal↓
Cog↓
Vanilla
84.61
85.31
51.00
30.50
3.20
w/o Select
85.69
85.89
51.85
25.25
2.20
w/ Select
85.70
86.05
51.00
25.00
2.15
S-stream †
85.38
85.73
50.50
25.25
2.30
V-stream †
85.55
85.91
51.15
25.50
2.20
Table 2 : Comparison of ablation results. ”Select” denotes the visual head selection mechanism. † denotes using only the modulation strategy corresponding to that stream. We report the average results across the three splits of POPE-COCO.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 : Visualization of different hallucination patterns. We asked multiple questions on the same sampled image to observe distinct hallucination patterns. In the figure, green boxes represent correct responses, purple boxes represent hallucinatory responses with PSD, and pink boxes represent hallucinatory responses with LF. Note: Some queries and labels are artificially designed and are intended solely to demonstrate the phenomenon of hallucination.
Figure 8 : Visualization of Attention Heads in Shikra. Due to the intrinsic spatial resolution of the 256-token encoding, we recommend viewing these visualizations at a reduced scale to better perceive the emergent patterns and structural features. In the figure, red text indicates the vision head.
Figure 9 : Visualization of Attention Heads in LLaVA-1.5. In the figure, red text indicates the vision head.
Figure 10 : Proportion of spectral energy in value and attention outputs across varying high-frequency thresholds. Thresholds represent the top percentile of frequency components (e.g., 0.1 denotes the top 10%). When conducting statistical experiments, we employ random sampling methods.
Figure 11 : Performance of Different Tasks on the MME Dataset.
Orders
λv
λs
k
POPE-Random
Accuracy ↑
F1 ↑
Precision ↑
Recall ↑
Greedy
/
/
/
88.97
88.90
89.41
88.40
1
1.0
0.9
5
89.63
89.37
91.72
87.13
2
1.2
0.9
5
90.03
89.77
92.20
87.47
3
1.4
0.9
5
89.97
89.69
92.25
87.27
4
1.6
0.9
5
89.83
89.52
92.41
86.80
Appendix
Table 3 : Comparison of performance across different hyperparameters. In the table, orders 1–5 represent sensitivity experiments for λv ; orders 6–10 represent sensitivity experiments for λs ; orders 11–15 represent sensitivity experiments for k . The green shaded rows indicate the hyperparameter combinations used in this work.
Figure 12 : Comparison of parameter sensitivity experiment results for key hyperparameters.
Orders
Modulation Space
POPE-Random
Logarithmic
Linear
Accuracy ↑
F1 ↑
Precision ↑
Recall ↑
1
V.
S.
90.03
89.77
92.20
87.47
2
S.
V.
89.77
89.51
91.80
87.33
3
V. + S.
/
90.00
89.74
92.13
87.47
4
/
V. + S.
89.70
89.43
91.85
87.13
Appendix
Table 4 : Comparison of modulation strategies in DSM. The V. and S. in the table represent V-stream and S-stream respectively. Logarithmic and Linear represent modulation of the spectrum in the logarithmic domain and linear domain, respectively.
Methods
POPE-R
POPE-P
POPE-A
Average
CHAIR
Times
Acc. ↑
F1 ↑
Acc. ↑
F1 ↑
Acc. ↑
F1 ↑
Acc. ↑
F1 ↑
C. S ↓
C. I ↓
LLaVA-1.5
88.97
88.90
85.63
86.03
79.23
80.99
84.61
85.31
48.10
12.75
1 ×
+VCD
89.07
88.99
85.60
85.99
79.27
81.00
84.65
85.33
48.80
12.85
1.98 ×
+PAI
89.30
89.27
86.07
86.45
79.23
81.06
84.87
85.59
47.80
12.35
1.87 ×
+SID
89.40
89.04
85.93
85.93
80.33
81.38
85.22
85.45
48.10
12.40
2.58 ×
+Ours
90.03
89.77
86.53
86.62
80.53
81.75
85.70
86.05
47.50
12.55
1.69 ×
Appendix
Table 5 : Comparison of performance and efficiency across different methods. In the table, Average represents the performance average across the three POPE-MSCOCO splits. Times reflects the multiple of the inference time of the baseline model (relative to the baseline method) when applying different methods on the POPE-bench dataset.
Figure 13 : Qualitative comparison of the ablation study results.
Figure 14 : Qualitative comparison of different methods for the generative task.
λs
λv
S-stream Layer
V-stream-layer
τ
k
Discrimination Task
LLaVA-1.5 7B
0.9
1.2
16-27
3-32
0.2
5
Shikra 7B
0.9
1.1
16-27
3-32
0.4
5
LLaVA-1.5 13B
0.9
5.0
16-27
3-40
0.2
5
Generation Task
LLaVA-1.5 7B
0.1
1.2
16-27
3-32
0.2
5
Appendix
Table 6 : Hyperparameters for different models. In the table, “V/S-stream layer” indicates the layer within the MHA where the corresponding modulation scheme is applied.
Center of Statistical Research, School of Statistics and Data Science, Southwestern University of Finance and Economics, Chengdu, China. · Department of Biomedical Engineering, College of Design and Engineering, National University of Singapore, Singapore.
1Nanyang Technological University · 2State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · 3Tsinghua University
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · Nanyang Technological University · University of Pennsylvania +4