cs.SDSep 22, 2026

REVE: Efficient Hallucination Correction for Large Audio-Language Models via Reused Encoder States

Authors: Hongjin Song, Jiasheng Kuang, Xinyu Yang, Qiuyu Fang, Ziyu Wu, Guowu Tan, Xiang Xie

Organizations: Beijing Institute of Technology, Zhuhai, Guangdong, China

Abstract

Large audio-language models may mention acoustic events that are absent from the input. A separate audio event detector can verify these mentions, but doing so requires a second audio encoder and a separate forward pass. We propose Reused Encoder States for Verifying Events (REVE), a lightweight method that uses states already computed by the target model. One readout summarizes class scores across audio frames, while another uses pooled states from four consecutive frame intervals. Class-aware score fusion combines their outputs to verify generated event mentions without encoding the audio again. On AudioSet, REVE removes 92.9% of label-unsupported mentions under a faithful-mention recall constraint. With fewer added parameters and no second audio-encoding pass, REVE achieves a reduction comparable to those of CED-Tiny and CED-Base. Its complete verification latency is about 1/18 of the CED-Base path. Results on controlled DESED mixtures and different target-model architectures further confirm the effectiveness of encoder-state reuse.

Figures & tables

Explore similar work

CardsList
  1. MISHAP-Bench: A Hallucination Benchmark for Large Audio-Language Models

    Sep 27, 2026Zhi Wen Soi, Giulio Segalini, Jian-Jia Chen +1Large Audio Language ModelsLarge Language Model Hallucination

  2. Silence is Golden: Mitigating Hallucinations in Large Audio-Language Models via Layer-Weighted Vector Steering

    Oct 14, 2025Tsung-En Lin, Kuan-Yi Lee, Hung-Yi LeeLarge Audio Language ModelsSteering