MISHAP-Bench: A Hallucination Benchmark for Large Audio-Language Models
Organizations: University of Neuchâtel · TU Dortmund University · RWTH Aachen University · Delft University of Technology
Abstract
Large audio-language models (LALMs) produce fluent responses about audio but often hallucinate by making plausible yet ungrounded claims. Existing audio hallucination benchmarks mainly measure response correctness, leaving it unclear whether an LALM hallucinates or simply fails to understand the audio. We challenge correctness-based evaluation by defining two hallucination categories: (i) context, where claims are not grounded in the audio; and (ii) knowledge, where claims about audio-related topics lack support from externally verifiable facts. We introduce MISHAP-Bench, a comprehensive benchmark with 12,000 challenging open-ended question-audio pairs and a rigorous evaluation pipeline covering both categories. To evaluate open-ended responses, we propose a groundedness judge that uses reference rubrics and judge prompts guided by human annotations. We extensively evaluate ten state-of-the-art LALMs and show that hallucination remains substantial. Even a frontier model such as Gemini 3.7 Flash reaches a hallucination rate of 36.5%. We further adapt and benchmark four mitigation methods from multiple domains for LALMs. Despite some improvements, effective hallucination mitigation remains an open challenge. Finally, we call on the community to evaluate hallucination and benchmark mitigation methods with MISHAP-Bench.
Figures & tables
| Music | Sound | Speech | ||||||||
| Model | Context | Knowledge | Overall | Context | Knowledge | Overall | Context | Knowledge | Overall | Overall |
| Open-Weight Model | ||||||||||
| Qwen2-Audio Instruct | 71.8 | 65.7 | 68.8 | 72.4 | 51.1 | 61.8 | 29.4 | 31.3 | 30.4 | 53.6 |
| Qwen2.5-Omni | 56.8 | 38.5 | 47.6 | 50.6 | 15.2 | 33.0 | 37.2 | 18.1 | 27.7 | 36.1 |
| Kimi-Audio Instruct | 70.5 | 62.3 | 66.4 | 71.8 | 32.2 | 52.0 | 30.9 | 35.4 | 33.1 | 50.5 |
| Audio Flamingo 3 | 77.7 | 79.2 | 78.4 | 69.4 | 46.5 | 58.0 | 48.5 | 63.8 | 56.2 | 64.2 |
| Music | Sound | Speech | ||||||||
| Method | Context | Knowledge | Overall | Context | Knowledge | Overall | Context | Knowledge | Overall | Overall |
| Qwen2.5-Omni | ||||||||||
| W/O | 56.8 | 38.5 | 47.6 | 50.6 | 15.2 | 33.0 | 37.2 | 18.1 | 27.7 | 36.1 |
| AAD | 54.2 | 35.1 | 44.6 | 64.4 | 16.1 | 40.3 | 25.2 | 20.1 | 22.7 | 35.9 |
| VCD | 63.4 | 39.3 | 51.4 | 58.9 | 16.4 | 37.7 | 29.8 | 22.1 | 25.9 | 38.3 |
| MTI | 50.3 | 36.1 | 43.2 | 50.0 | 16.4 | 33.1 | 41.8 | 19.9 | 30.9 | 35.7 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Domain | Dataset | Download Link |
| Music | Song Describer Dataset | zenodo.org/records/10072001 |
| Mridangam Stroke 1.5 | zenodo.org/records/4068196 | |
| Sound | Clotho v2.1 | zenodo.org/records/4783391 |
| FSD50K | zenodo.org/records/4060432 | |
| Speech | LibriSpeech | openslr.org/12 |
| Common Voice 22.0 | fsicoli/common_voice_22_0 |
| Model | Access | Checkpoint or Endpoint |
| Qwen2-Audio Instruct | Hugging Face | Qwen/Qwen2-Audio-7B-Instruct |
| Qwen2.5-Omni | Hugging Face | Qwen/Qwen2.5-Omni-7B |
| Kimi-Audio Instruct | Hugging Face | moonshotai/Kimi-Audio-7B-Instruct |
| Audio Flamingo 3 | Hugging Face | nvidia/audio-flamingo-3-hf |
| Step-Audio 2 mini | Hugging Face | stepfun-ai/Step-Audio-2-mini |
| MiMo-Audio Instruct | Hugging Face | XiaomiMiMo/MiMo-Audio-7B-Instruct |
| Model | Decoding | Temp. | Top- | Top- | Rep. Pen. | Max Tokens |
| Open-Weight Model | ||||||
| Qwen2-Audio Instruct | Greedy | – | – | – | 1.10 | 512 |
| Qwen2.5-Omni | Greedy | – | – | – | 1.00 | 512 |
| Kimi-Audio Instruct | Greedy | – | – | – | 1.00 | 512 |
| Audio Flamingo 3 | Sampled | 0.7 | 0.80 | 20 | 1.05 | 512 |
| Step-Audio 2 mini | Greedy | – | – | – | 1.05 | 512 |
| Method | Reference Branch | Hyperparameter | Value |
| AAD | Audio removed | Original weight | 2.0 |
| Reference weight | 1.0 | ||
| VCD | Diffused audio | Original weight | 2.0 |
| Reference weight | 1.0 | ||
| Plausibility threshold | 0.1 | ||
| Noise rate | 0.1 |
| Benchmark | Context | Knowledge | Question Format | Evaluator | Human Annot. | Mitigation | # Pairs | # Audio |
| USMQ | ✓ | ✗ | Binary; captioning | Yes / No parsing; ECHO / GPT-4 | ✗ | ✓ | 30,220 | N/R |
| CompA | ✓ | ✗ | Audio–caption matching | Similarity scores | ✗ | ✗ | 1,300 | 1,300 |
| MATCH | ✓ | ✗ | Binary | Exact label matching | ✗ | ✓ | 15,530 | N/R |
| AVHBench | ✓ | ✗ | Binary; captioning | Label / caption metrics; GPT-4 | ✗ | ✓ | 6,408 | 2,136 |
| AHa-Bench | ✓ | ✗ | Binary; transcription | GPT-4o label mapping; WER | ✗ | ✗ | 906 | 396 |
| HalluAudio | ✓ | ✗ | Binary; MCQ | Keyword / rule matching | ✗ | ✗ | 5,720 | N/R |
| Music | Sound | Speech | ||||||||
| Method | Context | Knowledge | Overall | Context | Knowledge | Overall | Context | Knowledge | Overall | Overall |
| Qwen2-Audio Instruct | ||||||||||
| W/O | 71.8 | 65.7 | 68.8 | 72.4 | 51.1 | 61.8 | 29.4 | 31.3 | 30.4 | 53.6 |
| AAD | 68.4 | 57.1 | 62.8 | 71.3 | 50.2 | 60.8 | 20.6 | 26.6 | 23.6 | 49.0 |
| VCD | 70.6 | 58.5 | 64.6 | 68.1 | 44.7 | 56.4 | 24.4 | 23.4 | 23.9 | 48.3 |
| MTI | 76.2 | 65.2 | 70.8 | 74.9 | 48.8 | 61.8 | 34.9 | 31.6 | 33.2 | 55.3 |
| Context | Knowledge | |||||||||
| Model | Contradicted Premise | Absent Premise | False Negation | Overall | Niche Specifics | Unknowable Provenance | False Authority | False Premise | Overall | Overall |
| Open-Weight Model | ||||||||||
| Qwen2-Audio Instruct | 64.9 | 57.6 | 39.3 | 57.9 | 47.7 | 32.0 | 61.2 | 59.7 | 49.4 | 53.6 |
| Qwen2.5-Omni | 49.3 | 49.6 | 44.0 | 48.2 | 26.9 | 1.4 | 19.3 | 54.6 | 23.9 | 36.1 |
| Kimi-Audio Instruct | 64.3 | 51.7 | 45.2 | 57.7 | 39.7 | 15.3 | 56.0 | 66.9 | 43.3 | 50.5 |
| Audio Flamingo 3 | 70.3 | 73.8 | 44.3 | 65.2 | 53.6 | 30.1 | 93.9 | 78.4 | 63.2 | 64.2 |
| Context | Knowledge | |||||||||
| Method | Contradicted Premise | Absent Premise | False Negation | Overall | Niche Specifics | Unknowable Provenance | False Authority | False Premise | Overall | Overall |
| Qwen2-Audio Instruct | ||||||||||
| W/O | 64.9 | 57.6 | 39.3 | 57.9 | 47.7 | 32.0 | 61.2 | 59.7 | 49.4 | 53.6 |
| AAD | 61.7 | 48.1 | 35.8 | 53.5 | 42.6 | 38.7 | 51.7 | 46.0 | 44.6 | 49.0 |
| VCD | 61.0 | 51.0 | 39.6 | 54.4 | 44.0 | 27.0 | 49.1 | 52.9 | 42.2 | 48.3 |
| MTI | 69.1 | 63.5 | 41.7 | 62.0 | 49.3 | 33.2 | 57.9 | 57.4 | 48.5 | 55.3 |