cs.SDSep 24, 2026

Broadening Uncertainty Estimation for Audio Question Answering Across Methods, Formats, and Inputs

Authors: Aaron Isidore Grace, Weiran Wang

Organizations: David R. Cheriton School of Computer Science, University of Waterloo, Waterloo, ON, Canada · Department of Computer Science, University of Iowa, Iowa City, IA, USA

Abstract

Audio-language models can produce confident answers unsupported by the audio, motivating uncertainty estimates that identify unreliable responses. We compare probability-based, sampling-based, self-verification, evidential, and contrastive measures across four open-weight models and five audio QA benchmarks. In multiple-choice evaluation, first-token measures are strongest overall, with top-1 probability achieving a mean AUROC of .740, compared with .708 for ten-sample discrete semantic entropy, while requiring no additional model calls. Across four benchmarks, shifting from multiple-choice to open-ended evaluation lowers mean accuracy from 57.6% to 36.6%, yet uncertainty remains predictive of errors: semantic entropy, maximum token entropy, and semantic agreement achieve mean AUROCs of .697, .694, and .693, respectively. To test whether uncertainty reflects the evidence available to answer the question, we perform input ablations that remove either the audio or the question. Across top-1 confidence, entropy, and sampling-based measures, removing audio reduces error-detection AUROC by .101 on average, compared with .010 when removing the question. Together, these results establish efficient uncertainty baselines and show that uncertainty in audio-language models depends substantially more on available audio evidence than on question text.

Figures & tables

Explore similar work

CardsList
  1. Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models

    Apr 28, 2026Chun-Yi Kuan, Wei-Ping Huang, Hung-yi LeeLarge Language Model UncertaintyLarge Audio Language Models

  2. ORCA: Open-ended Response Correctness Assessment for Audio Question Answering

    Nov 28, 2025Šimon Sedláček, Sara Barahona, Bolaji Yusuf +9Large Audio Language ModelsOpen-Ended Responses