eess.ASSep 29, 2026

Louder, Longer, Livelier: Acoustic Shortcuts and Underspecified Rationales in Speech LLM Judges

Authors: Mingyue Huo, Shivam Mehta, Bhavin Jawade, Yinghong Lan, Haoqi Li

Organizations: University of Illinois Urbana-Champaign · Netflix

Abstract

LLM-as-a-judge is widely used for evaluating text, but extending this paradigm to speech requires models to interpret acoustic as well as linguistic evidence. This introduces a modality-specific risk: a speech judge may treat a perceptually salient cue as evidence of quality even when that cue is irrelevant to the target criterion or receives more weight than human listeners give it. We call this behavior an acoustic shortcut. To study it, we audit six speech LLM judges using controlled manipulations of intensity, content richness, and emotional delivery. We evaluate both pointwise scoring and pairwise comparison, using human preference calibration to interpret the results. The judges consistently reward louder audio, prefer content-rich speech more strongly than human listeners do, and map emotional delivery into quality preferences. These effects are most visible in pairwise comparison, while pointwise scores often obscure them. More concerningly, the accompanying rationales rarely identify the acoustic cue that changes a judgment and instead repeatedly rely on a limited vocabulary, leaving them acoustically underspecified. Together, these findings show that reliable speech judges must both resist acoustic shortcuts and ground their rationales in the acoustic evidence behind their decisions. To support reproducibility and future audits, we also release SpeechJudgeAudit, the controlled stimuli and evaluation tools used in this study.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation

    Jul 15, 2026Joonyong Park, David M. Chan, Yuki Saito +1Large Audio Language ModelsLarge Language Model Judges

  2. SpeechCritic: Learning a Diagnostic Speech Judge from Limited Human Preferences

    Sep 28, 2026Mingyue Huo, Shivam Mehta, Bhavin Jawade +2Large Audio Language ModelsCritic

  3. ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge

    Jun 23, 2026Jisu Jeon, Seungyeon Jwa, Joosung Lee +6Paralinguistic CuesLarge Audio Language Models