Louder, Longer, Livelier: Acoustic Shortcuts and Underspecified Rationales in Speech LLM Judges
Organizations: University of Illinois Urbana-Champaign · Netflix
Abstract
LLM-as-a-judge is widely used for evaluating text, but extending this paradigm to speech requires models to interpret acoustic as well as linguistic evidence. This introduces a modality-specific risk: a speech judge may treat a perceptually salient cue as evidence of quality even when that cue is irrelevant to the target criterion or receives more weight than human listeners give it. We call this behavior an acoustic shortcut. To study it, we audit six speech LLM judges using controlled manipulations of intensity, content richness, and emotional delivery. We evaluate both pointwise scoring and pairwise comparison, using human preference calibration to interpret the results. The judges consistently reward louder audio, prefer content-rich speech more strongly than human listeners do, and map emotional delivery into quality preferences. These effects are most visible in pairwise comparison, while pointwise scores often obscure them. More concerningly, the accompanying rationales rarely identify the acoustic cue that changes a judgment and instead repeatedly rely on a limited vocabulary, leaving them acoustically underspecified. Together, these findings show that reliable speech judges must both resist acoustic shortcuts and ground their rationales in the acoustic evidence behind their decisions. To support reproducibility and future audits, we also release SpeechJudgeAudit, the controlled stimuli and evaluation tools used in this study.
Figures & tables
| Judge | Backbone | Quality criterion | Pairwise | Pointwise | Scale | Rationale |
|---|---|---|---|---|---|---|
| SpeechJudge-BTRM ( Zhang et al., 2025 ) | Qwen2.5-Omni | Naturalness | ✓ | |||
| SpeechJudge-GRM ( Zhang et al., 2025 ) | Qwen2.5-Omni | Naturalness | ✓ | 1–10 | ✓ | |
| UniSRM ( Wang et al., 2026b ) | Qwen2.5-Omni | Overall quality | ✓ | ✓ | 1–5 | ✓ |
| SQ-LLM ( Wang et al., 2026a ) | Qwen2.5-Omni | Overall quality | ✓ | ✓ | 1–5 | ✓ |
| Gemini-2.5-pro ( Gemini Team, Google, 2025 ) | Gemini | Overall quality | ✓ | ✓ | 1–10 | ✓ |
| Gemini-3.1-pro ( Gemini Team, 2026 ) | Gemini | Overall quality | ✓ | ✓ | 1–10 | ✓ |
| Comparison | What changes | Mean preference |
|---|---|---|
| Silence padding vs. concise | Duration only | 0.41 |
| Repetition vs. concise | Duplicated speech | 0.46 |
| Richer vs. concise | Content and varied speech material | 0.71 |
| Richer vs. duration-matched padding | Content richness at matched duration | 0.70 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| System | Generation structure | Representation |
|---|---|---|
| CosyVoice2 | AR LLM + flow matching | Codec tokens |
| Qwen3-TTS | AR LLM + codec decoder | Codec tokens |
| F5-TTS | Non-AR DiT flow matching | Mel-spectrogram |
| StyleTTS2 | Non-AR diffusion + GAN | Mel-spectrogram |
| Judge | Condition | Pairwise N | Pairwise pref. | Pairwise sig. | Pointwise N | Pointwise | Pointwise sig. | Notes |
|---|---|---|---|---|---|---|---|---|
| BTRM | dB vs. ref | – | – | – | 100 | -0.54 | ** | quiet penalty |
| BTRM | dB vs. ref | – | – | – | 99 | -0.42 | ** | quiet penalty |
| BTRM | dB vs. ref | – | – | – | 99 | +0.05 | n.s. | |
| BTRM | dB vs. ref | – | – | – | 99 | -0.12 | n.s. | |
| GRM | dB vs. ref | 80 | 0.500 | n.s. | – | – | – | pairwise-only judge |
| GRM | dB vs. ref | 80 | 0.475 | n.s. | – | – | – |
| Judge | Condition | Pairwise N | Pairwise pref. | Pairwise sig. | Pointwise N | Pointwise | Pointwise sig. | Notes |
|---|---|---|---|---|---|---|---|---|
| BTRM | pad500 vs. base | – | – | – | 80 | -1.59 | pad penalty | |
| BTRM | pad1000 vs. base | – | – | – | 80 | -2.87 | pad penalty grows | |
| UniSRM | pad500 vs. base | 80 | 0.250 | 80 | -0.01 | n.s. | pointwise 84% identical | |
| UniSRM | pad1000 vs. base | 80 | 0.069 | 80 | +0.01 | n.s. | pointwise 86% identical | |
| SQ-LLM | pad500 vs. base | 80 | 0.537 | n.s. | 80 | +0.01 | n.s. | 96% identical scores |
| SQ-LLM | pad1000 vs. base | 80 | 0.422 | n.s. | 80 | -0.04 | n.s. | trend anti-pad |
| Setting | Richer preference | Takeaway |
|---|---|---|
| Baseline | 0.701 | reference |
| Position/length instruction | 0.695 | preference remains |
| Judge | Richer concise | Sig. | Pad concise | Sig. | Repeat concise | Sig. | Richer pad | Sig. |
|---|---|---|---|---|---|---|---|---|
| UniSRM | 0.846 | 0.082 | 0.447 | ** | 0.954 | |||
| GRM | 0.570 | 0.527 | n.s. | 0.362 | 0.567 | |||
| SQ-LLM | 0.755 | 0.424 | 0.515 | n.s. | 0.765 | |||
| Gemini-2.5 | 0.664 | 0.507 | n.s. | 0.515 | n.s. | 0.544 | ** | |
| Gemini-3.1 | 0.701 | 0.502 | n.s. | 0.471 | 0.661 |
| Judge | Setting | Text fidelity? | Richer pref. | What it shows |
|---|---|---|---|---|
| BTRM | Concise target transcript supplied | yes | 0.705 | target transcript changes the outcome |
| BTRM | Richer target transcript supplied | yes | 0.262 | target mismatch reverses the direction |
| BTRM | Concise and richer targets averaged | yes | 0.483 | target effects cancel; no target-free analogue |
| UniSRM | Target transcript plus text-fidelity dimension | yes | – | text fidelity masks content-richness preference |
| UniSRM | Target-free prompt variant | no | – | output became template-like, so excluded |
| Gemini-3.1 | Anti-position/length instruction only | no | 0.695 | preference remains despite the instruction |
| Judge | Emotion | Pairwise N | Pairwise pref. | Pairwise sig. | Pointwise | Pointwise sig. | Overall sig. | Notes |
|---|---|---|---|---|---|---|---|---|
| BTRM | Happy | – | – | – | -0.085 | n.s. | ** | |
| BTRM | Sad | – | – | – | -0.205 | ** | lower than Neutral | |
| BTRM | Angry | – | – | – | -0.129 | ** | lower than Neutral | |
| GRM | Happy | 400 | 0.478 | n.s. | – | – | – | flat |
| GRM | Sad | 400 | 0.392 | – | – | – | Neutral wins | |
| GRM | Angry | 400 | 0.395 | – | – | – | Neutral wins |
| Cue / example | Pointwise evidence | Pairwise evidence | Takeaway |
|---|---|---|---|
| Content richness | no consistent richer-speech advantage; Gemini-2.5 leans concise | richer preference 0.570–0.846 | pairwise comparison reveals a consistent preference |
| Content richness correlations | Gemini-3.1 ; SQ-LLM | same items, pairwise preferences | per-item pointwise and pairwise verdicts are weakly related |
| Silence / UniSRM | , 84–86% identical scores | pad1000 preference 0.069 | integer scale hides anti-pad preference |
| Loudness / SQ-LLM | loud6 , degenerate | loud6 preference 0.591 | pointwise resolution ceiling hides pro-loud preference |
| Loudness / Gemini | all pointwise deltas n.s. | quiet6 loses, loud6 wins | pointwise noise hides directional level preference |
| Emotional delivery | Kendall across pointwise ranks | Kendall across pairwise ranks | pairwise rankings are more consistent across judges |
| Setting | Position effect observed | Role in this audit |
|---|---|---|
| Content richness | GRM strongly favors the second clip in decisive comparisons (B wins 91.6%); Gemini-2.5 also favors B (71.8% of trials), while UniSRM and SQ-LLM lean toward A (64.8% and 56.4%). | We average the two presentation orders before computing the richer-speech preference, so the reported effect is not a single-order artifact. |
| Loudness | GRM shows order asymmetries of roughly 10–15 percentage points across loudness conditions even when its order-averaged loudness preference is near chance. | Order averaging prevents a positional preference from being interpreted as evidence for or against a loudness preference. |
| Silence padding | GRM and SQ-LLM show clear AB/BA differences, while their order-averaged padded-clip preferences remain near chance. | This supports using order-averaged item preferences and treating the strongest silence result as judge-specific rather than universal. |
| Emotional delivery | The average AB–BA gap is about 0.10 for SQ-LLM, 0.48 for GRM, 0.61 for Gemini-2.5, 0.68 for Gemini-3.1, and 0.80 for UniSRM. | The emotion preferences in Table 9 are interpretable because they are averaged across both orders. |
| Manipulation | Judge(s) audited | Explicit acoustic-term rate | Generic rubric terms |
|---|---|---|---|
| Content richness / length | Gemini-3.1, GRM, UniSRM | length 0–3%; richer/original 0% | rubric/prosody terms frequent |
| Loudness | UniSRM | loud/loudness/volume 0%; amplitude/dB/level 0.94% | prosody/naturalness/quality 100% |
| Silence padding | UniSRM | silence 0%; duration 0%; pad/leading/trailing 1.25% | prosody/naturalness/quality 100% |
| Emotion | rationale-producing judges | emotion-specific terms rarely identify the true emotion (3/16 pointwise cells) | generic prosody terms excluded from lexicon |
| Repetition | GRM | repetition keywords 22.1% for rep-vs-concise | partial exception; detection incomplete |
| Judge | Mentions rep. | ||
|---|---|---|---|
| Gemini-2.5 | 5.7% | 11.8% | 13.9% |
| Gemini-3.1 | 27.9% | 16.9% | 2.3% |
| GRM | 21.8% | 78.5% | 59.7% |
| UniSRM | 10.5% | 80.6% | 45.1% |
| Condition | Items | N | Human mean [CI] | Human sig. | Judge mean [range] | Sig. judges | Residual | Human entropy | Human variance |
|---|---|---|---|---|---|---|---|---|---|
| +6 dB vs. ref | 12 | 88 | 0.612 [0.548, 0.671] | 0.578 [0.541, 0.616] | 5/5 | -0.035 | 1.009 | 0.071 | |
| Richer vs. concise | 18 | 134 | 0.547 [0.464, 0.627] | n.s. | 0.707 [0.570, 0.846] | 5/5 | +0.160 | 1.260 | 0.111 |
| Sad vs. Neutral | 15 | 115 | 0.217 [0.129, 0.320] | 0.403 [0.328, 0.555] | 5/5 | +0.186 | 0.756 | 0.098 | |
| Happy vs. Neutral | 15 | 112 | 0.407 [0.326, 0.488] | n.s. | 0.531 [0.470, 0.598] | 3/5 | +0.124 | 1.305 | 0.155 |
| Rater | Valid N | Mean pref. | Tie rate | Majority agreement | Item-mean MAE |
|---|---|---|---|---|---|
| Lowest-agreement rater (rater_14) | 29 | 0.379 | 0.207 | 0.345 | 0.478 |
| Most extreme rater (rater_11) | 30 | 0.450 | 0.100 | 0.433 | 0.465 |
| All raters mean | 29.9 | 0.441 | 0.343 | 0.481 | 0.311 |
| Analysis | Statistic | Result |
|---|---|---|
| Speaker gender | female vs. male speakers | no preference (n.s.; all ) |
| Speaker accent | North-American vs. British speakers | no preference (n.s.) |
| Emotion gender interaction | gender emotion | no interaction (n.s.) |
| Low-level descriptors | descriptor–score correlations | no corrected correlation (n.s.) |
| Model self-preference | Qwen judges on Qwen3-TTS vs. other TTS | no self-preference (n.s.) |
| Gemini TTS preference | TTS-system effect | preference differs by system (***) |