SpeechCritic: Learning a Diagnostic Speech Judge from Limited Human Preferences
Organizations: University of Illinois Urbana-Champaign · Netflix
Abstract
Human speech conveys rich perceptual information, such as emotion and speaker identity, yet most automatic speech quality judges reduce it to a single naturalness score. We study diagnostic speech judges: given two candidates, a diagnostic judge decides which is better, along which perceptual dimensions (e.g., timbre, emotion, timing) they differ, and which audible cues support its decision. Learning such judges is challenging: expert annotation is costly, and simply prompting a frontier audio-language model to produce labels is unreliable: our probing reveals substantial errors and unstable instruction following. We introduce SpeechCritic, which learns a diagnostic judge in a reference-conditioned cross-lingual setting from only about 300 human-labeled comparisons. Rather than replacing the frontier model, SpeechCritic calibrates it with these labels: for each dimension, it selects the acoustic measurements that agree with human judgments, maps them to A/Tie/B probabilities, and passes these to the model as non-binding hints alongside the audio. Compared with the same model labeling without hints, this raises dimension-level agreement with humans by 6.3 points and cuts the mismatch with human Tie rates by 10.4 points. We then train a 7B judge on this supervision and find that different training signals shape different judge behaviors: SFT establishes the task, OPD transfers the teacher's dimension-level strengths and weaknesses, and RL helps most on clear-cut comparisons where human raters agree. Notably, human listeners also find that RL makes rationales cite more specific, localized acoustic cues, although it never directly rewards rationale text. Finally, we show that the pipeline is language-pair agnostic by instantiating it on both English-Japanese and English-Spanish. Together, these results demonstrate a path from limited human preferences to a diagnostic speech judge.
Figures & tables
| Dimension | Speaker | Emotion | Timing | Pronunciation | Artifacts |
|---|---|---|---|---|---|
| Retained metric(s) | Speaker similarity | Arousal mismatch | Duration deviation + envelope DTW | Character error rate or Language ID ∗ | None |
| OOF macro-F1 (%) | No improv. |
| Dev | Test | |||||
|---|---|---|---|---|---|---|
| Strategy | Overall Acc. % | Dim. Macro-F1 % | Tie MAE ∗ | Overall Acc. % | Dim. Macro-F1 % | Tie MAE |
| Direct labeling | 72.6 | 44.1 | 29.0 | 71.3 | 42.3 | 35.2 |
| Calibrated hints | 75.8 | 50.9 | 22.1 | 69.9 | 48.5 | 24.8 |
| Configuration | Overall Acc. (%) | High-Consensus Acc. (%) | Low-Consensus Acc. (%) | Dim. Macro-F1 (%) |
|---|---|---|---|---|
| Baseline | ||||
| Qwen2.5-Omni-7B (zero-shot) | 13.10 | 13.53 | 12.50 | 28.76 |
| A. SFT | ||||
| SFT | 69.54\,{\color[rgb]{0.35,0.35,0.35}\pm 0.80} | 81.76\,{\color[rgb]{0.35,0.35,0.35}\pm 1.02} | 52.22\,{\color[rgb]{0.35,0.35,0.35}\pm 3.37} | 52.37\,{\color[rgb]{0.35,0.35,0.35}\pm 0.10} |
| B. OPD: teacher choice | ||||
| Privileged: SFT teacher | 71.72\,{\color[rgb]{0.35,0.35,0.35}\pm 1.58} | 85.10\,{\color[rgb]{0.35,0.35,0.35}\pm 1.89} | 52.78\,{\color[rgb]{0.35,0.35,0.35}\pm 3.85} | 52.15\,{\color[rgb]{0.35,0.35,0.35}\pm 1.26} |
| Training | SFT data (comparisons) | JA Overall Acc. (%) | JA Dim. Macro-F1 (%) | ES Overall Acc. (%) | ES Dim. Macro-F1 (%) |
|---|---|---|---|---|---|
| JA only | 7,840 | 69.54\,{\color[rgb]{0.35,0.35,0.35}\pm 0.80} | 52.37\,{\color[rgb]{0.35,0.35,0.35}\pm 0.10} | 65.91\,{\color[rgb]{0.35,0.35,0.35}\pm 2.17} | 48.23\,{\color[rgb]{0.35,0.35,0.35}\pm 2.03} |
| ES only | 8,000 | 70.92\,{\color[rgb]{0.35,0.35,0.35}\pm 0.80} | 50.52\,{\color[rgb]{0.35,0.35,0.35}\pm 1.17} | 66.02\,{\color[rgb]{0.35,0.35,0.35}\pm 1.32} | 50.00\,{\color[rgb]{0.35,0.35,0.35}\pm 0.89} |
| JA+ES combined | 15,840 | 71.38\,{\color[rgb]{0.35,0.35,0.35}\pm 1.72} | 50.10\,{\color[rgb]{0.35,0.35,0.35}\pm 1.30} | 66.67\,{\color[rgb]{0.35,0.35,0.35}\pm 1.35} | 49.75\,{\color[rgb]{0.35,0.35,0.35}\pm 0.56} |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| English–Japanese | English–Spanish | |||
|---|---|---|---|---|
| Dev | Test | Dev | Test | |
| Comparisons | 315 | 290 | 307 | 307 |
| High-consensus | 180 | 160 | 139 | 136 |
| Low-consensus | 122 | 116 | 153 | 156 |
| Dimension | Human-majority Tie (%) | Direct Gemini Tie (%) | Calibrated Gemini Tie (%) | OPD judge Tie (%) |
|---|---|---|---|---|
| D1: Speaker | 52.03 | 9.59 | 18.08 | 34.32 |
| D2: Emotion | 25.38 | 3.08 | 4.62 | 4.36 |
| D3: Timing | 39.63 | 7.41 | 12.59 | 4.57 |
| D4: Pronunciation and accent | 70.42 | 21.48 | 46.83 | 72.89 |
| D5: Audio artifacts | 87.89 | 57.44 | 64.71 | 87.77 |
| Dimension | Metric implementation | Human validation and use |
| Speaker | WeSpeaker w2vbert2_mfa embedding cosine to the target-language reference ( Wang et al., 2023 ) | Retained. OOF macro-F1 improves from 20.9 to 29.1. |
| The same WeSpeaker embedding cosine to the reference | Rejected as the calibrated feature; retained only as a sanity check. | |
| Emotion | Arousal mismatch from audEERING wav2vec2 MSP-DIM ( Wagner et al., 2023 ) | Retained. Strongest screened emotion signal; OOF macro-F1 improves from 19.2 to 36.0. |
| Valence mismatch from the same MSP-DIM model | Rejected; weaker than arousal in feature screening. | |
| emotion2vec+ large embedding similarity ( Ma et al., 2024 ) | Rejected; unreliable on target-language speech in expert listening. | |
| F0/pitch-contour match | Rejected; no improvement over arousal in a 30-pair expert pilot. |
| Configuration | Eval. | Runs | Overall Acc. (%) | High-Consensus Acc. (%) | Low-Consensus Acc. (%) | Dim. Macro-F1 (%) |
| A. Zero-shot audio-language models | ||||||
| Gemini-2.5-Pro | JA | 1 | 62.41 | 67.65 | 55.00 | 48.58 |
| Gemini-3.5-Flash | JA | 1 | 67.59 | 74.12 | 58.33 | 39.67 |
| Qwen3-Omni-30B | JA | 1 | 57.24 | 61.18 | 51.67 | 41.75 |
| Qwen2.5-Omni-7B | JA | 1 | 13.10 | 13.53 | 12.50 | 28.76 |
| Step-Audio-2-mini | JA | 1 | 49.66 | 50.00 | 49.17 | 13.58 |
| System | D1 Voice | D2 Emotion | D3 Naturalness ∗ | D4 Pronun. | D5 Audio | Mean |
|---|---|---|---|---|---|---|
| Human upper reference | 72.3 | 78.4 | 75.8 | 77.8 | 79.7 | 76.8 |
| Human lower reference | 50.0 | 56.4 | 49.0 | 64.3 | 65.5 | 57.0 |
| Always Tie | 29.5 | 11.4 | 22.6 | 25.9 | 19.9 | 21.8 |
| Random | 25.8 | 31.1 | 31.9 | 29.6 | 32.5 | 30.2 |
| Gemini 3.1 Pro | 26.1 | 45.1 | 35.9 | 45.0 | 48.2 | 40.0 |
| EN–JA audio-adapted SFT | 19.8 | 38.9 | 36.2 | 51.8 | 22.5 | 33.8 |
| Models | Overall | Communication | Grounding | Consistency | Dim. hygiene |
|---|---|---|---|---|---|
| SFT vs. SFT+RL | 3/ 11 /14 | 2/3/23 | 1/ 11 /16 | 1/ 6 /21 | 2/5/21 |
| SFT vs. Gemini 3.1 Pro | 10 /7/9 | 3/6/17 | 6/6/14 | 1/1/24 | 9 /1/16 |
| SFT vs. audio-adapted SFT | 10 /6/13 | 1/3/25 | 6 /3/20 | 2/1/26 | 5 /4/20 |
| OPD (Qwen3 teacher) vs. OPD (SFT teacher) | 15 /5/5 | 5 /4/16 | 13 /2/10 | 5 /4/16 | 6 /5/14 |