Human speech conveys rich perceptual information, such as emotion and speaker identity, yet most automatic speech quality judges reduce it to a single naturalness score. We study diagnostic speech judges: given two candidates, a diagnostic judge decides which is better, along which perceptual dimensions (e.g., timbre, emotion, timing) they differ, and which audible cues support its decision. Learning such judges is challenging: expert annotation is costly, and simply prompting a frontier audio-language model to produce labels is unreliable: our probing reveals substantial errors and unstable instruction following. We introduce SpeechCritic, which learns a diagnostic judge in a reference-conditioned cross-lingual setting from only about 300 human-labeled comparisons. Rather than replacing the frontier model, SpeechCritic calibrates it with these labels: for each dimension, it selects the acoustic measurements that agree with human judgments, maps them to A/Tie/B probabilities, and passes these to the model as non-binding hints alongside the audio. Compared with the same model labeling without hints, this raises dimension-level agreement with humans by 6.3 points and cuts the mismatch with human Tie rates by 10.4 points. We then train a 7B judge on this supervision and find that different training signals shape different judge behaviors: SFT establishes the task, OPD transfers the teacher's dimension-level strengths and weaknesses, and RL helps most on clear-cut comparisons where human raters agree. Notably, human listeners also find that RL makes rationales cite more specific, localized acoustic cues, although it never directly rewards rationale text. Finally, we show that the pipeline is language-pair agnostic by instantiating it on both English-Japanese and English-Spanish. Together, these results demonstrate a path from limited human preferences to a diagnostic speech judge.
Figures & tables
Figure 1: In a reference-conditioned cross-lingual comparison, a diagnostic speech judge produces multidimensional verdicts and evidence-grounded rationales that identify audible differences between two candidates. Our SpeechCritic framework covers the end-to-end data, training, and evaluation pipeline. The central human-calibration method is introduced in Section 4 .
Figure 2: Zero-shot probing of three proprietary and five open-source audio-language models on our constructed English–Japanese test set (conducted only after all model choices were fixed.) The results reveal a clear gap to reliable human-aligned diagnostic judging.
Figure 3: Human-calibrated domain hints. Direct machine labeling exhibits systematic disagreement with human. Our select–calibrate–scale pipeline uses limited human judgments to convert reliable domain metrics into probabilistic hints for scalable and acoustically grounded supervision.
Dimension
Speaker
Emotion
Timing
Pronunciation
Artifacts
Retained metric(s)
Speaker similarity
Arousal mismatch
Duration deviation + envelope DTW
Character error rate or Language ID ∗
None
OOF macro-F1 (%)
20.9→29.1
19.2→36.0
18.6→54.7
26.7→50.6
No improv.
Table 1: Retained domain metrics and grouped OOF macro-F1 on the human-labeled dev set.
Dev
Test
Strategy
Overall Acc. % ↑
Dim. Macro-F1 % ↑
Tie MAE ∗ ↓
Overall Acc. % ↑
Dim. Macro-F1 % ↑
Tie MAE ↓
Direct labeling
72.6
44.1
29.0
71.3
42.3
35.2
+ Calibrated hints
75.8
50.9
22.1
69.9
48.5
24.8
Table 2: Human-calibrated domain hints improve the dimensional alignment and Tie -rate calibration of Gemini 3.1 Pro as a machine labeler, with no detectable change in Overall accuracy.
Configuration
Overall Acc. (%)
High-Consensus Acc. (%)
Low-Consensus Acc. (%)
Dim. Macro-F1 (%)
Baseline
Qwen2.5-Omni-7B (zero-shot)
13.10
13.53
12.50
28.76
A. SFT
SFT
69.54\,{\color[rgb]{0.35,0.35,0.35}\pm 0.80}
81.76\,{\color[rgb]{0.35,0.35,0.35}\pm 1.02}
52.22\,{\color[rgb]{0.35,0.35,0.35}\pm 3.37}
52.37\,{\color[rgb]{0.35,0.35,0.35}\pm 0.10}
B. OPD: teacher choice
Privileged: SFT teacher
71.72\,{\color[rgb]{0.35,0.35,0.35}\pm 1.58}
85.10\,{\color[rgb]{0.35,0.35,0.35}\pm 1.89}
52.78\,{\color[rgb]{0.35,0.35,0.35}\pm 3.85}
52.15\,{\color[rgb]{0.35,0.35,0.35}\pm 1.26}
Table 3: Main results: SFT, OPD, and RL produce distinct training signals to shape a diagnostic speech judge, rather than a uniform accuracy ladder.
Figure 7
Figure 5: RL can improve rationale grounding without direct rationale rewards. Given identical verdicts, SFT+RL cites more specific and better-localized acoustic evidence than SFT.
Training
SFT data (comparisons)
JA Overall Acc. (%)
JA Dim. Macro-F1 (%)
ES Overall Acc. (%)
ES Dim. Macro-F1 (%)
JA only
7,840
69.54\,{\color[rgb]{0.35,0.35,0.35}\pm 0.80}
52.37\,{\color[rgb]{0.35,0.35,0.35}\pm 0.10}
65.91\,{\color[rgb]{0.35,0.35,0.35}\pm 2.17}
48.23\,{\color[rgb]{0.35,0.35,0.35}\pm 2.03}
ES only
8,000
70.92\,{\color[rgb]{0.35,0.35,0.35}\pm 0.80}
50.52\,{\color[rgb]{0.35,0.35,0.35}\pm 1.17}
66.02\,{\color[rgb]{0.35,0.35,0.35}\pm 1.32}
50.00\,{\color[rgb]{0.35,0.35,0.35}\pm 0.89}
JA+ES combined
15,840
71.38\,{\color[rgb]{0.35,0.35,0.35}\pm 1.72}
50.10\,{\color[rgb]{0.35,0.35,0.35}\pm 1.30}
66.67\,{\color[rgb]{0.35,0.35,0.35}\pm 1.35}
49.75\,{\color[rgb]{0.35,0.35,0.35}\pm 0.56}
Table 5: Cross-lingual transfer of diagnostic speech judges. All models adapt only the LLM with LoRA. Monolingual judges transfer strongly across target languages, as highlighted in red.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
English–Japanese
English–Spanish
Dev
Test
Dev
Test
Comparisons
315
290
307
307
High-consensus
180
160
139
136
Low-consensus
122
116
153
156
Appendix
Table 6: Human benchmark used in our experiments. Counts include only development and test comparisons with a strict majority for Overall. Consensus counts are computed among standard three-rater comparisons: high-consensus is 3 – 0 and low-consensus is 2 – 1 .
Dimension
Human-majority Tie (%)
Direct Gemini Tie (%)
Calibrated Gemini Tie (%)
OPD judge Tie (%)
D1: Speaker
52.03
9.59
18.08
34.32
D2: Emotion
25.38
3.08
4.62
4.36
D3: Timing
39.63
7.41
12.59
4.57
D4: Pronunciation and accent
70.42
21.48
46.83
72.89
D5: Audio artifacts
87.89
57.44
64.71
87.77
Appendix
Table 7: Dimensional Tie prevalence on our constructed English–Japanese test set. Human values are the percentages of comparisons whose panel-majority label is Tie ; model values are the percentages predicted as Tie . The OPD judge uses an SFT-trained teacher; its results are averaged over three seeds.
Dimension
Metric implementation
Human validation and use
Speaker
WeSpeaker w2vbert2_mfa embedding cosine to the target-language reference ( Wang et al., 2023 )
Retained. OOF macro-F1 improves from 20.9 to 29.1.
The same WeSpeaker embedding cosine to the reference
Rejected as the calibrated feature; retained only as a sanity check.
Emotion
Arousal mismatch from audEERING wav2vec2 MSP-DIM ( Wagner et al., 2023 )
Retained. Strongest screened emotion signal; OOF macro-F1 improves from 19.2 to 36.0.
Valence mismatch from the same MSP-DIM model
Rejected; weaker than arousal in feature screening.
emotion2vec+ large embedding similarity ( Ma et al., 2024 )
Rejected; unreliable on target-language speech in expert listening.
F0/pitch-contour match
Rejected; no improvement over arousal in a 30-pair expert pilot.
Appendix
Table 8: Candidate domain metrics, human validation, and use in the final soft hints. OOF denotes grouped out-of-fold evaluation against human development-set verdicts; pilot listening studies were conducted by linguist experts. Retained metrics enter the probability mapping, auxiliary metrics support construction or auditing, and rejected metrics are not exposed to the machine labeler.
Configuration
Eval.
Runs
Overall Acc. (%)
High-Consensus Acc. (%)
Low-Consensus Acc. (%)
Dim. Macro-F1 (%)
A. Zero-shot audio-language models
Gemini-2.5-Pro
JA
1
62.41
67.65
55.00
48.58
Gemini-3.5-Flash
JA
1
67.59
74.12
58.33
39.67
Qwen3-Omni-30B
JA
1
57.24
61.18
51.67
41.75
Qwen2.5-Omni-7B
JA
1
13.10
13.53
12.50
28.76
Step-Audio-2-mini
JA
1
49.66
50.00
49.17
13.58
Appendix
Table 9: Additional experimental results and controls. Unless marked ES, results use the English–Japanese test set. Three-run results are mean ± standard deviation across training seeds; single-run results are zero-shot evaluations or targeted ablations. Trained checkpoints are selected by development-set Overall accuracy.
System
D1 Voice
D2 Emotion
D3 Naturalness ∗
D4 Pronun.
D5 Audio
Mean
Human upper reference
72.3
78.4
75.8
77.8
79.7
76.8
Human lower reference
50.0
56.4
49.0
64.3
65.5
57.0
Always Tie
29.5
11.4
22.6
25.9
19.9
21.8
Random
25.8
31.1
31.9
29.6
32.5
30.2
Gemini 3.1 Pro
26.1
45.1
35.9
45.0
48.2
40.0
EN–JA audio-adapted SFT
19.8
38.9
36.2
51.8
22.5
33.8
Appendix
Table 10: External evaluation on English-to-Spanish VOX-DUB. Values are dimensional macro-F1 (%); D3 (Naturalness) is an approximate match to our Timing dimension.
Figure 15
Models
Overall
Communication
Grounding
Consistency
Dim. hygiene
SFT vs. SFT+RL
3/ 11 /14
2/3/23
1/ 11 /16
1/ 6 /21
2/5/21
SFT vs. Gemini 3.1 Pro
10 /7/9
3/6/17
6/6/14
1/1/24
9 /1/16
SFT vs. audio-adapted SFT
10 /6/13
1/3/25
6 /3/20
2/1/26
5 /4/20
OPD (Qwen3 teacher) vs. OPD (SFT teacher)
15 /5/5
5 /4/16
13 /2/10
5 /4/16
6 /5/14
Appendix
Table 11: Human pairwise rationale evaluation. Each cell is first model better / second model better / no difference ; all paired models predict the same Overall verdict.
Graduate School of Information Science and Technology, The University of Tokyo, Tokyo, Japan · Berkeley AI Research (BAIR), University of California, Berkeley, CA, USA