Non-intrusive MOS predictors are widely used instead of subjective listening tests to evaluate and rank speech enhancement (SE) systems. If they accurately reflect perceived quality, raising their scores should lead to higher-quality speech. We present the first comprehensive analysis of test-time optimization for the SE task, which directly modifies the enhanced signal to raise the average of multiple MOS predictor scores. On seven systems from the URGENT 2026 challenge, we find that 1)~all the optimized predicted scores increase while reference-based metrics remain nearly unchanged, 2)~a non-optimized predicted score does not increase, and 3)~a MUSHRA listening test shows no improvement in perceived quality. These findings reveal a risk that such optimization can distort evaluations, e.g., biasing comparisons of SE systems regardless of their perceived quality. We believe these findings can inform future evaluation practices: they suggest that predictors used for optimization should not be used for evaluation, and that challenges should keep the predictors used for ranking undisclosed.
Figures & tables
Non-intrusive (Optimized)
Unseen
Intrusive
Task-indep.
Task-dependent
Rank
System
λ
DNSMOS
NISQA
UTMOS
SCOREQ
Distill
PESQ
ESTOI
SBERT
LPS
SpkSim
EmoSim
LID
CAcc
non-int.
overall
WR
−
3.06
3.37
2.77
3.67
3.68
2.81
0.86
0.89
0.83
0.76
0.99
0.96
88.16
3
1
0.005
+0.24
+1.19
+0.15
+0.12
+0.02
0.00
0.00
0.00
0.00
0.00
0.00
0.00
+0.05
3 → 1
1 → 1
0.05
+0.50
+1.77
+0.54
+0.28
+0.01
−0.06
−0.01
−0.01
0.00
0.00
0.00
0.00
+0.43
3 → 1
1 → 1
subatom.
−
2.96
3.38
2.41
3.32
3.36
2.71
0.84
0.88
0.79
0.76
0.99
0.95
87.76
6
2
0.005
+0.49
+1.21
+0.21
+0.19
+0.03
−0.01
0.00
−0.01
0.00
0.00
0.00
0.00
−0.66
6 → 2
2 → 2
Table 1: Scores and ranks of the enhanced signal x^ (shaded rows), and score differences and rank changes between the modified signal x~ and x^ . The enhanced signal is jointly optimized for the four predictors, except in the rows marked (S), where it is optimized for DNSMOS alone.
Comparison
Δ MUSHRA
95% CI
p
Equiv. bound
RICK − baird
+8.48
[+2.5,+14.5]
0.011
13.4
baird, λ=0.005
−0.22
[−2.2,+1.7]
0.81
1.8
baird, λ=0.05
−1.20
[−5.0,+2.6]
0.50
4.3
RICK, λ=0.005
−0.65
[−2.6,+1.3]
0.47
2.2
Table 2: Differences in the MUSHRA scores, with a 95% confidence interval and the smallest equivalence bound at the 5% level.
We investigate whether predicted Mean Opinion Scores (MOS) can reliably support system-level comparisons of speech enhancement (SE) methods by introducing system-level preference accuracy (SPA). Although MOS prediction models are widely used to evaluate SE systems, their performance is typically assessed by correlation with human-rated MOS, which does not guarantee agreement on which system is better. SPA addresses this gap by directly evaluating whether predicted and human-rated MOS yield the same system preferences. Using SPA, we systematically evaluate three settings: single prediction models, ensembling, and domain adaptation. Through experiments, SPA varies substantially across single prediction models, from 9.4% to 76.8%. Even the best model disagrees with human judgments in approximately 23% of system comparisons. Ensembling yields only limited improvement, while domain adaptation tends to substantially improve SPA in the closed condition but brings only modest gains in the more practical open condition, where neither the target systems nor the speakers are known. These results suggest that SPA can reveal errors correlation-based evaluation alone does not expose, and that predicted MOS alone can lead to unreliable conclusions in practical SE system comparison.
Mean opinion scores (MOS) are widely used for speech quality assessment, yet scalar labels are sensitive to rater variability and listening test differences. This introduces labeling noise, which limits the reliability of MOS prediction. Preference prediction reduces this variability as listeners compare signals directly, producing cleaner labels. We study MOS-free preference prediction and propose PrefSQA, which incorporates uncertainty-aware logits, an impairment attention head, and a module based on non-matching-reference comparisons. We use and refine five datasets, including MOS-derived and low-noise simulated sets with matching and non-matching content, experiment with human preference sets, and test on unseen data. Experiments show small improvements on MOS-derived data, while other sets reveal clear improvement over the baselines, highlighting the value of high-quality preference data and demonstrating the effectiveness of the proposed method.
Junyi Fan, Donald S. Williamson
Department of Computer Science and Engineering, The Ohio State University, USA
Mean opinion score (MOS) prediction models are widely used as proxy metrics in text-to-speech (TTS) research, yet their ability to capture quality differences beyond acoustic fidelity remains unclear. We investigate this via controlled perturbations on speech: acoustic degradation, prosodic errors, and manipulation of speaker-specific characteristics such as pitch and speaking rate. We obtained MOS predictions for these speech samples from both human listeners and the model, and analyzed the differences in their perceptual characteristics. Results show that most models track acoustic degradation well, while all are insensitive to prosodic errors despite large subjective score drops. For speaker characteristics, models exhibit a double dissociation: strong mean fundamental frequency (F0) biases absent in human ratings, yet insensitivity to speaking rate and F0 variability that humans notice. These findings highlight limitations of scalar MOS prediction beyond acoustic fidelity.
Masato Takagi, Masaya Kawamura, Reo Shimizu +1
Nagoya Institute of Technology, Japan · LY Corporation, Japan