We investigate whether predicted Mean Opinion Scores (MOS) can reliably support system-level comparisons of speech enhancement (SE) methods by introducing system-level preference accuracy (SPA). Although MOS prediction models are widely used to evaluate SE systems, their performance is typically assessed by correlation with human-rated MOS, which does not guarantee agreement on which system is better. SPA addresses this gap by directly evaluating whether predicted and human-rated MOS yield the same system preferences. Using SPA, we systematically evaluate three settings: single prediction models, ensembling, and domain adaptation. Through experiments, SPA varies substantially across single prediction models, from 9.4% to 76.8%. Even the best model disagrees with human judgments in approximately 23% of system comparisons. Ensembling yields only limited improvement, while domain adaptation tends to substantially improve SPA in the closed condition but brings only modest gains in the more practical open condition, where neither the target systems nor the speakers are known. These results suggest that SPA can reveal errors correlation-based evaluation alone does not expose, and that predicted MOS alone can lead to unreliable conclusions in practical SE system comparison.
Figures & tables
CHiME-7 UDASE
URGENT24
Model
#M
sys-SRCC
SPA
sys-SRCC
SPA
UniVERSA.
1
0.900
62.0−12.0+8.0
0.938
76.8−3.7+4.2
SCOREQ
1
0.900
65.7−15.7+4.3
0.926
73.6−3.6+3.1
Distill-MOS
1
0.700
49.4−9.4+0.6
0.755
73.1−3.5+4.0
UTMOS
1
0.200
9.4−9.4+0.6
0.902
72.8−4.0+3.9
NISQA
1
0.200
40.9−0.9+9.1
0.702
67.1−3.4+3.3
Table 1: System-level SRCC (“sys-SRCC”) and SPA [%] with 95% CIs. #M indicates number of models used for ensemble. For Ave (Best), #M = 2 on CHiME-7 UDASE and #M = 3 on URGENT24.
CHiME-7 UDASE
UniVERSA.
SCOREQ
UTMOS
Distill-MOS
NISQA
n
closed
open
closed
open
closed
open
closed
open
closed
open
0
62.0
65.7
9.4
49.4
40.9
100
83.0
71.7
92.4
86.1
81.2
50.6
88.7
40.4
62.1
40.4
300
96.7
81.4
98.2
85.9
87.7
64.0
98.8
44.2
85.7
40.3
URGENT24
Table 2: SPA [%] of adapted models for closed and open conditions.
Figure 1: Confusion matrix of human-rated MOS vs. predicted MOS on URGENT24.
Non-intrusive MOS predictors are widely used instead of subjective listening tests to evaluate and rank speech enhancement (SE) systems. If they accurately reflect perceived quality, raising their scores should lead to higher-quality speech. We present the first comprehensive analysis of test-time optimization for the SE task, which directly modifies the enhanced signal to raise the average of multiple MOS predictor scores. On seven systems from the URGENT 2026 challenge, we find that 1)~all the optimized predicted scores increase while reference-based metrics remain nearly unchanged, 2)~a non-optimized predicted score does not increase, and 3)~a MUSHRA listening test shows no improvement in perceived quality. These findings reveal a risk that such optimization can distort evaluations, e.g., biasing comparisons of SE systems regardless of their perceived quality. We believe these findings can inform future evaluation practices: they suggest that predictors used for optimization should not be used for evaluation, and that challenges should keep the predictors used for ranking undisclosed.
Mean opinion scores (MOS) are widely used for speech quality assessment, yet scalar labels are sensitive to rater variability and listening test differences. This introduces labeling noise, which limits the reliability of MOS prediction. Preference prediction reduces this variability as listeners compare signals directly, producing cleaner labels. We study MOS-free preference prediction and propose PrefSQA, which incorporates uncertainty-aware logits, an impairment attention head, and a module based on non-matching-reference comparisons. We use and refine five datasets, including MOS-derived and low-noise simulated sets with matching and non-matching content, experiment with human preference sets, and test on unseen data. Experiments show small improvements on MOS-derived data, while other sets reveal clear improvement over the baselines, highlighting the value of high-quality preference data and demonstrating the effectiveness of the proposed method.
Junyi Fan, Donald S. Williamson
Department of Computer Science and Engineering, The Ohio State University, USA
Mean Opinion Score (MOS) is the gold standard for evaluating synthesized speech naturalness. However, current automatic MOS predictors are dominated by self-supervised learning (SSL) models that prioritize high-level semantics, potentially compromising their ability to capture critical acoustic details. In this paper, we systematically investigate representations from three paradigms: SSLs, acoustic-only neural audio codecs (NACs), and unified NACs that integrate semantics into reconstruction-based architectures. Extensive benchmarking on the standard BVCC and multiple out-of-domain (OOD) datasets demonstrates that features synergizing semantic understanding with fine-grained acoustic modeling achieve a higher performance upper bound in speech quality assessment. Ultimately, our findings highlight that semantics alone are not enough; a dual focus on semantic content and acoustic fidelity is essential for robust MOS prediction.
Tianyu Lan, Yufei Shi, Yang Ai +3
National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China