Non-intrusive MOS predictors are widely used instead of subjective listening tests to evaluate and rank speech enhancement (SE) systems. If they accurately reflect perceived quality, raising their scores should lead to higher-quality speech. We present the first comprehensive analysis of test-time optimization for the SE task, which directly modifies the enhanced signal to raise the average of multiple MOS predictor scores. On seven systems from the URGENT 2026 challenge, we find that 1)~all the optimized predicted scores increase while reference-based metrics remain nearly unchanged, 2)~a non-optimized predicted score does not increase, and 3)~a MUSHRA listening test shows no improvement in perceived quality. These findings reveal a risk that such optimization can distort evaluations, e.g., biasing comparisons of SE systems regardless of their perceived quality. We believe these findings can inform future evaluation practices: they suggest that predictors used for optimization should not be used for evaluation, and that challenges should keep the predictors used for ranking undisclosed.
Figures & tables
Non-intrusive (Optimized)
Unseen
Intrusive
Task-indep.
Task-dependent
Rank
System
λ
DNSMOS
NISQA
UTMOS
SCOREQ
Distill
PESQ
ESTOI
SBERT
LPS
SpkSim
EmoSim
LID
CAcc
non-int.
overall
WR
−
3.06
3.37
2.77
3.67
3.68
2.81
0.86
0.89
0.83
0.76
0.99
0.96
88.16
3
1
0.005
+0.24
+1.19
+0.15
+0.12
+0.02
0.00
0.00
0.00
0.00
0.00
0.00
0.00
+0.05
3 → 1
1 → 1
0.05
+0.50
+1.77
+0.54
+0.28
+0.01
−0.06
−0.01
−0.01
0.00
0.00
0.00
0.00
+0.43
3 → 1
1 → 1
subatom.
−
2.96
3.38
2.41
3.32
3.36
2.71
0.84
0.88
0.79
0.76
0.99
0.95
87.76
6
2
0.005
+0.49
+1.21
+0.21
+0.19
+0.03
−0.01
0.00
−0.01
0.00
0.00
0.00
0.00
−0.66
6 → 2
2 → 2
Table 1: Scores and ranks of the enhanced signal x^ (shaded rows), and score differences and rank changes between the modified signal x~ and x^ . The enhanced signal is jointly optimized for the four predictors, except in the rows marked (S), where it is optimized for DNSMOS alone.
Comparison
Δ MUSHRA
95% CI
p
Equiv. bound
RICK − baird
+8.48
[+2.5,+14.5]
0.011
13.4
baird, λ=0.005
−0.22
[−2.2,+1.7]
0.81
1.8
baird, λ=0.05
−1.20
[−5.0,+2.6]
0.50
4.3
RICK, λ=0.005
−0.65
[−2.6,+1.3]
0.47
2.2
Table 2: Differences in the MUSHRA scores, with a 95% confidence interval and the smallest equivalence bound at the 5% level.