Speech quality predictors are increasingly used as rewards, yet no agreed measure of their hackability exists. The usual measurement has two flaws. First, the perturbation reaches the predictor through a processing chain -- here a neural codec -- that shifts the score on its own, which scoring against the raw input charges to the attack. Referencing the unperturbed round trip instead changes measured hackability by up to a factor of four (0.31 to 0.08 for one defence). Second, one trained attacker is a sample, not a measurement: five attackers differing only in random seed reach success rates from 0.00 to 0.38 against one fixed predictor, so a defence claim needs the worst case over several. Under this protocol, four published predictors differ widely: NISQA is hacked on 90% of utterances, SSL-MOS on 21%, DNSMOS on 14% and UTMOS on 6%. We then audit a closed attack-detect-patch loop. It hardens the predictor only in its own attack space, by less than the spread between attackers; a random-perturbation baseline matches it; and it costs up to 0.30 system SRCC out of domain. Enhancers post-trained against patched predictors hack them far less (PESQ -0.03 versus -0.23). Code, preregistration and run outputs are released.
Figures & tables
Fig. 1: The attack–detect–patch loop. A GRPO policy perturbs codec latents at a fixed size; an output is flagged only if its score rises above that of the zero-perturbation round trip x^ (dashed) and an independent panel degrades relative to x^ . Flagged pairs patch the predictor; the next round retrains the attacker.
Fig. 2: (a) Benchmark of four common predictors at one perturbation size under the naive ( x ) and corrected ( x^ ) ceiling references; error bars span the five attackers, and the reference moves DNSMOS and NISQA in opposite directions. (b) In-loop attack success rate (BVCC train, corrected reference) of the round- k attacker against the predictor patched in round k−1 , one line per seed: the single-space loop converges in its own space; under multi-space training DAC and PGD fall while the EnCodec attacker stays flat ( +0.01 ).
Predictor
offset
mean ASR
worst ASR
mean gain
SSL-MOS (BVCC)
-0.10
0.205
0.252
+0.191
UTMOS22
-0.07
0.057
0.076
+0.144
NISQA
+0.14
0.896
0.916
+0.954
DNSMOS OVRL
+0.02
0.136
0.200
+0.124
TABLE I: Hackability of common predictors at one size ( ϵ=0.5 , EnCodec latents), five independently trained attackers each, 500 BVCC-test utterances; “worst” is the largest of the five.
offset
ASR by attack space
human agreement
Predictor
R(x^)−R(x)
EnCodec
DAC
PGD
BVCC
OOD
R0 (no defense)
-0.09 ± 0.02
0.226 ± 0.205
0.065 ± 0.113
0.999 ± 0.002
0.922 ± 0.004
0.731 ± 0.006
R1
-0.06 ± 0.03
0.050 ± 0.048
0.161 ± 0.033
1.000 ± 0.001
0.911 ± 0.009
0.690 ± 0.014
R2
0.00 ± 0.07
0.017 ± 0.006
0.141 ± 0.025
1.000 ± 0.000
0.918 ± 0.003
0.682 ± 0.008
R3 (EnCodec loop)
-0.27 ± 0.25
0.137 ± 0.179
0.127 ± 0.031
1.000 ± 0.000
0.920 ± 0.001
0.637 ± 0.046
R3multi (multi-space)
0.95 ± 0.16
0.082 ± 0.107
0.016 ± 0.026
0.841 ± 0.089
0.914 ± 0.003
0.552 ± 0.015
TABLE II: Auditing the loop. Corrected-reference ASR per attack space (fresh attacker, BVCC test) and agreement with human MOS (BVCC system SRCC; out-of-domain mean over 12 sets), mean ± sd over 3 seeds.
Measure
R0
R3
random-latent pairs
Δ own reward
0.263 ± 0.010
0.269 ± 0.028
0.334 ± 0.051
Δ PESQ
-0.226 ± 0.170
-0.034 ± 0.126
0.019 ± 0.050
Δ ESTOI
-0.057 ± 0.010
-0.038 ± 0.013
-0.043 ± 0.006
Δ SI-SDR (dB)
-6.1 ± 1.4
-5.8 ± 1.4
-5.7 ± 0.6
Δ DNSMOS
-0.075 ± 0.031
-0.076 ± 0.056
-0.075 ± 0.043
Δ NISQA
0.335 ± 0.186
0.300 ± 0.072
0.468 ± 0.059
TABLE III: Downstream GRPO enhancement. Change against the supervised pre-trained enhancer on VoiceBank-DEMAND test, per reward predictor. Mean ± sd over 3 seeds ( R2 : one seed).
UTMOS has become one of the most commonly used deep neural network-based speech quality assessment (SQA) metrics in speech processing research. In this paper, we attack UTMOS to probe its robustness. Starting from high-quality speech samples, we optimize the input in two directions: a score-preserving attack, which degrades perceived quality while maintaining the predicted score, and a quality-preserving attack, which lowers the predicted score while maintaining perceived quality. We consider three input spaces: raw waveform, mel spectrogram with a HiFi-GAN vocoder, and the latent space of EnCodec, a neural audio codec. Experimental results show that score-preserving attacks are effective against UTMOS. Although perfect quality-preserving attacks are more difficult, optimization in the EnCodec latent space provides the best chance of success. These results reveal failure modes of UTMOS and highlight the importance of robustness analysis for DNN-based SQA metrics.
Non-intrusive MOS predictors are widely used instead of subjective listening tests to evaluate and rank speech enhancement (SE) systems. If they accurately reflect perceived quality, raising their scores should lead to higher-quality speech. We present the first comprehensive analysis of test-time optimization for the SE task, which directly modifies the enhanced signal to raise the average of multiple MOS predictor scores. On seven systems from the URGENT 2026 challenge, we find that 1)~all the optimized predicted scores increase while reference-based metrics remain nearly unchanged, 2)~a non-optimized predicted score does not increase, and 3)~a MUSHRA listening test shows no improvement in perceived quality. These findings reveal a risk that such optimization can distort evaluations, e.g., biasing comparisons of SE systems regardless of their perceived quality. We believe these findings can inform future evaluation practices: they suggest that predictors used for optimization should not be used for evaluation, and that challenges should keep the predictors used for ranking undisclosed.
Mean opinion score (MOS) prediction models are widely used as proxy metrics in text-to-speech (TTS) research, yet their ability to capture quality differences beyond acoustic fidelity remains unclear. We investigate this via controlled perturbations on speech: acoustic degradation, prosodic errors, and manipulation of speaker-specific characteristics such as pitch and speaking rate. We obtained MOS predictions for these speech samples from both human listeners and the model, and analyzed the differences in their perceptual characteristics. Results show that most models track acoustic degradation well, while all are insensitive to prosodic errors despite large subjective score drops. For speaker characteristics, models exhibit a double dissociation: strong mean fundamental frequency (F0) biases absent in human ratings, yet insensitivity to speaking rate and F0 variability that humans notice. These findings highlight limitations of scalar MOS prediction beyond acoustic fidelity.
Masato Takagi, Masaya Kawamura, Reo Shimizu +1
Nagoya Institute of Technology, Japan · LY Corporation, Japan