Speech quality predictors are increasingly used as rewards, yet no agreed measure of their hackability exists. The usual measurement has two flaws. First, the perturbation reaches the predictor through a processing chain -- here a neural codec -- that shifts the score on its own, which scoring against the raw input charges to the attack. Referencing the unperturbed round trip instead changes measured hackability by up to a factor of four (0.31 to 0.08 for one defence). Second, one trained attacker is a sample, not a measurement: five attackers differing only in random seed reach success rates from 0.00 to 0.38 against one fixed predictor, so a defence claim needs the worst case over several. Under this protocol, four published predictors differ widely: NISQA is hacked on 90% of utterances, SSL-MOS on 21%, DNSMOS on 14% and UTMOS on 6%. We then audit a closed attack-detect-patch loop. It hardens the predictor only in its own attack space, by less than the spread between attackers; a random-perturbation baseline matches it; and it costs up to 0.30 system SRCC out of domain. Enhancers post-trained against patched predictors hack them far less (PESQ -0.03 versus -0.23). Code, preregistration and run outputs are released.
Figures & tables
Fig. 1: The attack–detect–patch loop. A GRPO policy perturbs codec latents at a fixed size; an output is flagged only if its score rises above that of the zero-perturbation round trip x^ (dashed) and an independent panel degrades relative to x^ . Flagged pairs patch the predictor; the next round retrains the attacker.
Fig. 2: (a) Benchmark of four common predictors at one perturbation size under the naive ( x ) and corrected ( x^ ) ceiling references; error bars span the five attackers, and the reference moves DNSMOS and NISQA in opposite directions. (b) In-loop attack success rate (BVCC train, corrected reference) of the round- k attacker against the predictor patched in round k−1 , one line per seed: the single-space loop converges in its own space; under multi-space training DAC and PGD fall while the EnCodec attacker stays flat ( +0.01 ).
Predictor
offset
mean ASR
worst ASR
mean gain
SSL-MOS (BVCC)
-0.10
0.205
0.252
+0.191
UTMOS22
-0.07
0.057
0.076
+0.144
NISQA
+0.14
0.896
0.916
+0.954
DNSMOS OVRL
+0.02
0.136
0.200
+0.124
TABLE I: Hackability of common predictors at one size ( ϵ=0.5 , EnCodec latents), five independently trained attackers each, 500 BVCC-test utterances; “worst” is the largest of the five.
offset
ASR by attack space
human agreement
Predictor
R(x^)−R(x)
EnCodec
DAC
PGD
BVCC
OOD
R0 (no defense)
-0.09 ± 0.02
0.226 ± 0.205
0.065 ± 0.113
0.999 ± 0.002
0.922 ± 0.004
0.731 ± 0.006
R1
-0.06 ± 0.03
0.050 ± 0.048
0.161 ± 0.033
1.000 ± 0.001
0.911 ± 0.009
0.690 ± 0.014
R2
0.00 ± 0.07
0.017 ± 0.006
0.141 ± 0.025
1.000 ± 0.000
0.918 ± 0.003
0.682 ± 0.008
R3 (EnCodec loop)
-0.27 ± 0.25
0.137 ± 0.179
0.127 ± 0.031
1.000 ± 0.000
0.920 ± 0.001
0.637 ± 0.046
R3multi (multi-space)
0.95 ± 0.16
0.082 ± 0.107
0.016 ± 0.026
0.841 ± 0.089
0.914 ± 0.003
0.552 ± 0.015
TABLE II: Auditing the loop. Corrected-reference ASR per attack space (fresh attacker, BVCC test) and agreement with human MOS (BVCC system SRCC; out-of-domain mean over 12 sets), mean ± sd over 3 seeds.
Measure
R0
R3
random-latent pairs
Δ own reward
0.263 ± 0.010
0.269 ± 0.028
0.334 ± 0.051
Δ PESQ
-0.226 ± 0.170
-0.034 ± 0.126
0.019 ± 0.050
Δ ESTOI
-0.057 ± 0.010
-0.038 ± 0.013
-0.043 ± 0.006
Δ SI-SDR (dB)
-6.1 ± 1.4
-5.8 ± 1.4
-5.7 ± 0.6
Δ DNSMOS
-0.075 ± 0.031
-0.076 ± 0.056
-0.075 ± 0.043
Δ NISQA
0.335 ± 0.186
0.300 ± 0.072
0.468 ± 0.059
TABLE III: Downstream GRPO enhancement. Change against the supervised pre-trained enhancer on VoiceBank-DEMAND test, per reward predictor. Mean ± sd over 3 seeds ( R2 : one seed).