Human mean opinion scores (MOS) are costly to collect, and non-intrusive MOS predictors degrade sharply outside their training domain. ProxyMOS turns a pool of public MOS predictors into a single stronger model without new human labels. Eight predictors are benchmarked against human ratings; the five most informative enter a subset search under uniform, correlation-weighted, error-weighted, MSE-optimised and adaptive per-utterance routing; and the best routed four-model ensemble labels 807k unlabeled utterances that train a wav2vec 2.0 student. On URGENT the student reaches Spearman ρ=0.802 against 0.773 for the best teacher. On mos260, a new Russian TTS benchmark of 4,600 utterances from 38 synthesis conditions, it reaches ρ=0.636 against 0.613 per utterance and 0.95 per condition, matching its own routed ensemble in one forward pass. Adaptive routing is the only rule that does not degrade when weak predictors are added. Model, ONNX exports and mos260 are released. It's about 950 characters; arXiv's limit is 1,920. I kept ρ because arXiv renders it on the abstract page. If you'd rather avoid math, replace ρ=0.802 with rho = 0.802 and do the same for the other ... values.
Figures & tables
Figure 1: Pipeline. Labeled sets rank the teachers (dashed) and fit the aggregation rule; teachers score the unlabeled corpora, the rule turns scores into targets y^ , and one student trained on them with MSE is tested against human MOS. ASP: attentive statistics pooling.
Corpus
Content
Utt.
Role
URGENT [ 36 ]
enhanced natural, En
6,900
rank, test
BVCC [ 10 ]
TTS / VC, En
3,000
router
SOMOS [ 21 ]
neural TTS, En
20,000
router
mos260 (ours)
TTS / codecs, Ru
4,600
held-out test
MLAAD [ 24 ]
TTS / VC, multilingual
200,000
targets
Balalaika [ 5 ]
natural speech, Ru
440,000
targets
Table 1: Corpora. Human MOS is used only for teacher ranking, router fitting and evaluation; unlabeled corpora receive ensemble targets.
URGENT ( N=6,900 )
mos260 ( N=4,355 )
Model
r↑
ρ↑
τ↑
RMSE ↓
MAE ↓
r↑
ρ↑
τ↑
RMSE ↓
MAE ↓
ProxyMOS (PyTorch)
0.806
0.802
0.620
0.471
0.368
0.691
0.636
0.474
0.897
0.692
ProxyMOS (ONNX FP32)
0.779
0.778
0.591
0.502
0.393
0.700
0.647
0.481
0.883
0.687
ProxyMOS (ONNX FP16)
0.779
0.778
0.591
0.502
0.393
0.700
0.647
0.481
0.883
0.687
Ensemble W+D+U+X, uniform
0.813
0.814
0.629
0.448
0.356
0.674
0.578
0.420
0.886
0.701
Ensemble W+D+U+X, router
0.819
0.816
0.632
0.439
0.348
0.706
0.636
0.471
0.822
0.633
Table 2: Utterance-level agreement with human MOS. Predictions are z-score aligned before RMSE/MAE. Best value per column in bold. Ensemble rows are the four-model teacher ensemble used to label the training corpora; the router is trained on SOMOS and BVCC, so both benchmarks are held out for every row. 95% CI half-width for ρ : ≤0.010 (URGENT), ≤0.025 (mos260).
URGENT
mos260
Rule
K=2
3
4
5
K=2
3
4
5
Best single
0.773 (W)
0.613 (D)
Uniform mean
.799
.809
.814
.804
.617
.602
.578
.510
ρ -weighted
.799
.809
.814
.814
.619
.608
.608
.588
1/RMSE-weighted
.799
.809
.814
.811
.619
.605
.583
.547
MSE-optimised
.795
.809
.813
.811
.617
.602
.578
.510
Table 3: Best Spearman ρ over all teacher subsets of size K under each aggregation rule. MSE-optimised weights are fit on URGENT (in-sample there); the router is trained on SOMOS and BVCC, so both benchmarks are held out for it. Winning subsets at K=4 : W+D+U+X for every rule except ρ -weighting on mos260 (D+U+X+M); D+U ( K=2 ) and D+U+X ( K=3 ) on mos260.
Figure 2: Best ρ versus ensemble size for each aggregation rule (Table 3 ); dotted line: best single teacher. Static rules peak at K=4 on URGENT and decay on mos260 as weaker teachers enter; the router is flat or rising in both domains.
Figure 3: mos260 per synthesis condition, sorted by human MOS: boxes show the distribution of ProxyMOS scores (quartiles, 1.5 IQR whiskers; counts in brackets; the 16 in-house configurations pooled by stage), diamonds the human MOS mean ± SE.
Family
Cond.
Utt.
Human MOS
ProxyMOS
Natural reference
1
200
4.45±0.62
3.99±0.17
Vocoder front-ends
6
1,200
4.05±0.78
3.88±0.24
Zero-shot TTS
5
1,000
3.52±0.80
3.63±0.33
Neural codecs
4
800
3.39±1.11
3.54±0.78
In-house TTS (4 stages)
16
800
2.83±0.90
3.02±0.37
Grad-TTS curriculum
6
600
1.91±0.98
2.69±0.77
Table 4: mos260 by family of synthesis condition: human MOS and ProxyMOS score (mean ± SD over utterances).
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S1: Condition-level agreement on mos260: mean human MOS versus mean ProxyMOS score for the 38 synthesis conditions (bars: standard error; dotted line: identity). Table S1 gives the scores per condition.
Figure S2: Spearman ρ of every model with 95% confidence intervals on URGENT and mos260. Shaded rows: the distilled student.
Figure S3: Best single teacher, the W+D+U+X ensemble under uniform and routed aggregation, and the distilled student (PyTorch), Spearman ρ on both benchmarks.
Figure S4: Distribution of ProxyMOS predictions on the two benchmarks.
Figure S5: Kernel density of z-score-aligned predictions of every model against the human MOS distribution (black) on URGENT (left) and mos260 (right).
Figure S6: Kernel density of raw (unaligned) predictions of every model on URGENT and mos260. Several models (HuBERT-MOS, DNSMOS, MOSNet) collapse to a narrow band, which explains their near-zero correlation.
Figure S7: URGENT: score distributions of the five best uniform-mean ensembles after z-score alignment (blue) against human MOS (grey).
Figure S8: mos260: score distributions of the five best uniform-mean ensembles after z-score alignment against human MOS.
Family
Condition
Utt.
Rated
Human MOS
Human Int-MOS
ProxyMOS
Reference
Natural reference
200
190
4.45±0.62
4.03±0.57
3.99±0.17
Vocoder
Mel + BigVGAN
200
191
4.30±0.58
3.95±0.70
3.99±0.18
Vocoder
DC-AE + Vocos
200
190
4.28±0.56
3.77±0.75
3.86±0.21
Vocoder
Mel + BigVGAN (ft)
200
184
4.22±0.71
3.95±0.62
3.99±0.17
Vocoder
Mel + Vocos
200
183
4.18±0.76
3.83±0.68
3.90±0.20
Neural codec
WavTokenizer
200
194
4.01±0.63
3.54±0.73
4.07±0.20
Appendix
Table S1: mos260 per synthesis condition (in-house configurations pooled by training stage): utterances, utterances with quality ratings, human MOS, human Int-MOS and ProxyMOS score (mean ± SD), sorted by human MOS.
URGENT
mos260
Model
ρ
95% CI
ρ
95% CI
ProxyMOS (PyTorch)
0.802
[0.792, 0.812]
0.636
[0.616, 0.655]
ProxyMOS (ONNX FP32)
0.778
[0.767, 0.788]
0.647
[0.628, 0.666]
ProxyMOS (ONNX FP16)
0.778
[0.767, 0.788]
0.647
[0.628, 0.666]
WhiSQA
0.773
[0.762, 0.784]
0.466
[0.442, 0.491]
DistillMOS
0.748
[0.736, 0.760]
0.613
[0.593, 0.633]
Appendix
\fnum@table : Spearman ρ with 95% confidence intervals (Fisher z with the Spearman correction of Bonett and Wright) for every model on both benchmarks.
URGENT
Subset
K
r
ρ
τ
RMSE
MAE
W+D+U+X
4
0.813
0.814
0.629
0.448
0.356
W+U+X
3
0.804
0.809
0.623
0.460
0.366
D+U+X
3
0.802
0.806
0.621
0.461
0.366
W+D+U
3
0.805
0.805
0.620
0.459
0.365
W+D+X
3
0.810
0.805
0.620
0.455
0.360
Appendix
\fnum@table : Teacher subsets under the uniform mean rule, 25 of 26, each side sorted by its own Spearman ρ .
URGENT
Subset
K
r
ρ
τ
RMSE
MAE
W+D+U+X+M
5
0.816
0.814
0.630
0.442
0.351
W+D+U+X
4
0.813
0.814
0.629
0.448
0.356
W+U+X+M
4
0.809
0.809
0.625
0.450
0.358
W+U+X
3
0.804
0.809
0.623
0.459
0.365
D+U+X+M
4
0.807
0.806
0.622
0.452
0.359
Appendix
\fnum@table : Teacher subsets under the ρ -weighted rule, 25 of 26, each side sorted by its own Spearman ρ .
URGENT
Subset
K
r
ρ
τ
RMSE
MAE
W+D+U+X
4
0.813
0.814
0.629
0.448
0.355
W+D+U+X+M
5
0.818
0.811
0.627
0.436
0.346
W+U+X
3
0.805
0.809
0.623
0.459
0.365
D+U+X
3
0.803
0.806
0.621
0.460
0.366
W+D+U
3
0.806
0.805
0.621
0.458
0.364
Appendix
\fnum@table : Teacher subsets under the 1/RMSE-weighted rule, 25 of 26, each side sorted by its own Spearman ρ .
URGENT
Subset
K
r
ρ
τ
RMSE
MAE
W+D+U+X
4
0.814
0.813
0.629
0.447
0.355
W+D+U+X+M
5
0.819
0.811
0.627
0.435
0.345
W+U+X
3
0.807
0.809
0.623
0.457
0.363
W+U+X+M
4
0.813
0.806
0.622
0.442
0.351
W+D+U
3
0.808
0.805
0.621
0.457
0.362
Appendix
\fnum@table : Teacher subsets under the MSE-optimised rule, 25 of 26, each side sorted by its own Spearman ρ .
URGENT
Subset
K
r
ρ
τ
RMSE
MAE
W+D+U+X+M
5
0.825
0.817
0.633
0.428
0.340
W+D+U+X
4
0.819
0.816
0.632
0.439
0.348
W+U+X
3
0.809
0.810
0.624
0.451
0.358
W+U+X+M
4
0.818
0.810
0.625
0.436
0.346
W+D+U+M
4
0.816
0.807
0.623
0.438
0.348
Appendix
\fnum@table : Teacher subsets under the adaptive router rule, 25 of 26, each side sorted by its own Spearman ρ .
Mean opinion score (MOS) prediction models are widely used as proxy metrics in text-to-speech (TTS) research, yet their ability to capture quality differences beyond acoustic fidelity remains unclear. We investigate this via controlled perturbations on speech: acoustic degradation, prosodic errors, and manipulation of speaker-specific characteristics such as pitch and speaking rate. We obtained MOS predictions for these speech samples from both human listeners and the model, and analyzed the differences in their perceptual characteristics. Results show that most models track acoustic degradation well, while all are insensitive to prosodic errors despite large subjective score drops. For speaker characteristics, models exhibit a double dissociation: strong mean fundamental frequency (F0) biases absent in human ratings, yet insensitivity to speaking rate and F0 variability that humans notice. These findings highlight limitations of scalar MOS prediction beyond acoustic fidelity.
Masato Takagi, Masaya Kawamura, Reo Shimizu +1
Nagoya Institute of Technology, Japan · LY Corporation, Japan
Mean opinion scores (MOS) are widely used for speech quality assessment, yet scalar labels are sensitive to rater variability and listening test differences. This introduces labeling noise, which limits the reliability of MOS prediction. Preference prediction reduces this variability as listeners compare signals directly, producing cleaner labels. We study MOS-free preference prediction and propose PrefSQA, which incorporates uncertainty-aware logits, an impairment attention head, and a module based on non-matching-reference comparisons. We use and refine five datasets, including MOS-derived and low-noise simulated sets with matching and non-matching content, experiment with human preference sets, and test on unseen data. Experiments show small improvements on MOS-derived data, while other sets reveal clear improvement over the baselines, highlighting the value of high-quality preference data and demonstrating the effectiveness of the proposed method.
Junyi Fan, Donald S. Williamson
Department of Computer Science and Engineering, The Ohio State University, USA
Mean Opinion Score (MOS) is the gold standard for evaluating synthesized speech naturalness. However, current automatic MOS predictors are dominated by self-supervised learning (SSL) models that prioritize high-level semantics, potentially compromising their ability to capture critical acoustic details. In this paper, we systematically investigate representations from three paradigms: SSLs, acoustic-only neural audio codecs (NACs), and unified NACs that integrate semantics into reconstruction-based architectures. Extensive benchmarking on the standard BVCC and multiple out-of-domain (OOD) datasets demonstrates that features synergizing semantic understanding with fine-grained acoustic modeling achieve a higher performance upper bound in speech quality assessment. Ultimately, our findings highlight that semantics alone are not enough; a dual focus on semantic content and acoustic fidelity is essential for robust MOS prediction.
Tianyu Lan, Yufei Shi, Yang Ai +3
National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China