Human mean opinion scores (MOS) are costly to collect, and non-intrusive MOS predictors degrade sharply outside their training domain. ProxyMOS turns a pool of public MOS predictors into a single stronger model without new human labels. Eight predictors are benchmarked against human ratings; the five most informative enter a subset search under uniform, correlation-weighted, error-weighted, MSE-optimised and adaptive per-utterance routing; and the best routed four-model ensemble labels 807k unlabeled utterances that train a wav2vec 2.0 student. On URGENT the student reaches Spearman ρ=0.802 against 0.773 for the best teacher. On mos260, a new Russian TTS benchmark of 4,600 utterances from 38 synthesis conditions, it reaches ρ=0.636 against 0.613 per utterance and 0.95 per condition, matching its own routed ensemble in one forward pass. Adaptive routing is the only rule that does not degrade when weak predictors are added. Model, ONNX exports and mos260 are released. It's about 950 characters; arXiv's limit is 1,920. I kept ρ because arXiv renders it on the abstract page. If you'd rather avoid math, replace ρ=0.802 with rho = 0.802 and do the same for the other ... values.
Figures & tables
Figure 1: Pipeline. Labeled sets rank the teachers (dashed) and fit the aggregation rule; teachers score the unlabeled corpora, the rule turns scores into targets y^ , and one student trained on them with MSE is tested against human MOS. ASP: attentive statistics pooling.
Corpus
Content
Utt.
Role
URGENT [ 36 ]
enhanced natural, En
6,900
rank, test
BVCC [ 10 ]
TTS / VC, En
3,000
router
SOMOS [ 21 ]
neural TTS, En
20,000
router
mos260 (ours)
TTS / codecs, Ru
4,600
held-out test
MLAAD [ 24 ]
TTS / VC, multilingual
200,000
targets
Balalaika [ 5 ]
natural speech, Ru
440,000
targets
Table 1: Corpora. Human MOS is used only for teacher ranking, router fitting and evaluation; unlabeled corpora receive ensemble targets.
URGENT ( N=6,900 )
mos260 ( N=4,355 )
Model
r↑
ρ↑
τ↑
RMSE ↓
MAE ↓
r↑
ρ↑
τ↑
RMSE ↓
MAE ↓
ProxyMOS (PyTorch)
0.806
0.802
0.620
0.471
0.368
0.691
0.636
0.474
0.897
0.692
ProxyMOS (ONNX FP32)
0.779
0.778
0.591
0.502
0.393
0.700
0.647
0.481
0.883
0.687
ProxyMOS (ONNX FP16)
0.779
0.778
0.591
0.502
0.393
0.700
0.647
0.481
0.883
0.687
Ensemble W+D+U+X, uniform
0.813
0.814
0.629
0.448
0.356
0.674
0.578
0.420
0.886
0.701
Ensemble W+D+U+X, router
0.819
0.816
0.632
0.439
0.348
0.706
0.636
0.471
0.822
0.633
Table 2: Utterance-level agreement with human MOS. Predictions are z-score aligned before RMSE/MAE. Best value per column in bold. Ensemble rows are the four-model teacher ensemble used to label the training corpora; the router is trained on SOMOS and BVCC, so both benchmarks are held out for every row. 95% CI half-width for ρ : ≤0.010 (URGENT), ≤0.025 (mos260).
URGENT
mos260
Rule
K=2
3
4
5
K=2
3
4
5
Best single
0.773 (W)
0.613 (D)
Uniform mean
.799
.809
.814
.804
.617
.602
.578
.510
ρ -weighted
.799
.809
.814
.814
.619
.608
.608
.588
1/RMSE-weighted
.799
.809
.814
.811
.619
.605
.583
.547
MSE-optimised
.795
.809
.813
.811
.617
.602
.578
.510
Table 3: Best Spearman ρ over all teacher subsets of size K under each aggregation rule. MSE-optimised weights are fit on URGENT (in-sample there); the router is trained on SOMOS and BVCC, so both benchmarks are held out for it. Winning subsets at K=4 : W+D+U+X for every rule except ρ -weighting on mos260 (D+U+X+M); D+U ( K=2 ) and D+U+X ( K=3 ) on mos260.
Figure 2: Best ρ versus ensemble size for each aggregation rule (Table 3 ); dotted line: best single teacher. Static rules peak at K=4 on URGENT and decay on mos260 as weaker teachers enter; the router is flat or rising in both domains.
Figure 3: mos260 per synthesis condition, sorted by human MOS: boxes show the distribution of ProxyMOS scores (quartiles, 1.5 IQR whiskers; counts in brackets; the 16 in-house configurations pooled by stage), diamonds the human MOS mean ± SE.
Family
Cond.
Utt.
Human MOS
ProxyMOS
Natural reference
1
200
4.45±0.62
3.99±0.17
Vocoder front-ends
6
1,200
4.05±0.78
3.88±0.24
Zero-shot TTS
5
1,000
3.52±0.80
3.63±0.33
Neural codecs
4
800
3.39±1.11
3.54±0.78
In-house TTS (4 stages)
16
800
2.83±0.90
3.02±0.37
Grad-TTS curriculum
6
600
1.91±0.98
2.69±0.77
Table 4: mos260 by family of synthesis condition: human MOS and ProxyMOS score (mean ± SD over utterances).
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S1: Condition-level agreement on mos260: mean human MOS versus mean ProxyMOS score for the 38 synthesis conditions (bars: standard error; dotted line: identity). Table S1 gives the scores per condition.
Figure S2: Spearman ρ of every model with 95% confidence intervals on URGENT and mos260. Shaded rows: the distilled student.
Figure S3: Best single teacher, the W+D+U+X ensemble under uniform and routed aggregation, and the distilled student (PyTorch), Spearman ρ on both benchmarks.
Figure S4: Distribution of ProxyMOS predictions on the two benchmarks.
Figure S5: Kernel density of z-score-aligned predictions of every model against the human MOS distribution (black) on URGENT (left) and mos260 (right).
Figure S6: Kernel density of raw (unaligned) predictions of every model on URGENT and mos260. Several models (HuBERT-MOS, DNSMOS, MOSNet) collapse to a narrow band, which explains their near-zero correlation.
Figure S7: URGENT: score distributions of the five best uniform-mean ensembles after z-score alignment (blue) against human MOS (grey).
Figure S8: mos260: score distributions of the five best uniform-mean ensembles after z-score alignment against human MOS.
Family
Condition
Utt.
Rated
Human MOS
Human Int-MOS
ProxyMOS
Reference
Natural reference
200
190
4.45±0.62
4.03±0.57
3.99±0.17
Vocoder
Mel + BigVGAN
200
191
4.30±0.58
3.95±0.70
3.99±0.18
Vocoder
DC-AE + Vocos
200
190
4.28±0.56
3.77±0.75
3.86±0.21
Vocoder
Mel + BigVGAN (ft)
200
184
4.22±0.71
3.95±0.62
3.99±0.17
Vocoder
Mel + Vocos
200
183
4.18±0.76
3.83±0.68
3.90±0.20
Neural codec
WavTokenizer
200
194
4.01±0.63
3.54±0.73
4.07±0.20
Appendix
Table S1: mos260 per synthesis condition (in-house configurations pooled by training stage): utterances, utterances with quality ratings, human MOS, human Int-MOS and ProxyMOS score (mean ± SD), sorted by human MOS.
URGENT
mos260
Model
ρ
95% CI
ρ
95% CI
ProxyMOS (PyTorch)
0.802
[0.792, 0.812]
0.636
[0.616, 0.655]
ProxyMOS (ONNX FP32)
0.778
[0.767, 0.788]
0.647
[0.628, 0.666]
ProxyMOS (ONNX FP16)
0.778
[0.767, 0.788]
0.647
[0.628, 0.666]
WhiSQA
0.773
[0.762, 0.784]
0.466
[0.442, 0.491]
DistillMOS
0.748
[0.736, 0.760]
0.613
[0.593, 0.633]
Appendix
\fnum@table : Spearman ρ with 95% confidence intervals (Fisher z with the Spearman correction of Bonett and Wright) for every model on both benchmarks.
URGENT
Subset
K
r
ρ
τ
RMSE
MAE
W+D+U+X
4
0.813
0.814
0.629
0.448
0.356
W+U+X
3
0.804
0.809
0.623
0.460
0.366
D+U+X
3
0.802
0.806
0.621
0.461
0.366
W+D+U
3
0.805
0.805
0.620
0.459
0.365
W+D+X
3
0.810
0.805
0.620
0.455
0.360
Appendix
\fnum@table : Teacher subsets under the uniform mean rule, 25 of 26, each side sorted by its own Spearman ρ .
URGENT
Subset
K
r
ρ
τ
RMSE
MAE
W+D+U+X+M
5
0.816
0.814
0.630
0.442
0.351
W+D+U+X
4
0.813
0.814
0.629
0.448
0.356
W+U+X+M
4
0.809
0.809
0.625
0.450
0.358
W+U+X
3
0.804
0.809
0.623
0.459
0.365
D+U+X+M
4
0.807
0.806
0.622
0.452
0.359
Appendix
\fnum@table : Teacher subsets under the ρ -weighted rule, 25 of 26, each side sorted by its own Spearman ρ .
URGENT
Subset
K
r
ρ
τ
RMSE
MAE
W+D+U+X
4
0.813
0.814
0.629
0.448
0.355
W+D+U+X+M
5
0.818
0.811
0.627
0.436
0.346
W+U+X
3
0.805
0.809
0.623
0.459
0.365
D+U+X
3
0.803
0.806
0.621
0.460
0.366
W+D+U
3
0.806
0.805
0.621
0.458
0.364
Appendix
\fnum@table : Teacher subsets under the 1/RMSE-weighted rule, 25 of 26, each side sorted by its own Spearman ρ .
URGENT
Subset
K
r
ρ
τ
RMSE
MAE
W+D+U+X
4
0.814
0.813
0.629
0.447
0.355
W+D+U+X+M
5
0.819
0.811
0.627
0.435
0.345
W+U+X
3
0.807
0.809
0.623
0.457
0.363
W+U+X+M
4
0.813
0.806
0.622
0.442
0.351
W+D+U
3
0.808
0.805
0.621
0.457
0.362
Appendix
\fnum@table : Teacher subsets under the MSE-optimised rule, 25 of 26, each side sorted by its own Spearman ρ .
URGENT
Subset
K
r
ρ
τ
RMSE
MAE
W+D+U+X+M
5
0.825
0.817
0.633
0.428
0.340
W+D+U+X
4
0.819
0.816
0.632
0.439
0.348
W+U+X
3
0.809
0.810
0.624
0.451
0.358
W+U+X+M
4
0.818
0.810
0.625
0.436
0.346
W+D+U+M
4
0.816
0.807
0.623
0.438
0.348
Appendix
\fnum@table : Teacher subsets under the adaptive router rule, 25 of 26, each side sorted by its own Spearman ρ .