cs.SDSep 30, 2026

ProxyMOS: Label-Free Speech Quality Assessment by Multi-Teacher Distillation with Adaptive Routing

Authors: Maxim Trokunov, Kirill Borodin, Nikita Vasiliev, Grach Mkrtchian

Organizations: Lab260, Yerevan, Armenia · BitmanagerAI, Dubai, UAE · MTUCI, Moscow, Russia

Abstract

Human mean opinion scores (MOS) are costly to collect, and non-intrusive MOS predictors degrade sharply outside their training domain. ProxyMOS turns a pool of public MOS predictors into a single stronger model without new human labels. Eight predictors are benchmarked against human ratings; the five most informative enter a subset search under uniform, correlation-weighted, error-weighted, MSE-optimised and adaptive per-utterance routing; and the best routed four-model ensemble labels 807k unlabeled utterances that train a wav2vec 2.0 student. On URGENT the student reaches Spearman ρ=0.802ρ=0.802 against 0.7730.773 for the best teacher. On mos260, a new Russian TTS benchmark of 4,600 utterances from 38 synthesis conditions, it reaches ρ=0.636ρ=0.636 against 0.6130.613 per utterance and 0.950.95 per condition, matching its own routed ensemble in one forward pass. Adaptive routing is the only rule that does not degrade when weak predictors are added. Model, ONNX exports and mos260 are released. It's about 950 characters; arXiv's limit is 1,920. I kept ρρ because arXiv renders it on the abstract page. If you'd rather avoid math, replace ρ=0.802ρ=0.802 with rho = 0.802 and do the same for the other ...... values.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations

    Jun 18, 2026Masato Takagi, Masaya Kawamura, Reo Shimizu +1Seed-Tts-Eval BenchmarkProsody

  2. PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets

    Jun 17, 2026Junyi Fan, Donald S. WilliamsonMean Opinion ScoresSpeaker

  3. Is Semantics Enough for Speech Mean Opinion Score Prediction?

    Sep 3, 2026Tianyu Lan, Yufei Shi, Yang Ai +3Mean Opinion ScoresSpeaker