cs.SDSep 30, 2026

How Reliable Are Predicted MOS for Reproducing Human System-Level Preferences in Speech Enhancement?

Authors: Nahomi Kusunoki, Tsubasa Ochiai, Naohiro Tawara, Marc Delcroix, Naoyuki Kamo, Tetsuji Ogawa, Shoko Araki

Organizations: Waseda University, Japan · NTT, Inc., Japan

Abstract

We investigate whether predicted Mean Opinion Scores (MOS) can reliably support system-level comparisons of speech enhancement (SE) methods by introducing system-level preference accuracy (SPA). Although MOS prediction models are widely used to evaluate SE systems, their performance is typically assessed by correlation with human-rated MOS, which does not guarantee agreement on which system is better. SPA addresses this gap by directly evaluating whether predicted and human-rated MOS yield the same system preferences. Using SPA, we systematically evaluate three settings: single prediction models, ensembling, and domain adaptation. Through experiments, SPA varies substantially across single prediction models, from 9.4% to 76.8%. Even the best model disagrees with human judgments in approximately 23% of system comparisons. Ensembling yields only limited improvement, while domain adaptation tends to substantially improve SPA in the closed condition but brings only modest gains in the more practical open condition, where neither the target systems nor the speakers are known. These results suggest that SPA can reveal errors correlation-based evaluation alone does not expose, and that predicted MOS alone can lead to unreliable conclusions in practical SE system comparison.

Figures & tables

Explore similar work

CardsList
  1. Improving Predicted MOS Scores, Not Perceived Quality: Multi-Predictor Test-Time Optimization of Enhanced Speech

    Sep 30, 2026Tsubasa Ochiai, Marc Delcroix, Nahomi Kusunoki +5Mean Opinion ScoresSpeech Enhancement

  2. PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets

    Jun 17, 2026Junyi Fan, Donald S. WilliamsonMean Opinion ScoresSpeaker

  3. Is Semantics Enough for Speech Mean Opinion Score Prediction?

    Sep 3, 2026Tianyu Lan, Yufei Shi, Yang Ai +3Mean Opinion ScoresSpeaker