eess.ASSep 30, 2026

Improving Predicted MOS Scores, Not Perceived Quality: Multi-Predictor Test-Time Optimization of Enhanced Speech

Authors: Tsubasa Ochiai, Marc Delcroix, Nahomi Kusunoki, Rintaro Ikeshita, Naohiro Tawara, Naoyuki Kamo, Tetsuji Ogawa, Shoko Araki

Organizations: NTT, Inc., Japan · Waseda University, Japan

Abstract

Non-intrusive MOS predictors are widely used instead of subjective listening tests to evaluate and rank speech enhancement (SE) systems. If they accurately reflect perceived quality, raising their scores should lead to higher-quality speech. We present the first comprehensive analysis of test-time optimization for the SE task, which directly modifies the enhanced signal to raise the average of multiple MOS predictor scores. On seven systems from the URGENT 2026 challenge, we find that 1)~all the optimized predicted scores increase while reference-based metrics remain nearly unchanged, 2)~a non-optimized predicted score does not increase, and 3)~a MUSHRA listening test shows no improvement in perceived quality. These findings reveal a risk that such optimization can distort evaluations, e.g., biasing comparisons of SE systems regardless of their perceived quality. We believe these findings can inform future evaluation practices: they suggest that predictors used for optimization should not be used for evaluation, and that challenges should keep the predictors used for ranking undisclosed.

Figures & tables

Explore similar work

CardsList
  1. How Reliable Are Predicted MOS for Reproducing Human System-Level Preferences in Speech Enhancement?

    Sep 30, 2026Nahomi Kusunoki, Tsubasa Ochiai, Naohiro Tawara +4Mean Opinion ScoresSpeech Enhancement

  2. PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets

    Jun 17, 2026Junyi Fan, Donald S. WilliamsonMean Opinion ScoresSpeaker

  3. Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations

    Jun 18, 2026Masato Takagi, Masaya Kawamura, Reo Shimizu +1Seed-Tts-Eval BenchmarkProsody