cs.SDOct 5, 2026

A Comprehensive Objective Evaluation of Modern Text-to-Speech for Turkish Using Speech Quality Assessment Models

Authors: Yunus Emre Ozkose, Alperen Kahraman, Ali Haznedaroglu

Organizations: Sestek, Ankara, Türkiye

Abstract

Modern text-to-speech (TTS) systems can clone a target speaker from a short reference clip or be fine-tuned on a target voice, yet their behaviour on morphologically rich, lower-resource languages such as Turkish remain under-characterised. We present a systematic benchmark of four contemporary systems (Chatterbox, CosyVoice, OmniVoice, and VoxCPM2) evaluated across fine-tuned and zero-shot configurations, contrasted with a conventional VITS baseline and anchored to natural gold speech. Each configuration is scored with eighteen complementary objective metrics spanning learned naturalness predictors (UTMOS v2, DNSMOS-Pro, SCOREQ, WhisQA, AudioBox-PQ, NatScore, SpeechLMScore), intelligibility and signal-quality estimators (SQUIM PESQ/SI-SDR/STOI, Brouhaha), speaker similarity, distributional fidelity (TTSDS) and low-level acoustic descriptors. We further analyse how quality varies with utterance length and quantify long-form temporal consistency through speaker-identity and naturalness drift over chunked utterances. We release our evaluation code to support reproducible TTS evaluation.

Figures & tables

Explore similar work

CardsList
  1. Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study

    Aug 3, 2026Ali Jafar, Amal Sarmad, Shifa Yousaf +1Seed-Tts-Eval BenchmarkFlow-Matching Text-To-Speech

  2. An Evaluation Framework for Text-to-Speech Voice Reconstruction

    Jun 19, 2026Ariadna Sanchez, Christoph Minixhofer, Korin Richmond +3Seed-Tts-Eval BenchmarkFlow-Matching Text-To-Speech