eess.ASSep 30, 2026

Intonation Perception in Real and Synthetic Speech across Varying Familiarity Levels: A Pilot Study of Equivalence Assessment

Authors: Hanrui Zhou, Gaoyuan Zhang, Yixiang Chen, Yujie Xing, Feng Xu, Xurong Xie, Hui Chen

Abstract

Language training relies on a corpus constructed by a large number linguistic materials. AI-powered voice clones provide a way to construct the corpus with relatively low cost. Singing voice conversion (SVC) model is used to generate synthetic voices. This study compares participants' performances on natural and synthetic speech in two experiments, similarity perception and intonation recognition. In the accuracy of similarity perception task, a significant interaction between speech type and intonation is found, suggesting that question may serve as a cue for speaker identification but may be influenced by synthetic features. In the accuracy of intonation recognition task, a significant interaction between speech type and familiarity is observed, indicating that speech type affects how much familiarity contributes to voice processing.

Explore similar work

Sep 30, 2026eess.AS

A barrier or a booster? Familiarity effects on Mandarin emotion prosody recognition using AI-powered voice cloning

Emotion prosody perception requires simultaneous processing of acoustic cues and speaker identity. While listeners effortlessly decode natural speech, AI synthetic voices introduce cognitive complexities due to subtle acoustic atypicalities. It remains unclear how these synthetic features interact with a listener's prior social knowledge and memory of a familiar speaker. This study investigated how speech sources (human vs. AI) and speaker familiarity affect emotion recognition accuracy and cognitive load. A within-subject task with Mandarin-speaking adults evaluated behavioral (accuracy, reaction time) and physiological data (heart rate variability). Results showed that human voices yielded significantly higher accuracy and faster processing times than AI voices, while HRV did not significantly differentiate between conditions. These findings show that decoding synthetic speech is gated by top-down social cognition, highlighting limitations in current AI synthesis technologies.
Sep 24, 2026cs.SD

Accent Analogy Guidance: More Speaker Similarity at Equal Accent in Cross-Lingual Voice Cloning

In cross-lingual zero-shot text-to-speech, the accent of the reference leaks into the target speech. We propose accent analogy guidance (AAG), a training-free sampler term that subtracts an accent direction estimated from the model's own predictions for one synthetic voice rendered in both languages, so the voice cancels and only the accent remains. By a blind LLM accent judge on real dubbing data, reweighting classifier-free guidance between reference and text, and its variants, stay near one identity-accent trade-off curve; we score a method by its speaker similarity above that curve at equal accent (ΔΔSIM). Across four open TTS models AAG lies above the curve: on OmniVoice ΔΔSIM is +0.11 to +0.27 on three test sets (accent 3.51 to 4.28 on a 1-5 scale at speaker similarity 0.29, where reweighting keeps 0.02); MaskGCT and CosyVoice 2 also lie above their curves, and on F5-TTS it is more native than any reweighting setting. An LLM-free language-ID measure and a twelve-listener panel agree. A premise test and the reach of a model's own curve indicate in advance whether and roughly how much AAG can gain, predicting the one model where it gains nothing (X-Voice).
Apr 2, 2026cs.SD

Acoustic and perceptual differences between standard and accented speech and their voice clones

Voice cloning is often evaluated in terms of overall quality, but less is known about accent preservation and its perceptual consequences. We compare standard and heavily accented Mandarin speech and their voice clones using a combined computational and perceptual design. Embedding-based analyses showed larger original-clone distances for accented speakers in several speaker-discriminative embedding spaces, but this difference disappeared after adjusting for each speaker's within-original baseline variability. In the perception study, clones are rated as more similar to their originals for standard than for accented speakers, and intelligibility increases from original to clone, with a larger gain for accented speech. These results show that accent variation can shape perceived identity match and intelligibility in voice cloning even when it is not observed in baseline-adjusted speaker-embedding distance, and they motivate treating accent preservation as an explicit component of speaker identity preservation, rather than assuming that it is fully captured by off-the-shelf speaker-discriminative embeddings.