eess.ASSep 30, 2026

A barrier or a booster? Familiarity effects on Mandarin emotion prosody recognition using AI-powered voice cloning

Authors: Feng Xu, Gaoyuan Zhang, Shanshan Xue, Yixiang Chen, Hanrui Zhou, Xurong Xie, Hui Chen

Organizations: Institute of Software, Chinese Academy of Sciences · Department of Linguistics, Macquarie University

Abstract

Emotion prosody perception requires simultaneous processing of acoustic cues and speaker identity. While listeners effortlessly decode natural speech, AI synthetic voices introduce cognitive complexities due to subtle acoustic atypicalities. It remains unclear how these synthetic features interact with a listener's prior social knowledge and memory of a familiar speaker. This study investigated how speech sources (human vs. AI) and speaker familiarity affect emotion recognition accuracy and cognitive load. A within-subject task with Mandarin-speaking adults evaluated behavioral (accuracy, reaction time) and physiological data (heart rate variability). Results showed that human voices yielded significantly higher accuracy and faster processing times than AI voices, while HRV did not significantly differentiate between conditions. These findings show that decoding synthetic speech is gated by top-down social cognition, highlighting limitations in current AI synthesis technologies.

Explore similar work

CardsList
  1. Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings

    May 4, 2026Vamshi Nallaguntla, Shruti Kshirsagar, Anderson R. AvilaAudio Deepfake DetectionGrapheme-To-Phoneme

  2. Information-Geometric Superposed Vowel Evaluation: Part 1. Moraic Syllabary (Japanese)

    Jul 5, 2026Yusei Tamura, Shigekazu Ishihara, Ken ItoConsonantsNatural Language Processing

  3. Real-Time Voice AI Hears but Does Not Listen

    Jun 24, 2026Martijn Bartelds, Federico Bianchi, James ZouVoice AgentsVocalizations