SEER: Source-Conditioned Emotion Enhancement via Retrieval for Cochlear-Implant Speech
Authors: Hsing-Hang Chou, Yun-Shao Lin, Ching-Chin Sung, Chi-Chun Lee
Organizations: Department of Electrical Engineering, National Tsing Hua University, Taiwan · Research Center for Information Technology Innovation, Academia Sinica, Taiwan
Cochlear implants (CIs) restore speech access but weaken cues needed for vocal emotion recognition. Prior CI-oriented enhancement requires parallel normal/strong recordings and intensity labels. We propose SEER, a retrieval-based framework that learns which same-emotion reference helps each source remain recognizable after CI processing. A source-conditioned retriever learns CI-aware utility from sampled emotional voice conversion outcomes, while uncertainty-guided exploration avoids exhaustive pair evaluation; neither parallel recordings nor intensity labels are required. SEER improves Source macro-F1 at N8 by 7.30 points on RAVDESS and 11.66 points on ESD, with significant ESD gains across N4/N8/N16. Sixteen-listener RAVDESS gains are significant across all conditions. Exhaustive analysis finds an aggregate benefit from stronger references but little effect from matching gender or content.
The Chinese University of Hong Kong, Hong Kong SAR, China · Institute of Software, Chinese Academy of Sciences, China · National Research Council Canada, Canada +1