Paper ID: 2209.02785

Read it to me: An emotionally aware Speech Narration Application

Rishibha Bansal

In this work we try to perform emotional style transfer on audios. In particular, MelGAN-VC architecture is explored for various emotion-pair transfers. The generated audio is then classified using an LSTM-based emotion classifier for audio. We find that "sad" audio is generated well as compared to "happy" or "anger" as people have similar expressions of sadness.

Submitted: Sep 6, 2022

Topics

Audio Driven
Emotion Transfer

Links

arXiv PDF