eess.ASOct 7, 2026

Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners

Authors: Xilin Jiang, Shun Zhang, Tejas Jayashankar, Yinghao Aaron Li, Osama Hanna

Organizations: Columbia University New York, NY, USA · Meta Superintelligence Labs Menlo Park, CA, USA

Abstract

We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts. Given a context and a response speech, CVAM describes salient moments that characterize the voice and predicts nine categorical attributes spanning gender, pitch, pacing, emotion, and delivery. The key challenge lies in perceptual fields such as emotion and delivery, which are inherently subjective and lack definitive ground truth. Therefore, we collect ~10 human annotations for each of 3k real and synthetic responses derived from the CANDOR corpus. CVAM is supervised finetuned on synthesized aesthetic descriptions and labels, then optimized with Group Relative Policy Optimization on human judgments. Experiments show that CVAM better agrees with human listeners than Gemini 3.1 Pro and open-source speech LLMs, and outperforms single-human-vs.-rest agreement. Together, we demonstrate the importance of grounding voice aesthetics in human perception and propose a principled framework for human alignment.

Figures & tables

Explore similar work

CardsList
  1. Ready to Speak: Aligning LLMs for TTS-Friendly Text Generation

    Sep 1, 2026Thibaut Thonet, Jos Rozen, Laurent BesacierLLM AlignmentTTS Evaluation

  2. Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech

    Oct 8, 2026Shiao Zhu, Lianbo Liu, Sizhen Lyu +3TTS SynthesisControllable Speech Generation

  3. Beyond Prompt Adherence: Auditing Attribute-Level Voice Control in Speech Generation

    Aug 1, 2026Xianhao Zhou, Jianghao WuSpeech Generation EvaluationTTS Evaluation