We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts. Given a context and a response speech, CVAM describes salient moments that characterize the voice and predicts nine categorical attributes spanning gender, pitch, pacing, emotion, and delivery. The key challenge lies in perceptual fields such as emotion and delivery, which are inherently subjective and lack definitive ground truth. Therefore, we collect ~10 human annotations for each of 3k real and synthetic responses derived from the CANDOR corpus. CVAM is supervised finetuned on synthesized aesthetic descriptions and labels, then optimized with Group Relative Policy Optimization on human judgments. Experiments show that CVAM better agrees with human listeners than Gemini 3.1 Pro and open-source speech LLMs, and outperforms single-human-vs.-rest agreement. Together, we demonstrate the importance of grounding voice aesthetics in human perception and propose a principled framework for human alignment.
Figures & tables
Figure 1: With supervised finetuning on synthetic annotation and reinforcement learning on human crowd judgments, CVAM better agrees with the human crowd on emotion and delivery labels than Gemini 3.1 Pro and evaluated open-source speech LLMs in both accuracy and macro F1 against human majority-voted labels. Notably, CVAM also exceeds the single-human-vs-rest agreement.
Figure 2: CVAM receives context & response speech and outputs JSON descriptions and labels for the response’s gender, pitch, pacing, emotion, and delivery. Colors indicate target sources: self-reported gender and tool-measured objective labels , LLM-written descriptions , and perceptual labels from an LLM for SFT or humans for GRPO .
Figure 3: Distributions of the objective labels (measured by speech analysis tools) and perceptual labels (human majority-voted) of 3k human-annotated speech responses from the CANDOR corpus.
System
Emotion
Arousal
Warmth
Confidence
Avg. Perceptual
Acc
F1
Acc
F1
Acc
F1
Acc
F1
Acc
F1
Baseline Speech Large Language Models
Qwen2.5-Omni [ 30 ]
70.7
27.6
64.1
38.1
42.1
28.5
43.5
30.3
55.1
31.1
Qwen3-Omni [ 31 ]
75.1
29.8
41.3
39.1
37.1
24.7
38.3
26.3
47.9
30.0
Step-Audio-2 mini [ 29 ]
51.7
24.1
61.3
34.7
43.9
29.7
53.4
35.2
52.6
30.9
Gemma-4-12B [ 8 ]
55.5
25.2
39.3
29.7
44.0
29.7
44.7
36.2
45.9
30.2
Table 1: Agreement with human listeners on the perceptual fields. Accuracy and macro-F1 are against the majority-voted ( N≈10 ) label for the perceptual fields. † Tools measure provides the pitch and pacing labels and the evidence log. Single Human scores each one of the N raters against the other N−1 and reports the average.
System
Gender
PitchHgt
PitchVar
Speed
Rhythm
Acc
F1
Acc
F1
Acc
F1
Acc
F1
Acc
F1
Baseline Speech Large Language Models
Qwen2.5-Omni
86.0
84.2
30.7
9.5
32.6
12.5
30.9
9.4
15.9
8.2
Qwen3-Omni
94.4
93.9
41.1
26.2
29.4
20.1
32.8
18.1
15.8
9.4
Step-Audio-2 mini
76.2
75.4
30.7
9.8
24.9
13.8
30.9
10.1
23.7
13.6
Gemma-4-12B
72.1
66.2
32.7
16.9
26.6
17.9
29.9
16.3
29.7
16.6
Table 2: Agreement with speaker self-reported gender and tool-measured objective fields. † Tools measure provides the pitch and pacing labels and the evidence log, so those rows reach 100% by construction. Only compare models without tool measure.
GRPO human Reward
Emotion
Arousal
Warmth
Confidence
Avg. Perceptual
Acc
F1
Acc
F1
Acc
F1
Acc
F1
Acc
F1
† SFT synth Baseline
72.9
36.1
55.0
46.0
60.1
47.1
47.5
44.8
58.9
43.5
Acc on single-human labels
84.7
35.2
77.7
53.9
74.7
51.7
60.3
52.9
74.4
48.4
Acc on majority-voted labels
86.6
38.5
77.6
48.5
78.2
51.7
64.4
49.5
76.7
47.0
Fraction of shared labels
88.1
33.0
78.8
52.5
78.7
53.6
65.7
51.1
77.8
47.6
TVD on votes’ distribution
77.2
37.4
74.1
58.7
65.3
53.7
52.5
49.1
67.3
49.7
Table 3: Study of GRPO rewards from the same SFT checkpoint. The first three are accuracy-based, and the last are distribution-based.
Version
Emotion
Arousal
Warmth
Confidence
Avg. Perceptual
Acc
F1
Acc
F1
Acc
F1
Acc
F1
Acc
F1
Real
82.0
34.0
76.0
52.3
68.6
45.9
51.0
45.0
69.4
44.3
Voice-Cloned
90.8
68.7
81.2
50.4
65.2
40.3
52.8
39.7
72.5
49.8
Voice-Designed
79.0
33.6
68.6
62.8
64.2
51.4
55.6
52.8
66.9
50.1
Table 4: Performance of CVAM ( † SFT synth + GRPO humanKL ) on real and synthetic speeches with the same (real human speech) context. Each version contributes 0.5k of the 1.5k test turns.
Figure 4: Confusion matrices and macro-F1 (above) for the perceptual labels on 1.5k test responses. Both SFT and GRPO stages improve macro-F1 over the base Qwen2.5-Omni. Predict Mode is included as a baseline, which always outputs the most frequent label.