Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters
Organizations: National Centre for Computer Animation, Bournemouth University, U.K.
Abstract
Much of expressive text-to-speech research rests on an untested assumption that written text carries enough information to select an appropriate prosodic style for its delivery. Text-predicted style models improve listener preference, and expressive-appropriateness evaluation presupposes that context constrains style, yet neither measures the assumption itself. This paper tests it as a falsifiable hypothesis against style labels derived from acoustics alone. For each of six speakers in a 1,200-hour conversational corpus, utterances are clustered in the spaces of five speech models, including a prosody-only control, and the cluster of held-out utterances is predicted from twelve text embedding models. Three controls are applied: utterance length is erased from the speech embeddings; accuracy is scored against the majority-class floor of unbalanced clusters rather than uniform chance; and a bag-of-words baseline measures word identity alone. Text predicts the cluster above that floor for all six speakers (+0.111 top-3 accuracy), but bag-of-words achieves three quarters of this. Sentence embeddings add only +0.026, largest for encoders not trained for sentence semantics and reversed by tree-based probes for all others. Acoustic clusters are not compact in text embedding space in any of 360 configurations. The prosody-only space weakens the association for five speakers, but not for the speaker showing it most strongly. Text thus informs these delivery clusters mainly through word choice, whether as a cue to prosody or as a marker of topic and recording situation, and reference-free style selection cannot assume more.
Figures & tables
| Model | Params | Input | Pretraining data | Objective |
|---|---|---|---|---|
| HuBERT Large [ 21 ] | 300M | waveform, 16 kHz | 60k h Libri-Light | masked prediction of K-means targets |
| HuBERT XLarge [ 21 ] | 1B | waveform, 16 kHz | 60k h Libri-Light | masked prediction of K-means targets |
| WavLM Large [ 22 ] | 300M | waveform, 16 kHz | 94k h mixed | masked prediction + denoising |
| wav2vec2 XLSR-53 [ 43 ] | 300M | waveform, 16 kHz | 56k h, 53 languages | contrastive |
| TIMBRESPACE (Section III ) | 52M | 6 prosodic features, 62.5 Hz | 860 h LibriSpeech | masked prediction of prosodic tokens |
| Emotion | Identity | ||||
|---|---|---|---|---|---|
| Encoder | dim | L1 | L7 | L1 | L7 |
| WavLM-Large | 1024 | 0.570 | 0.586 | 0.984 | 0.949 |
| wav2vec2-XLSR53 | 1024 | 0.580 | 0.578 | 0.986 | 0.966 |
| HuBERT-XLarge | 1280 | 0.537 | 0.582 | 0.983 | 0.935 |
| HuBERT-Large | 1024 | 0.545 | 0.526 | 0.982 | 0.941 |
| TIMBRESPACE | 768 | 0.417 | 0.397 | 0.814 | 0.834 |
| Emotion | Identity | |||||
|---|---|---|---|---|---|---|
| Encoder | L1 | L7 | null | L1 | L7 | ARI |
| WavLM-Large | 0.056 | 0.055 | 21 | 0.808 | 0.144 | 0.009 |
| wav2vec2-XLSR53 | 0.056 | 0.035 | 14 | 0.470 | 0.019 | 0.007 |
| HuBERT-Large | 0.058 | 0.036 | 14 | 0.550 | 0.039 | 0.013 |
| HuBERT-XLarge | 0.049 | 0.041 | 16 | 0.367 | 0.059 | 0.014 |
| TIMBRESPACE | 0.086 | 0.166 | 69 | 0.169 | 0.161 | 0.024 |
| Spk | Floor | Len | TF-IDF | Sem | Sem–TFIDF | worst |
|---|---|---|---|---|---|---|
| F1 | 0.233 | 0.231 | 0.297 | 0.308 | ||
| F2 | 0.250 | 0.260 | 0.321 | 0.351 | ||
| F3 | 0.248 | 0.256 | 0.484 | 0.503 | ||
| M1 | 0.309 | 0.272 | 0.330 | 0.352 | ||
| M2 | 0.213 | 0.203 | 0.234 | 0.276 | ||
| M3 | 0.276 | 0.302 | 0.375 | 0.406 |
| All 12 | MLM | Sim. | |||
|---|---|---|---|---|---|
| Fixed classifier | Sem–TFIDF | Spk | |||
| K-Nearest Neighbors | 6/6 | .016 | |||
| Linear Discriminant | 6/6 | .016 | |||
| Logistic Regression | 6/6 | .016 | |||
| Ridge Classifier | 6/6 | .016 | |||
| Random Forest | 0/6 | 1.00 | |||
| Layer index | 1 | 3 | 7 | |
|---|---|---|---|---|
| Majority-top-3 floor | 0.219 | 0.220 | 0.255 | .011 |
| Embedding-arm top-3 | 0.343 | 0.361 | 0.389 | .030 |
| Lift over floor | 0.124 | 0.141 | 0.134 | .042 |
| excluding XLSR-53 | ||||
| Floor | 0.217 | 0.221 | 0.219 | .61 |
| Lift over floor | 0.125 | 0.143 | 0.169 | .030 |