Much of expressive text-to-speech research rests on an untested assumption that written text carries enough information to select an appropriate prosodic style for its delivery. Text-predicted style models improve listener preference, and expressive-appropriateness evaluation presupposes that context constrains style, yet neither measures the assumption itself. This paper tests it as a falsifiable hypothesis against style labels derived from acoustics alone. For each of six speakers in a 1,200-hour conversational corpus, utterances are clustered in the spaces of five speech models, including a prosody-only control, and the cluster of held-out utterances is predicted from twelve text embedding models. Three controls are applied: utterance length is erased from the speech embeddings; accuracy is scored against the majority-class floor of unbalanced clusters rather than uniform chance; and a bag-of-words baseline measures word identity alone. Text predicts the cluster above that floor for all six speakers (+0.111 top-3 accuracy), but bag-of-words achieves three quarters of this. Sentence embeddings add only +0.026, largest for encoders not trained for sentence semantics and reversed by tree-based probes for all others. Acoustic clusters are not compact in text embedding space in any of 360 configurations. The prosody-only space weakens the association for five speakers, but not for the speaker showing it most strongly. Text thus informs these delivery clusters mainly through word choice, whether as a cue to prosody or as a marker of topic and recording situation, and reference-free style selection cannot assume more.
Figures & tables
Fig. 1: The cross-modal evaluation framework. Utterances are clustered using speech representation models (Table I ); the cluster labels are then predicted from text embeddings of the transcriptions, and compared against length-only and bag-of-words arms that see the same labels and the same held-out utterances.
Model
Params
Input
Pretraining data
Objective
HuBERT Large [ 21 ]
300M
waveform, 16 kHz
60k h Libri-Light
masked prediction of K-means targets
HuBERT XLarge [ 21 ]
1B
waveform, 16 kHz
60k h Libri-Light
masked prediction of K-means targets
WavLM Large [ 22 ]
300M
waveform, 16 kHz
94k h mixed
masked prediction + denoising
wav2vec2 XLSR-53 [ 43 ]
300M
waveform, 16 kHz
56k h, 53 languages
contrastive
TIMBRESPACE (Section III )
52M
6 prosodic features, 62.5 Hz
860 h LibriSpeech
masked prediction of prosodic tokens
TABLE I: Speech Representation Models
Emotion
Identity
Encoder
dim
L1
L7
L1
L7
WavLM-Large
1024
0.570
0.586
0.984
0.949
wav2vec2-XLSR53
1024
0.580
0.578
0.986
0.966
HuBERT-XLarge
1280
0.537
0.582
0.983
0.935
HuBERT-Large
1024
0.545
0.526
0.982
0.941
TIMBRESPACE
768
0.417
0.397
0.814
0.834
TABLE II: Speaker-Independent Probe Performance on IEMOCAP
Emotion ηˉ2
Identity ηˉ2
Encoder
L1
L7
× null
L1
L7
ARI
WavLM-Large
0.056
0.055
21
0.808
0.144
0.009
wav2vec2-XLSR53
0.056
0.035
14
0.470
0.019
0.007
HuBERT-Large
0.058
0.036
14
0.550
0.039
0.013
HuBERT-XLarge
0.049
0.041
16
0.367
0.059
0.014
TIMBRESPACE
0.086
0.166
69
0.169
0.161
0.024
TABLE III: Unsupervised Structure on IEMOCAP
Fig. 2: Top-3 accuracy of the three text arms per speaker, averaged over the twelve text models and five speech models (error bars: standard deviation over those cells). The dashed rule is each speaker’s majority-top-3 floor; uniform chance (0.10) is shown for comparison only.
Spk
Floor
Len
TF-IDF
Sem
Sem–TFIDF
worst
F1
0.233
0.231
0.297
0.308
+0.011
−0.058
F2
0.250
0.260
0.321
0.351
+0.030
−0.037
F3
0.248
0.256
0.484
0.503
+0.019
−0.038
M1
0.309
0.272
0.330
0.352
+0.022
−0.040
M2
0.213
0.203
0.234
0.276
+0.042
+0.004
M3
0.276
0.302
0.375
0.406
+0.030
−0.020
TABLE IV: Per-Speaker Top-3 Accuracy of the Three Text Arms
All 12
MLM
Sim.
Fixed classifier
Sem–TFIDF
Spk >0
p
K-Nearest Neighbors
+0.066
6/6
.016
+0.080
+0.063
Linear Discriminant
+0.032
6/6
.016
+0.051
+0.028
Logistic Regression
+0.020
6/6
.016
+0.037
+0.017
Ridge Classifier
+0.019
6/6
.016
+0.036
+0.015
Random Forest
−0.013
0/6
1.00
+0.014
−0.020
TABLE V: Increment of Text Embeddings over Bag-of-Words With the Classifier Held Fixed
Fig. 3: Top-3 accuracy above the TF-IDF baseline for each text embedding model, averaged over speakers and speech models. Highlighted: the two MLM encoders, included as weak-encoder controls. para-multiling-MiniLM-L12 and para-distilroberta-v2 abbreviate paraphrase-multilingual-MiniLM-L12-v2 and paraphrase-distilroberta-base-v2 .
Fig. 4: Top-3 accuracy of the three text arms per speech representation model, averaged over speakers and the twelve text models, against the corresponding majority-top-3 floor (dashed).
Fig. 5: Top-3 accuracy above TF-IDF per speech representation model. Open circles are individual speakers (mean over the twelve text models); filled circles are the mean over the six speakers.
Layer index
1
3
7
p
Majority-top-3 floor
0.219
0.220
0.255
.011
Embedding-arm top-3
0.343
0.361
0.389
.030
Lift over floor
0.124
0.141
0.134
.042
excluding XLSR-53
Floor
0.217
0.221
0.219
.61
Lift over floor
0.125
0.143
0.169
.030
TABLE VI: Sensitivity to Encoder Depth
Fig. 6: Top-3 accuracy above each configuration’s majority floor, per speech representation model and encoder depth (index 1 is the first seventh of the model, 7 the final layer), averaged over the six speakers and the three text models of Table VI .