This work presents SelfTTS, a text-to-speech (TTS) model designed for cross-speaker style transfer that eliminates the need for external pre-trained speaker or emotion encoders. The architecture achieves emotional expressivity in neutral speakers through an explicit disentanglement strategy utilizing Gradient Reversal Layers (GRL) combined with cosine similarity loss to decouple speaker and emotion information. We introduce Multi Positive Contrastive Learning (MPCL) to induce clustered representations of speaker and emotion embeddings based on their respective labels. Furthermore, SelfTTS employs a self-refinement strategy via Self-Augmentation, exploiting the model's voice conversion capabilities to enhance the naturalness of synthesized speech. Experimental results demonstrate that SelfTTS achieves superior emotional naturalness (eMOS) and robust stability in target timbre and emotion compared to state-of-the-art baselines.
Figures & tables
Fig. 1: SelfTTS model architecture. The emotion and speaker encoders receive mel-spectrogram slices of the reference waveform as input. Each encoder is optimized using the MPCL loss, while their embeddings are disentangled through a cosine-based GRL applied on top of the Linear Processor output for each encoder. The final forward step of the Normalizing Flows ( zp ) is also disentangled using a cosine-based GRL, through a Convolutional Processor that predicts the corresponding emotion or speaker embedding. SDP stands for Stochastic Duration Predictor . Purple dashed arrows indicate the proposed Self-Augmentation pipeline.
Fig. 2: SelfTTS architecture at inference time. The model follows the VITS inference pipeline augmented with disentangled speaker and emotion embeddings. Both embeddings can be obtained either from a reference audio or from post-hoc analysis of the training style space via centroid prototypes.
Model
Speaker Encoder
Emotion Encoder
Subjective Evaluation
Objective Evaluation
nMOS ↑
eMOS ↑
sMOS ↑
UTMOS ↑
WER ↓
SECS ↑
EECS ↑
GT
-
-
3.638±0.106
3.541 ± 0.122
4.009 ± 0.107
3.6560
0.0718
0.8297
0.9659
E3-VITS
Lookup
StyleSpeech
2.763 ± 0.111
2.237 ± 0.108
2.815 ± 0.116
3.9614
0.1222
0.8069
0.5367
VECL
HSAP
CNN-based
2.707 ± 0.107
2.556 ± 0.108
2.804 ± 0.114
3.7959
0.1146
0.7806
0.6929
SelfTTS w/o Self-Aug.
RE
RE
2.228 ± 0.105
2.849 ± 0.108
3.121 ± 0.112
3.4899
0.1952
0.8103
0.8793
SelfTTS
RE
RE
2.746 ± 0.109
2.853 ± 0.107
3.112 ± 0.114
3.6104
0.1636
0.8163
0.8423
TABLE I: Subjective and objective performance comparison between SelfTTS and baseline models. Subjective metrics include a 95% confidence interval. For this and the subsequent tables, the following visualization pattern will be adopted: ↑ higher is better. ↓ lower is better. Best values are bolded , second-best values are underlined and proposed model is highlighted in green.
Fig. 3: Emotional similarity (eMOS) per emotion category across different models. The models are ordered as: GT (ground-truth), SelfTTS, SelfTTS w/o Self-Aug., VECL, and E3-VITS.
Fig. 4: UMAP projections of the emotion style spaces for all evaluated models. Embeddings are colored by emotion: Neutral (blue), Happy (green), Angry (red), Sad (pink), and Surprise (orange). The emotion style space of SelfTTS clearly forms well-separated emotional clusters. VECL generates highly concentrated clusters, but there is still overlap between different emotions. E3-VITS is not capable of producing clustered representations.
Encoder Loss
GRL
TTS
VC
CKA ↓
LK-CKA
UTMOS ↑
WER ↓
SECS ↑
EECS ↑
UTMOS ↑
WER ↓
SECS ↑
EECS ↑
Speaker ↑
Emotion ↑
CE
None
3.8585
0.1521
0.8064
0.7391
3.8324
0.0798
0.8120
0.7061
0.3148
0.8704
0.6964
MPCL
3.6268
0.1734
0.8165
0.7903
3.6968
0.0866
0.8213
0.7425
0.0176
0.9588
0.9648
CE
CE
3.6931
0.2046
0.8019
0.7276
3.7239
0.0896
0.8087
0.7043
0.0235
0.9373
0.9282
MPCL
3.4996
0.1665
0.7704
0.7899
3.6247
0.1013
0.7722
0.7545
0.0336
0.9514
0.9458
CE
Cosine
3.4737
0.1775
0.8075
0.8474
3.5507
0.0965
0.8095
0.8107
0.0205
0.9376
0.9024
TABLE II: Objective metrics of TTS and VC samples generated varying the encoder loss (MPCL × CE) and Gradient Reversal Layer configuration (None GRL applied, Cross-Entropy loss and proposed Cosine loss). The table also reports speaker-emotion embedding representations similarity (CKA) and speaker and emotion label alignment (LK-CKA).
Emotion Encoder
20 Mel Bins
TP
UTMOS ↑
WER ↓
SECS ↑
EECS ↑
RE
×
×
3.4899
0.1952
0.8103
0.8793
×
✓
3.5171
0.1768
0.8196
0.7771
✓
×
3.5367
0.1703
0.8102
0.7500
✓
✓
3.7208
0.2612
0.8080
0.5933
StyleSpeech
×
×
3.6500
0.1993
0.8114
0.7621
×
✓
3.7438
0.1676
0.8063
0.6439
TABLE III: Objective performance comparison of different emotion encoder configurations and input perturbation methods. ✓ indicates the feature is used, while × indicates it is not.
Module
Mode
LK-CKA Speaker
LK-CKA Emotion
Flow Step 1
Forward
0.4830
0.0004
Flow Step 2
Forward
0.4849
0.0006
Flow Step 3
Forward
0.2517
0.0014
Flow Step 4
Forward
0.0270
0.0032
Flow Step 5
Reverse
0.0916
0.1193
Flow Step 6
Reverse
0.1273
0.1810
TABLE IV: LK-CKA analysis per module for speaker and emotion label alignment in the VC path (SelfTTS).
Module
Mode
LK-CKA Speaker
LK-CKA Emotion
Flow Step 1
Forward
0.5324
0.0005
Flow Step 2
Forward
0.5207
0.0007
Flow Step 3
Forward
0.2692
0.0015
Flow Step 4
Forward
0.0552
0.0030
Flow Step 5
Reverse
0.1306
0.1218
Flow Step 6
Reverse
0.1653
0.1703
TABLE V: LK-CKA analysis per module for speaker and emotion label alignment in the VC path (w/ CE GRL).
Self Aug.
UTMOS ↑
WER ↓
SECS ↑
EECS ↑
config
GT
3.2509
0.1784
0.7950
0.8896
ENC
3.6461
0.1554
0.8162
0.8027
BOTH
3.4108
0.1524
0.8024
0.8425
TABLE VI: Objective performance comparison of different Self-Augmentation configurations.
Self Aug.
UTMOS ↑
WER ↓
SECS ↑
EECS ↑
proportion
0.00
3.5335
0.1538
0.8150
0.8778
0.25
3.6104
0.1636
0.8163
0.8423
0.50
3.6461
0.1554
0.8162
0.8027
0.75
3.6738
0.1538
0.8226
0.7651
1.00
3.7663
0.1564
0.8155
0.6624
TABLE VII: Impact of different Self-Augmentation proportions on objective metrics.
Model
Speaker
UTMOS ↑
WER ↓
SECS ↑
EECS ↑
SelfTTS w/o Self Aug.
p226
3.1898
0.5464
0.8964
0.6583
p231
3.1937
0.5358
0.8714
0.6890
LJ
3.1800
0.3482
0.8355
0.6133
SelfTTS
p226
3.3806
0.5176
0.9114
0.6074
p231
3.3806
0.5399
0.8914
0.6125
LJ
3.3951
0.3327
0.8587
0.6006
TABLE VIII: Objective performance comparison across LJspeech, p226 and p231 speakers from LJspeech and VCTK in cross-corpus scenario for SelfTTS.