This work presents SelfTTS, a text-to-speech (TTS) model designed for cross-speaker style transfer that eliminates the need for external pre-trained speaker or emotion encoders. The architecture achieves emotional expressivity in neutral speakers through an explicit disentanglement strategy utilizing Gradient Reversal Layers (GRL) combined with cosine similarity loss to decouple speaker and emotion information. We introduce Multi Positive Contrastive Learning (MPCL) to induce clustered representations of speaker and emotion embeddings based on their respective labels. Furthermore, SelfTTS employs a self-refinement strategy via Self-Augmentation, exploiting the model's voice conversion capabilities to enhance the naturalness of synthesized speech. Experimental results demonstrate that SelfTTS achieves superior emotional naturalness (eMOS) and robust stability in target timbre and emotion compared to state-of-the-art baselines.
Figures & tables
Fig. 1: SelfTTS model architecture. The emotion and speaker encoders receive mel-spectrogram slices of the reference waveform as input. Each encoder is optimized using the MPCL loss, while their embeddings are disentangled through a cosine-based GRL applied on top of the Linear Processor output for each encoder. The final forward step of the Normalizing Flows ( zp ) is also disentangled using a cosine-based GRL, through a Convolutional Processor that predicts the corresponding emotion or speaker embedding. SDP stands for Stochastic Duration Predictor . Purple dashed arrows indicate the proposed Self-Augmentation pipeline.
Fig. 2: SelfTTS architecture at inference time. The model follows the VITS inference pipeline augmented with disentangled speaker and emotion embeddings. Both embeddings can be obtained either from a reference audio or from post-hoc analysis of the training style space via centroid prototypes.
Model
Speaker Encoder
Emotion Encoder
Subjective Evaluation
Objective Evaluation
nMOS ↑
eMOS ↑
sMOS ↑
UTMOS ↑
WER ↓
SECS ↑
EECS ↑
GT
-
-
3.638±0.106
3.541 ± 0.122
4.009 ± 0.107
3.6560
0.0718
0.8297
0.9659
E3-VITS
Lookup
StyleSpeech
2.763 ± 0.111
2.237 ± 0.108
2.815 ± 0.116
3.9614
0.1222
0.8069
0.5367
VECL
HSAP
CNN-based
2.707 ± 0.107
2.556 ± 0.108
2.804 ± 0.114
3.7959
0.1146
0.7806
0.6929
SelfTTS w/o Self-Aug.
RE
RE
2.228 ± 0.105
2.849 ± 0.108
3.121 ± 0.112
3.4899
0.1952
0.8103
0.8793
SelfTTS
RE
RE
2.746 ± 0.109
2.853 ± 0.107
3.112 ± 0.114
3.6104
0.1636
0.8163
0.8423
TABLE I: Subjective and objective performance comparison between SelfTTS and baseline models. Subjective metrics include a 95% confidence interval. For this and the subsequent tables, the following visualization pattern will be adopted: ↑ higher is better. ↓ lower is better. Best values are bolded , second-best values are underlined and proposed model is highlighted in green.
Fig. 3: Emotional similarity (eMOS) per emotion category across different models. The models are ordered as: GT (ground-truth), SelfTTS, SelfTTS w/o Self-Aug., VECL, and E3-VITS.
Fig. 4: UMAP projections of the emotion style spaces for all evaluated models. Embeddings are colored by emotion: Neutral (blue), Happy (green), Angry (red), Sad (pink), and Surprise (orange). The emotion style space of SelfTTS clearly forms well-separated emotional clusters. VECL generates highly concentrated clusters, but there is still overlap between different emotions. E3-VITS is not capable of producing clustered representations.
Encoder Loss
GRL
TTS
VC
CKA ↓
LK-CKA
UTMOS ↑
WER ↓
SECS ↑
EECS ↑
UTMOS ↑
WER ↓
SECS ↑
EECS ↑
Speaker ↑
Emotion ↑
CE
None
3.8585
0.1521
0.8064
0.7391
3.8324
0.0798
0.8120
0.7061
0.3148
0.8704
0.6964
MPCL
3.6268
0.1734
0.8165
0.7903
3.6968
0.0866
0.8213
0.7425
0.0176
0.9588
0.9648
CE
CE
3.6931
0.2046
0.8019
0.7276
3.7239
0.0896
0.8087
0.7043
0.0235
0.9373
0.9282
MPCL
3.4996
0.1665
0.7704
0.7899
3.6247
0.1013
0.7722
0.7545
0.0336
0.9514
0.9458
CE
Cosine
3.4737
0.1775
0.8075
0.8474
3.5507
0.0965
0.8095
0.8107
0.0205
0.9376
0.9024
TABLE II: Objective metrics of TTS and VC samples generated varying the encoder loss (MPCL × CE) and Gradient Reversal Layer configuration (None GRL applied, Cross-Entropy loss and proposed Cosine loss). The table also reports speaker-emotion embedding representations similarity (CKA) and speaker and emotion label alignment (LK-CKA).
Emotion Encoder
20 Mel Bins
TP
UTMOS ↑
WER ↓
SECS ↑
EECS ↑
RE
×
×
3.4899
0.1952
0.8103
0.8793
×
✓
3.5171
0.1768
0.8196
0.7771
✓
×
3.5367
0.1703
0.8102
0.7500
✓
✓
3.7208
0.2612
0.8080
0.5933
StyleSpeech
×
×
3.6500
0.1993
0.8114
0.7621
×
✓
3.7438
0.1676
0.8063
0.6439
TABLE III: Objective performance comparison of different emotion encoder configurations and input perturbation methods. ✓ indicates the feature is used, while × indicates it is not.
Module
Mode
LK-CKA Speaker
LK-CKA Emotion
Flow Step 1
Forward
0.4830
0.0004
Flow Step 2
Forward
0.4849
0.0006
Flow Step 3
Forward
0.2517
0.0014
Flow Step 4
Forward
0.0270
0.0032
Flow Step 5
Reverse
0.0916
0.1193
Flow Step 6
Reverse
0.1273
0.1810
TABLE IV: LK-CKA analysis per module for speaker and emotion label alignment in the VC path (SelfTTS).
Module
Mode
LK-CKA Speaker
LK-CKA Emotion
Flow Step 1
Forward
0.5324
0.0005
Flow Step 2
Forward
0.5207
0.0007
Flow Step 3
Forward
0.2692
0.0015
Flow Step 4
Forward
0.0552
0.0030
Flow Step 5
Reverse
0.1306
0.1218
Flow Step 6
Reverse
0.1653
0.1703
TABLE V: LK-CKA analysis per module for speaker and emotion label alignment in the VC path (w/ CE GRL).
Self Aug.
UTMOS ↑
WER ↓
SECS ↑
EECS ↑
config
GT
3.2509
0.1784
0.7950
0.8896
ENC
3.6461
0.1554
0.8162
0.8027
BOTH
3.4108
0.1524
0.8024
0.8425
TABLE VI: Objective performance comparison of different Self-Augmentation configurations.
Self Aug.
UTMOS ↑
WER ↓
SECS ↑
EECS ↑
proportion
0.00
3.5335
0.1538
0.8150
0.8778
0.25
3.6104
0.1636
0.8163
0.8423
0.50
3.6461
0.1554
0.8162
0.8027
0.75
3.6738
0.1538
0.8226
0.7651
1.00
3.7663
0.1564
0.8155
0.6624
TABLE VII: Impact of different Self-Augmentation proportions on objective metrics.
Model
Speaker
UTMOS ↑
WER ↓
SECS ↑
EECS ↑
SelfTTS w/o Self Aug.
p226
3.1898
0.5464
0.8964
0.6583
p231
3.1937
0.5358
0.8714
0.6890
LJ
3.1800
0.3482
0.8355
0.6133
SelfTTS
p226
3.3806
0.5176
0.9114
0.6074
p231
3.3806
0.5399
0.8914
0.6125
LJ
3.3951
0.3327
0.8587
0.6006
TABLE VIII: Objective performance comparison across LJspeech, p226 and p231 speakers from LJspeech and VCTK in cross-corpus scenario for SelfTTS.
Audiobook narration, conversational agents, and audiovisual dubbing require speech that conveys changing emotions and adapts its pacing within a single utterance. But most existing TTS systems typically rely on utterance-level style conditioning, making such fine-grained control difficult to achieve. In light of this, and inspired by the success of post-training in large language models, we propose a unified post-training framework that equips pretrained text-to-speech models with natural-language control over segment-level emotion and duration. Supervised fine-tuning establishes instruction-conditioned speech generation, while reinforcement learning with group relative policy optimization refines control accuracy using emotion and duration rewards alongside content and speaker preservation objectives. By reusing the pretrained architecture, our approach avoids additional inference-time control modules. Experiments demonstrate significantly improved fine-grained controllability while maintaining speech intelligibility and speaker identity, highlighting post-training as a practical approach to extending existing speech synthesis models.
Lianru Gao, Yujie Guo, Yong Qin
TMCC, College of Computer Science, Nankai University, Tianjin, China
For the last couple of years, the field of speech synthesis has improved dramatically thanks to deep learning. There are more and more deep learning-based TTS systems developed to make it possible to produce voices with high intelligibility and naturalness. Meanwhile, controlling the expressiveness is yet a big deal, generating speech in different styles or manners has received a lot of attention from community recently. This paper aims to give our solutions to deal with the task emotional speech synthesis (ESS) at VLSP 2022 which allows to generate humanlike natural-sounding voice from a given input text with desired emotional expression. By integrating speaker embedding, prosody bottleneck into FastSpeech 2, our systems can promisingly generate emotional speech of a single speaker (Sub-task 1), transfer speaking styles from another speaker to the target speaker with neutral non-expressive data while retaining the target speaker's identity (Sub-task 2).
Building state-of-the-art text-to-speech (TTS) systems typically demands millions of hours of proprietary data and complex multi-stage architectures, creating substantial barriers for resource-constrained research teams. In this report, we present PilotTTS, a lightweight autoregressive TTS system that achieves competitive performance through minimalist architecture and rigorous data engineering. PilotTTS is trained on only 200K hours of data processed entirely with open-source tools. Specifically, our contributions are: (1) a reproducible multi-stage data processing pipeline covering quality assessment, label annotation, and filtering, and (2) a compact model architecture that employs Q-Former-based conditioning to decouple speaker identity from speaking style via cross-sample paired training. Within a unified framework, PilotTTS supports zero-shot voice cloning, emotion synthesis (11 categories), paralinguistic synthesis (4 categories), and Chinese dialect synthesis (14 dialects). On the Seed-TTS Eval benchmark, PilotTTS achieves the lowest WER of 1.50% on test-en, a CER of 0.87% on test-zh, and the highest speaker similarity on both test sets (0.862 and 0.815), outperforming systems trained on significantly larger datasets. We release the complete data pipeline recipe, pretrained weights, and code at https://github.com/AMAPVOICE/PilotTTS.
Bowen Li, Shaotong Guo, Zhen Wang +11
Amap, Alibaba Group · The Chinese University of Hong Kong, Shenzhen