Zero-shot text-to-speech (TTS) can reproduce an unseen speaker from a short reference recording, but typically entangles speaker identity and accent within the same reference. We introduce DEFINE, an end-to-end framework that decouples these factors by conditioning speaker identity and target accent on separate audio exemplars. A single inference-time guidance weight continuously controls accent strength without retraining. Built on F5-TTS with parameter-efficient LoRA adaptation, DEFINE maps short accent exemplars into a conditioning space using an exemplar encoder supervised through learned accent prototypes, requiring neither accent labels at inference time nor post-synthesis waveform conversion. On seen accents, increasing accent guidance improves accent-probe accuracy from 6.5% to 19.6%. More importantly, a single DEFINE model generalizes accent control beyond its training accent set: on seen and out-of-domain accents, though not on held-out accents, it matches the accent transfer performance of a two-model TTS-voice-conversion cascade while achieving higher speaker similarity and comparable predicted speech quality. These results demonstrate that speaker identity and accent can be independently controlled from audio exemplars within a single zero-shot TTS model, including for accents unseen during training.
Figures & tables
Figure 1: Define overview. (a) Accent exemplars are encoded by a frozen XLS-R and pooled into one accent vector; (b) a learned per-accent prototype table supervises the encoder during training. The accent vector shifts the timestep and text embeddings of a frozen F5-TTS backbone adapted with LoRA. At inference the prototypes are discarded and a single weight w sets accent strength. Snowflake: frozen, flame: trained.
System
ACC ↑ (95% CI)
SPK ↑
WER ↓
UTMOS ↑
Seen accents , n=1134 items
Real speech (reference)
0.436
[0.406, 0.464]
--
--
3.08
F5-TTS (stock)
0.079
[0.063, 0.096]
0.732
0.071
3.99
Cascade F5-TTS → Seed-VC
0.188
[0.166, 0.211]
0.600
0.056
3.92
Define , accent off
0.065
[0.051, 0.079]
0.694
0.063
4.06
Define (exemplar)
0.196
[0.173, 0.219]
0.624
0.067
3.93
Table 1: Accent control with 95% bootstrap intervals over items. Define is reported at w=15 , selected on validation items disjoint from the evaluation grid; it infers the accent from exemplar clips and uses no accent label. † Prototype lookup is instead given the accent label and reads the corresponding row of the learned prototype table; it therefore cannot address accents outside the training set; it is also shown at w=11 , the operating point of the listening study. The reference row is real speech by a different speaker reading different text, so SPK and WER against the prompt and target sentence do not apply ( -- ).
Seen accents
Held-out accents
Out-of-domain
w
ACC ex
SPK
ACC lk
ACC ex
SPK
ACC ex
0
0.065
0.694
0.065
0.066
0.698
0.048
7
0.156
0.667
0.272
0.079
0.683
0.048
9
0.168
0.655
0.311
0.098
0.675
0.056
11
0.185
0.644
0.326
0.108
0.669
0.063
15
0.196
0.624
0.354
0.127
0.654
0.087
Table 2: Guidance sweep. w=0 disables the accent term; the LoRA adapters remain active, so this is not identical to F5-TTS; larger w pushes further toward the accent. ACC ex and SPK are exemplar-driven; ACC lk is the prototype-lookup bound and exists for seen accents only.
Variant
w∗
ACC
SPK
UTMOS
Define (full)
15
0.127
0.654
4.01
− contrastive prototype term
9
0.074
0.662
4.04
− prototype anchoring (free encoder)
0
0.079
0.711
4.03
− flow-matching gradient into the accent path
0
0.074
0.717
4.01
+ encoder-consistency loss
0
0.074
0.698
4.04
cross-attention conditioning
0
0.085
0.693
4.01
Table 3: Ablation on the 378 held-out-accent set. Each variant is reported at the largest guidance weight w∗ keeping UTMOS within 0.15 and speaker similarity within 0.10 of the same variant with the accent switched off, selected on validation items: variants tolerate guidance very differently, so a common w would reward one that degrades into noise. A variant whose w∗ is 0 gains no accent control at any weight; its high speaker similarity and predicted quality simply reflect the accent being switched off.
Figure 2: Listening test on seen accents. All Define samples use prototype-lookup, i.e. accents given as a label, at w=11 , where the accent change is clearly perceptible. (a) Define against F5-TTS and the cascade; (b) the same model at three guidance weights. 11 listeners, 10 trials per part, giving 110 votes per part. Bars show the share of votes each option received, with 95% Wilson intervals.
In cross-lingual zero-shot text-to-speech, the accent of the reference leaks into the target speech. We propose accent analogy guidance (AAG), a training-free sampler term that subtracts an accent direction estimated from the model's own predictions for one synthetic voice rendered in both languages, so the voice cancels and only the accent remains. By a blind LLM accent judge on real dubbing data, reweighting classifier-free guidance between reference and text, and its variants, stay near one identity-accent trade-off curve; we score a method by its speaker similarity above that curve at equal accent (ΔSIM). Across four open TTS models AAG lies above the curve: on OmniVoice ΔSIM is +0.11 to +0.27 on three test sets (accent 3.51 to 4.28 on a 1-5 scale at speaker similarity 0.29, where reweighting keeps 0.02); MaskGCT and CosyVoice 2 also lie above their curves, and on F5-TTS it is more native than any reweighting setting. An LLM-free language-ID measure and a twelve-listener panel agree. A premise test and the reach of a model's own curve indicate in advance whether and roughly how much AAG can gain, predicting the one model where it gains nothing (X-Voice).
Accent text-to-speech (TTS) aims to synthesize speech with a target accent while preserving speaker identity, but faces two key challenges: disentangling accent from speaker characteristics and effectively conditioning speech generation on the two disentangled factors. In this paper, we propose Joycent, a diffusion-based accent TTS framework that addresses both challenges. Our key idea is to separate accent and speaker information in both representation learning and TTS conditioning. Joycent uses WhisAID, a Whisper-based accent encoder with gradient reversal to learn speaker-disentangled accent representations, and introduces layer-specific conditional layer normalization to inject accent and speaker information at different stages of the text encoder. We evaluate Joycent on the Mandarin Regional Accent Corpus (MRAC) with seen and unseen speakers, including a challenging cross-accent setting where the speaker and accent prompts come from different accents. Experimental results show that Joycent improves accent similarity over existing methods while maintaining strong speaker similarity, with consistent gains under the challenging cross-accent setting. Subjective evaluation further confirms improved naturalness, accent similarity, and speaker preservation. The audio samples are available at https://oshindow.github.io/joycent/.
Xintong Wang, Junchuan Zhao, Ye Wang
School of Computing, National University of Singapore, Singapore
Accent conversion and controllability remain fundamental challenges in cross-lingual text-to-speech (TTS), particularly for low-resource and phonetically diverse Indic languages. While recent large language model (LLM)-based TTS systems exhibit strong cross-lingual generalization, they provide limited explicit control over accent characteristics and intensity. In this paper, we propose CrossAccentTTS, a framework that enables both accent control and conversion while preserving speaker identity. Specifically, we introduce an Accent Intensity Controller (AIC) that injects weighted language embeddings into the accent subspace, allowing smooth interpolation between accents and fine-grained modulation of accent strength at inference time. Experiments on the Indic Multilingual and L2-arctic datasets shows that CrossAccent-TTS achieves precise control of accent intensity, outperforming strong baselines in accent similarity and controllability by maintaining speaker similarity and naturalness.
Ram Annamdevula, Ankit Tatawat, Ashishkumar P. Gudmalwar +2