Zero-shot text-to-speech (TTS) can reproduce an unseen speaker from a short reference recording, but typically entangles speaker identity and accent within the same reference. We introduce DEFINE, an end-to-end framework that decouples these factors by conditioning speaker identity and target accent on separate audio exemplars. A single inference-time guidance weight continuously controls accent strength without retraining. Built on F5-TTS with parameter-efficient LoRA adaptation, DEFINE maps short accent exemplars into a conditioning space using an exemplar encoder supervised through learned accent prototypes, requiring neither accent labels at inference time nor post-synthesis waveform conversion. On seen accents, increasing accent guidance improves accent-probe accuracy from 6.5% to 19.6%. More importantly, a single DEFINE model generalizes accent control beyond its training accent set: on seen and out-of-domain accents, though not on held-out accents, it matches the accent transfer performance of a two-model TTS-voice-conversion cascade while achieving higher speaker similarity and comparable predicted speech quality. These results demonstrate that speaker identity and accent can be independently controlled from audio exemplars within a single zero-shot TTS model, including for accents unseen during training.
Figures & tables
Figure 1: Define overview. (a) Accent exemplars are encoded by a frozen XLS-R and pooled into one accent vector; (b) a learned per-accent prototype table supervises the encoder during training. The accent vector shifts the timestep and text embeddings of a frozen F5-TTS backbone adapted with LoRA. At inference the prototypes are discarded and a single weight w sets accent strength. Snowflake: frozen, flame: trained.
System
ACC ↑ (95% CI)
SPK ↑
WER ↓
UTMOS ↑
Seen accents , n=1134 items
Real speech (reference)
0.436
[0.406, 0.464]
--
--
3.08
F5-TTS (stock)
0.079
[0.063, 0.096]
0.732
0.071
3.99
Cascade F5-TTS → Seed-VC
0.188
[0.166, 0.211]
0.600
0.056
3.92
Define , accent off
0.065
[0.051, 0.079]
0.694
0.063
4.06
Define (exemplar)
0.196
[0.173, 0.219]
0.624
0.067
3.93
Table 1: Accent control with 95% bootstrap intervals over items. Define is reported at w=15 , selected on validation items disjoint from the evaluation grid; it infers the accent from exemplar clips and uses no accent label. † Prototype lookup is instead given the accent label and reads the corresponding row of the learned prototype table; it therefore cannot address accents outside the training set; it is also shown at w=11 , the operating point of the listening study. The reference row is real speech by a different speaker reading different text, so SPK and WER against the prompt and target sentence do not apply ( -- ).
Seen accents
Held-out accents
Out-of-domain
w
ACC ex
SPK
ACC lk
ACC ex
SPK
ACC ex
0
0.065
0.694
0.065
0.066
0.698
0.048
7
0.156
0.667
0.272
0.079
0.683
0.048
9
0.168
0.655
0.311
0.098
0.675
0.056
11
0.185
0.644
0.326
0.108
0.669
0.063
15
0.196
0.624
0.354
0.127
0.654
0.087
Table 2: Guidance sweep. w=0 disables the accent term; the LoRA adapters remain active, so this is not identical to F5-TTS; larger w pushes further toward the accent. ACC ex and SPK are exemplar-driven; ACC lk is the prototype-lookup bound and exists for seen accents only.
Variant
w∗
ACC
SPK
UTMOS
Define (full)
15
0.127
0.654
4.01
− contrastive prototype term
9
0.074
0.662
4.04
− prototype anchoring (free encoder)
0
0.079
0.711
4.03
− flow-matching gradient into the accent path
0
0.074
0.717
4.01
+ encoder-consistency loss
0
0.074
0.698
4.04
cross-attention conditioning
0
0.085
0.693
4.01
Table 3: Ablation on the 378 held-out-accent set. Each variant is reported at the largest guidance weight w∗ keeping UTMOS within 0.15 and speaker similarity within 0.10 of the same variant with the accent switched off, selected on validation items: variants tolerate guidance very differently, so a common w would reward one that degrades into noise. A variant whose w∗ is 0 gains no accent control at any weight; its high speaker similarity and predicted quality simply reflect the accent being switched off.
Figure 2: Listening test on seen accents. All Define samples use prototype-lookup, i.e. accents given as a label, at w=11 , where the accent change is clearly perceptible. (a) Define against F5-TTS and the cascade; (b) the same model at three guidance weights. 11 listeners, 10 trials per part, giving 110 votes per part. Bars show the share of votes each option received, with 95% Wilson intervals.