Recent non-autoregressive (NAR) zero-shot text-to-speech (TTS) models generate in parallel but typically require the target sequence length to be specified before generation. We introduce EditVoice, to our knowledge the first variable-length NAR zero-shot TTS model, which uses Edit Flows to jointly update speech content and sequence length through insertions, deletions, and substitutions. EditVoice adopts speech-infilling training, which unifies zero-shot TTS and text-based speech editing and allows both prefix and suffix speech prompt placements at inference. We introduce Complementary Prompt Sampling (CPS) to leverage the complementary Edit Flow predictions induced by the two prompt placements. We further find that EditVoice can edit source and model-generated speech beyond its training sources. We use this generalization for end-to-end editing and training-free post-generation refinement. With the Edit Flow model trained on 10K h of GigaSpeech, EditVoice demonstrates competitive zero-shot TTS performance on Seed-TTS Eval EN and LibriSpeech-PC and speech editing performance on RealEdit.
Figures & tables
Model
Params.
Train (h)
Seed-TTS EN
LibriSpeech-PC
NMOS ↑
SMOS ↑
T2S RTF ↓
WER ↓
SIM-o ↑
UTMOS ↑
WER ↓
SIM-o ↑
UTMOS ↑
Ground Truth
–
–
–
–
–
–
–
–
3.77
4.11
–
Autoregressive
CosyVoice2 † [ 5 ]
0.5B
167K
2.58
0.659
4.17
1.80
0.655
4.39
3.88
3.84
0.5313
CosyVoice3 [ 6 ]
0.5B
1M
2.00
0.697
3.97
1.77
0.695
4.30
–
–
–
VoxCPM [ 30 ]
0.5B
1.8M
1.86
0.729
3.82
1.92
0.719
4.21
–
–
–
Table 1: Zero-shot TTS results. For EditVoice , 10K h denotes the GigaSpeech data used to train the Edit Flow model. † denotes systems using the same S3Tokenizer2 tokenizer and CosyVoice 2 detokenizer as EditVoice ; bold and underline denote the best and second-best synthesized results. T2S RTF measures only text-to-semantic-token generation.
Model
Type
WER ↓
SIM-o ↑
MOSN ↑
UTMOS ↑
Ground Truth
–
5.59
–
3.22
3.40
Cascade editing
FluentSpeech [ 13 ]
NAR
5.80
0.949
3.18
2.68
VoiceCraft [ 20 ]
AR
5.89
0.973
3.14
3.34
SSR-Speech [ 24 ]
AR
4.75
0.986
3.20
3.36
EditVoice
NAR
4.40
0.981
3.22
3.35
Table 2: Objective RealEdit results; bold and underline denote the best and second-best results within each setting.
Inference
WER ↓
SIM-o ↑
UTMOS ↑
56 generation NFE
Prefix
1.65%
0.653
4.08
Suffix
1.66%
0.656
4.08
CPS w/o Prompt Consistency
1.61%
0.658
4.08
CPS
1.54%
0.659
4.11
56 generation + 8 refinement NFE
Table 3: CPS and refinement ablation on Seed-TTS Eval EN.