Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Organizations: Dolby Laboratories
Abstract
TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. Our model replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation. Across two English and two Mandarin datasets, Tacit-TTS achieves competitive zero-shot quality while generating speech over 10x faster than IndexTTS2 for utterances longer than 5 seconds. Its transcript-free conditioning further supports cross-lingual and non-lexical references. We validate this capability using references from eight other languages, infant babble, and synthetic gibberish, where transcript-dependent systems often degrade or fail due to unreliable ASR transcripts.
Figures & tables
| LibriSpeech test-clean | SeedTTS test-en | SeedTTS test-zh | AISHELL-1 test | |||||
| Model | SS | WER | SS | WER | SS | CER | SS | CER |
| Autoregressive | ||||||||
| CosyVoice2 | 0.843 | 5.999 | 0.794 | 3.277 | 0.846 | 1.451 | 0.834 | 1.967 |
| SparkTTS | 0.756 | 8.843 | 0.755 | 1.543 | 0.683 | 2.636 | 0.593 | 1.743 |
| Non-autoregressive | ||||||||
| MaskGCT | 0.790 | 7.759 | 0.824 | 2.530 | 0.807 | 2.447 | 0.598 | 4.930 |
| English targets | Chinese targets | |||||
| Model | SS | WER | Fail | SS | CER | Fail |
| Transcript-dependent | ||||||
| F5-TTS | 0.663 0.120 | 24.56 22.28 | 0.0 0.0 | 0.622 0.100 | 25.01 20.89 | 0.0 0.0 |
| MaskGCT † | 0.701 0.065 | 20.16 14.68 | 33.3 47.1 | 0.700 0.033 | 28.26 14.32 | 33.3 47.1 |
| CosyVoice2 | 0.692 0.081 | 45.95 42.64 | 0.0 0.0 | 0.680 0.093 | 62.41 47.33 | 0.0 0.0 |
| SparkTTS | 0.611 0.091 | 47.27 42.23 | 20.0 21.5 | 0.578 0.092 | 60.42 48.06 | 12.9 16.9 |
| Infant babble | Synthetic gibberish | ||||||||||
| English targets | Chinese targets | English targets | Chinese targets | ||||||||
| Model | SS | WER | SS | CER | SS | WER | SS | CER | |||
| Transcript-dependent | |||||||||||
| SparkTTS | 0.278 | 99.00 | 0.310 | 96.00 | 0.648 | 18.00 | 0.639 | 20.30 | |||
| CosyVoice2 | 0.442 | 31.03 | 0.441 | 87.53 | 0.721 | 49.24 | 0.792 | 86.94 | |||
| MaskGCT † | 0.184 | 75.63 | 0.217 | 45.67 | 0.743 | 46.84 | 0.729 | 32.81 | |||
| Latency (ms) | Speed ( real-time) | |||||||
| Model | Cond. | ASR | T2S | S2A | Total | Cond. | Gen. | Total |
| Transcript-dependent | ||||||||
| CosyVoice2 | 1082 | 918 | 3214 | 604 | 5818 | 3.03 | 1.56 | |
| MaskGCT | 125 | 918 | 2428 | 2359 | 5830 | 5.88 | 1.27 | |
| SparkTTS ‡ | 57 | 918 | 7041 | 32 | 8049 | 6.67 | 0.92 | |
| F5-TTS † | 599 | 315 | 1122 | 2036 | 5.88 | 4.76 | ||
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Prompt recording | Pseudo-word sequence (target text) |
|---|---|
| voice_01 | kalene kuguva ge yu mu nola vi pesa fa zula hukogu zuma ni na loze zela |
| voice_02 | tuneho nide so gese wohele lu fine wadape vugofu yiba likive biye kiyume |
| voice_03 | tuya sobazi de gawape wutiyu mumi zoyo ne sa nunaga ga guni wepuwe veka |
| voice_04 | vivoha kafike mu poti nahu du ge vo mode bolo vuwe re di du yuta geguge |
| voice_05 | kunu dugoyi mini pihi bo kamule gezi hu ruyara la pi mimi widavi bi le |
| voice_06 | huvuba pu zu popari ze nazuvu peteva so nitafa nehi nebe sigi yosa le |
| LibriSpeech test-clean | SeedTTS test-en | SeedTTS test-zh | AISHELL-1 test | |||||
| Model | UTMOS | DNSMOS | UTMOS | DNSMOS | UTMOS | DNSMOS | UTMOS | DNSMOS |
| Autoregressive | ||||||||
| CosyVoice2 | 4.344 | 3.348 | 4.115 | 3.253 | 3.395 | 3.382 | 2.980 | 3.251 |
| SparkTTS | 4.017 | 3.065 | 3.903 | 3.133 | 3.271 | 3.266 | 2.864 | 3.081 |
| Non-autoregressive | ||||||||
| MaskGCT | 3.903 | 3.296 | 3.484 | 3.106 | 2.549 | 3.241 | 2.343 | 3.218 |
| English | Chinese | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| LibriSpeech test-clean | SeedTTS test-en | SeedTTS test-zh | AISHELL-1 test | |||||||
| T2S | S2Mel | SS | WER | SS | WER | SS | CER | SS | CER | |
| 8 | 1 | 0.710 | 7.11 | 0.682 | 2.32 | 0.570 | 3.25 | 0.530 | 3.37 | 17.9 |
| 8 | 4 | 0.860 | 6.56 | 0.837 | 2.05 | 0.800 | 3.12 | 0.784 | 3.24 | 18.5 |
| 8 | 8 | 0.872 | 6.64 | 0.849 | 2.13 | 0.825 | 3.14 | 0.799 | 3.17 | 14.3 |
| 12 | 1 | 0.711 | 5.78 | 0.684 | 2.36 | 0.575 | 1.81 | 0.535 | 2.39 | 19.2 |
| Japanese | Korean | |||||||||||
| EN | ZH | EN | ZH | |||||||||
| Model | SS | WER | Fail | SS | CER | Fail | SS | WER | Fail | SS | CER | Fail |
| Transcript-dependent | ||||||||||||
| F5-TTS | 0.408 | 38.86 | 0.0 | 0.406 | 40.99 | 0.0 | 0.648 | 31.33 | 0.0 | 0.643 | 23.59 | 0.0 |
| MaskGCT | 0.635 | 30.07 | 0.0 | 0.656 | 30.27 | 0.0 | 0.637 | 25.28 | 0.0 | 0.677 | 30.01 | 0.0 |
| CosyVoice2 | 0.629 | 25.04 | 0.0 | 0.603 | 78.62 | 0.0 | 0.746 | 3.81 | 0.0 | 0.819 | 8.21 | 0.0 |