Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model
Abstract
We distill Kokoro-82M, a widely used open text-to-speech model with 54 voices, into Paradee, an 8.07M-parameter model that speaks one of them. Paradee keeps Kokoro's architecture with much narrower layers, and each of its two halves is trained separately against the frozen teacher. It has 10x fewer parameters and needs 15x less compute. We first synthesize a corpus with the teacher and keep its durations, pitch, energy and phoneme features. We then train a small text side to predict these values, and a small decoder to turn the teacher's saved values into the teacher's audio, first with spectral losses and then adversarially. Finally, we connect the two halves and quantize the weights to int8. It needs no alignment learning and no joint training, and it runs on one laptop. Stored in int8, Paradee is 8.5 MB, runs 25x faster than real time on one CPU thread, and scores 4.41 on UTMOS against the teacher's 4.52. The student initially kept a slight buzz, which we trace to the phase of voiced speech between 2 and 8 kHz. A phase-locking filter applied after synthesis removes most of it, with no training and no extra parameters. Code, model files and audio samples are at https://github.com/sahilmahendrakar/paradee
Figures & tables
| Subsystem | Parameters | Share |
|---|---|---|
| Decoder AdaIN encode/decode blocks (1024 channels) | 33.6M | 41% |
| Decoder iSTFTNet generator | 19.7M | 24% |
| LSTMs (text encoder and prosody predictor) | 11.2M | 14% |
| ALBERT (12 shared layers, 768 hidden) and its output projection | 6.7M | 8% |
| F0, energy and duration heads | 6.6M | 8% |
| Text encoder CNN and phoneme embedding | 4.0M | 5% |
| Kokoro-82M (teacher) | Paradee (student) | |
|---|---|---|
| Parameters | 81.8M | 8.07M (10.1 fewer) |
| Weights on disk | 325 MB fp32 ONNX; 82 MB int8 | 8.45 MB int8 |
| Compute per second of audio | 55 GFLOP | 3.4 GFLOP (15 less) |
| Speed, PyTorch, 1 CPU thread | 7.6 real time | 25.0 real time |
| Decoder alone, onnxruntime, 1 CPU thread | 20.5 real time | |
| UTMOS, fp32 | 4.52 0.003 | 4.41 0.01 |
| Model (voice) | Params | File size | Speed | UTMOS | WER |
|---|---|---|---|---|---|
| Kokoro-82M, teacher (af_heart) | 81.8M | 325 MB | 7.6 | 4.52 0.003 | 5.7% |
| Paradee (af_heart) | 8.07M | 8.45 MB | 25.0 | 4.41 0.01 | 5.7% |
| Paradee, released ONNX file | 8.07M | 9.0 MB | 17.8 | 4.41 | 6.0% |
| Kokoro-7M-Distill (af_msa) | 7.48M | 30.1 MB | 35.5 | 4.18 0.03 | 7.4% |
| Piper, en_US-lessac-medium | 15.7M | 63.2 MB | 15.4 | 4.36 0.02 | 8.8% |
| KittenTTS nano 0.8 (Bella) | 14.0M | 56.8 MB | 10.7 | 4.01 0.04 | 5.5% |
| Text side, rendered by the teacher decoder | Params | DTW log-mel | UTMOS |
|---|---|---|---|
| Teacher text side | 28M | 0.12 (floor) | 4.52 |
| Student s, linear projection | 4.0M | 0.97 | 4.15 |
| Student s, MLP projection | 4.2M | 0.91 | 4.36 |
| Student m, MLP projection | 7.1M | 0.89 | 4.35 |
| Student s, MLP, trained through the teacher decoder | 4.2M | 0.99 | 4.52 |
| Decoder A (3.85M) training | Decoder alone | Full student |
|---|---|---|
| Spectral losses only, 50k steps | 2.98 | 3.00 |
| + adversarial, spectral weight 45, 3k steps | 3.02 | |
| + adversarial, spectral weight 10, 5k steps | 4.29 | 4.32 |
| + adversarial, spectral weight 3, 5k more steps (stage two) | 4.37 | 4.39 |
| + phase-locking filter (Paradee) | 4.41 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Reference point | DTW log-mel | Duration change |
|---|---|---|
| Teacher against itself, second run (random excitation) | 0.09 | 0 |
| All weights fp16 | 0.11 | 0 |
| ONNX fp32 export against PyTorch | 0.21 | 0 |
| All weights int8 per channel, activations fp32 | 0.28 | 25 ms |
| Teacher at speed 0.97 (timing change only) | 0.36 | 130 ms |
| Teacher on CPU against Apple MPS | 0.39 | 0 |
| Part | Params | 8-bit | 4-bit (groups) | 4-bit | Timing shift, 8-bit |
|---|---|---|---|---|---|
| ALBERT | 6.3M | 0.263 | 0.495 | 0.555 | 10 ms |
| Text encoder | 5.6M | 0.153 | 0.266 | 0.319 | 5 ms |
| Prosody predictor | 16.2M | 0.312 | 0.379 | 0.402 | 25 ms |
| Decoder AdaIN | 33.6M | 0.090 | 0.126 | 0.179 | 0 |
| Generator | 19.7M | 0.103 | 0.354 | 0.537 | 0 |
| All weights | 81.8M | 0.270 | 0.624 | 25 ms |
| Preset | Widths and kernels | Params | Generator | GFLOP/s audio |
|---|---|---|---|---|
| Teacher | 1024 / 512 / 3,7,11 | 53.3M | 19.7M | 51.9 |
| A (Paradee) | 256 / 128 / 3,7,11 | 3.85M | 1.19M | 3.36 |
| B | 256 / 128 / 3,7 | 3.49M | 0.83M | 2.28 |
| C | 192 / 96 / 3,7 | 2.16M | 0.48M | 1.31 |
| D | 128 / 64 / 3,7 | 1.15M | 0.23M | 0.61 |
| W2 (wider final stage) | 128 / 256 / 3,7 | 4.39M | 3.12M | 8.2 |