Autoregressive decoding (AR) with Transformer models is memory bandwidth bound at single stream inference, the typical deployment regime for on device text to speech (TTS). Real time streaming with Qwen3-TTS requires more than 200 sequential model calls per second, dominated by the inner loop MultiCodeDecoder that emits the 15 residual vector quantization (RVQ) codes per 80 ms audio frame. We propose RVQ position aware speculative decoding for the MultiCodeDecoder, attaining 2.47 accepted tokens per model call at 5×10−4 percent added parameters and 10 to 20 percent per round speculation/verification overhead, reducing real time synthesis from 200 to 88 sequential model calls per second. The scheme is distributionally lossless under the deployed top-k sampling, and WER parity with the original system is consistent with this guarantee. We deliver 2 to 2.2x speedup for RVQ token generation with Qwen3-TTS 0.6B on recent iPhone and Apple Silicon Mac devices.
Figures & tables
Stage
Dataset
Language
Samples
Hours
Frames
Training
VoxPopuli
English
155,513
∼ 462
20.8M
Validation
VoxPopuli
English
17,200
∼ 56
2.53M
Calibration
LibriSpeech
English
1,804
∼ 4.4
200K
Evaluation
Fleurs
English
1,242
∼ 3.0
135K
German
874
∼ 2.5
112K
Spanish
1,216
∼ 4.3
192K
Table 1: Dataset specifications. Stage-wise corpora used to train the position-aware residual offsets, calibrate the sparse tree, and measure acceptance. English VoxPopuli [ 26 ] is split into training and validation; LibriSpeech [ 24 ] provides the held-out set on which the sparse tree configuration is selected; Fleurs [ 9 ] supplies multilingual evaluation across six languages. Frames count CodeDecoder steps at 80 ms/frame, each frame with 15 MultiCodeDecoder steps.
Drafter
Toks/step
Added params
Reused LM heads
1.70
—
Medusa heads [ 5 ]
2.60
∼94 M ( ∼+16% )
Ours (pos. offsets)
2.47
3,072 ( ∼+5×10−4% )
Table 2: Medusa-style drafter ablation. Three drafter choices on top of the frozen Qwen3-TTS-0.6B backbone. (i) Reused LM heads : the K=3 neighbouring codebook heads LMHk+1,…,LMHk+3 draft from the position- k hidden state; no new parameters. (ii) Medusa heads [ 5 ] : K=3 fresh Medusa heads per codebook position ( 45 new heads total). (iii) Position-aware offsets (ours): K=3 residual bias vectors b1,…,bK added to the hidden state; LM heads frozen and reused as drafters.
Language
Audio (min)
Autoregressive
Speculative
Toks/step
WER/CER
Toks/step
WER/CER
English
180.2
1.00
4.25
2.45
3.99
German
149.0
1.00
6.4
2.46
5.58
Spanish
255.6
1.00
7.28
2.44
5.98
Russian
139.8
1.00
9.55
2.46
8.26
Japanese
136.2
1.00
8.55
2.46
7.71
Table 3: Per-language performance on Fleurs. Acceptance rate ( Toks/step ) and WER (CER for Mandarin and Japanese), in percent, of the generated audio for the autoregressive baseline and our speculative decoder, evaluated on Fleurs subsets in six languages. WER is computed via WhisperKit [ 14 ] , using whisper-large-v3-turbo [ 25 ] for Mandarin and Japanese and parakeet-v3 [ 22 ] otherwise. Transcription and WER evaluation use our open source OpenBench suite [ 2 , 13 ] . As shown in prior work [ 17 , 7 ] , speculative decoding is distributionally lossless; the WER columns are consistent with this.