Autoregressive decoding (AR) with Transformer models is memory bandwidth bound at single stream inference, the typical deployment regime for on device text to speech (TTS). Real time streaming with Qwen3-TTS requires more than 200 sequential model calls per second, dominated by the inner loop MultiCodeDecoder that emits the 15 residual vector quantization (RVQ) codes per 80 ms audio frame. We propose RVQ position aware speculative decoding for the MultiCodeDecoder, attaining 2.47 accepted tokens per model call at 5×10−4 percent added parameters and 10 to 20 percent per round speculation/verification overhead, reducing real time synthesis from 200 to 88 sequential model calls per second. The scheme is distributionally lossless under the deployed top-k sampling, and WER parity with the original system is consistent with this guarantee. We deliver 2 to 2.2x speedup for RVQ token generation with Qwen3-TTS 0.6B on recent iPhone and Apple Silicon Mac devices.
Figures & tables
Stage
Dataset
Language
Samples
Hours
Frames
Training
VoxPopuli
English
155,513
∼ 462
20.8M
Validation
VoxPopuli
English
17,200
∼ 56
2.53M
Calibration
LibriSpeech
English
1,804
∼ 4.4
200K
Evaluation
Fleurs
English
1,242
∼ 3.0
135K
German
874
∼ 2.5
112K
Spanish
1,216
∼ 4.3
192K
Table 1: Dataset specifications. Stage-wise corpora used to train the position-aware residual offsets, calibrate the sparse tree, and measure acceptance. English VoxPopuli [ 26 ] is split into training and validation; LibriSpeech [ 24 ] provides the held-out set on which the sparse tree configuration is selected; Fleurs [ 9 ] supplies multilingual evaluation across six languages. Frames count CodeDecoder steps at 80 ms/frame, each frame with 15 MultiCodeDecoder steps.
Drafter
Toks/step
Added params
Reused LM heads
1.70
—
Medusa heads [ 5 ]
2.60
∼94 M ( ∼+16% )
Ours (pos. offsets)
2.47
3,072 ( ∼+5×10−4% )
Table 2: Medusa-style drafter ablation. Three drafter choices on top of the frozen Qwen3-TTS-0.6B backbone. (i) Reused LM heads : the K=3 neighbouring codebook heads LMHk+1,…,LMHk+3 draft from the position- k hidden state; no new parameters. (ii) Medusa heads [ 5 ] : K=3 fresh Medusa heads per codebook position ( 45 new heads total). (iii) Position-aware offsets (ours): K=3 residual bias vectors b1,…,bK added to the hidden state; LM heads frozen and reused as drafters.
Language
Audio (min)
Autoregressive
Speculative
Toks/step
WER/CER
Toks/step
WER/CER
English
180.2
1.00
4.25
2.45
3.99
German
149.0
1.00
6.4
2.46
5.58
Spanish
255.6
1.00
7.28
2.44
5.98
Russian
139.8
1.00
9.55
2.46
8.26
Japanese
136.2
1.00
8.55
2.46
7.71
Table 3: Per-language performance on Fleurs. Acceptance rate ( Toks/step ) and WER (CER for Mandarin and Japanese), in percent, of the generated audio for the autoregressive baseline and our speculative decoder, evaluated on Fleurs subsets in six languages. WER is computed via WhisperKit [ 14 ] , using whisper-large-v3-turbo [ 25 ] for Mandarin and Japanese and parakeet-v3 [ 22 ] otherwise. Transcription and WER evaluation use our open source OpenBench suite [ 2 , 13 ] . As shown in prior work [ 17 , 7 ] , speculative decoding is distributionally lossless; the WER columns are consistent with this.
Speculative decoding speeds up generation by letting a cheap draft propose several tokens that a target model checks in one pass. In the single-model form, the draft is a lightweight module attached to the target rather than a separate model. Applying this design to Automatic Speech Recognition (ASR) introduces an extra problem. The draft can read the whole audio at every step, yet its proposals get worse as it runs on its own. Access is not localization. The accepted text keeps the transcript position explicit, but the draft must also track the changing audio position. In the primary matched comparison, per-step audio access changes the first proposal modestly but roughly doubles later-proposal acceptance. Fixed-width windows show that the audio position explains part of this gap. A correctly placed window recovers continuation, while an equally narrow window at the wrong position reduces it. Late-draft median error reaches 21 frames in the hardest reported condition, while target attention during verification stays within a 2-frame median. We test two ways to reduce this drift. The first reads the audio position from verification attention and uses it to guide the next draft round. It saves time only when the extra accepted tokens offset the readout cost. The second is AnchorDraft, which teaches the draft to track the audio position during training without changing the inference graph. The trained draft improves end-to-end speed at both tested target scales. These results show that ASR self-speculation depends on token prediction, audio-position tracking, and draft cost.
Codec-based autoregressive (AR) speech language models have achieved strong text-to-speech (TTS) quality by modeling speech as sequences of discrete audio tokens with large pretrained backbones. However, this token-level formulation creates a structural efficiency bottleneck: speech-token sequences are much longer than text sequences, requiring the AR backbone to perform causal computation at every token position and maintain a KV cache that grows with the sequence length. We introduce TLDR, a patch-based autoregressive framework that accelerates codec-based AR-TTS by shifting the causal modeling from token-level speech sequences to patch-level sequences. TLDR groups consecutive codec tokens into compact latent patches using a lightweight compressor, models the resulting shorter patch sequence with a frozen pretrained AR-TTS backbone adapted by LoRA, and reconstructs fine-grained speech tokens within each patch using a speaker-conditioned extractor. With a patch size of 4, TLDR achieves a 1.8x inference speedup over the baseline AR-TTS model and reduces global KV-cache memory by up to 75%. Experimental results indicate that patch-level global causal modeling can be a practical way to reduce the inference cost of pretrained codec-based AR-TTS systems without replacing the existing modules.
Autoregressive (AR) text-to-speech (TTS) models generate discrete speech tokens sequentially, which makes inference slow and can degrade robustness by propagating local errors and hallucinations. This limitation stems from their left-to-right AR commitment: each token must be determined before future speech-token context is available. However, such ordering is not an inherent requirement for TTS, as the full input text is available before synthesis. In this paper, we introduce DELTA-TTS, a lightweight LoRA-based adaptation framework that converts a pretrained AR TTS model into a discrete diffusion language model (dLLM) for confidence-ordered speech-token decoding. To better capture the local structure of speech, DELTA-TTS incorporates a convolution module that injects local acoustic context, together with a 1/t-weighted training objective and a time-shifted inference schedule that defer low-confidence positions to later steps. Trained on only 585 hours of LibriTTS, DELTA-TTS achieves a 1.75% WER on Seed-TTS test-en, outperforming its AR backbone while generating tokens 3.3× faster. Further analysis shows that DELTA-TTS produces sharper text--speech alignment, increases overall decoding confidence, and mitigates hallucinations observed in AR generation.