cs.SDSep 29, 2026

RVQ Position Aware Speculative Decoding for On Device Text to Speech

Authors: Berkin Durmus, Eduardo Pacheco, Zach Nagengast, Atila Orhon

Organizations: Argmax, Inc., USA

Abstract

Autoregressive decoding (AR) with Transformer models is memory bandwidth bound at single stream inference, the typical deployment regime for on device text to speech (TTS). Real time streaming with Qwen3-TTS requires more than 200 sequential model calls per second, dominated by the inner loop MultiCodeDecoder that emits the 15 residual vector quantization (RVQ) codes per 80 ms audio frame. We propose RVQ position aware speculative decoding for the MultiCodeDecoder, attaining 2.47 accepted tokens per model call at 5×10−45\times10^{-4} percent added parameters and 10 to 20 percent per round speculation/verification overhead, reducing real time synthesis from 200 to 88 sequential model calls per second. The scheme is distributionally lossless under the deployed top-k sampling, and WER parity with the original system is consistent with this guarantee. We deliver 2 to 2.2x speedup for RVQ token generation with Qwen3-TTS 0.6B on recent iPhone and Apple Silicon Mac devices.

Figures & tables

Explore similar work

CardsList
  1. TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech

    Jun 8, 2026Yejin Lee, Junwon Moon, Hyoeun Kim +3Autoregressive Text-To-SpeechAutoregressive Decoding

  2. DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech

    Jul 5, 2026Junwon Moon, Seungbeom Kim, Yejin Lee +4Autoregressive Text-To-SpeechDiffusion Language Models