Discrete Speech Representation Learning

Latest papers 31

All topics
CardsList
  1. Interpreting and Evaluating Dynamic-Rate Speech Codec Boundaries

    Sep 29, 2026Han Wang, Jiaqi Li, Yingda Shen +2Neural Audio CodecsAudio Tokenization

  2. SPEAR-Gen: Generation-Aware Pre-training for Unified Speech Representations

    Sep 28, 2026Xiaoyu Yang, Arthur Hinsvark, Antonios Alexos +3Representation LearningSelf-Supervised Speech Representation Learning

  3. Beyond Encoder Fusion: Multi-View Discrete Token Augmentation for LLM-Based ASR

    Sep 20, 2026Paul Moïse Gangbadja, Mickael Rouvier, Fabrice LefèvreMulti-View LearningAutomatic Speech Recognition

  4. Kraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning

    Sep 14, 2026Hayato Futami, Hassan Shahmohammadi, Tushar Dhyani +5Speech TranslationSpeech Language Models

  5. StreamAlign: Streaming Text-Aligned Speech Tokenization

    Sep 9, 2026Kang-wook Kim, Jinyoung Park, Jinsoo Kim +3Audio TokenizationSpeech Language Models

  6. BiMTokenizer: Preserving Semantic-Acoustic Balance in Low-Bitrate Speech Tokenization via Bidirectional State-Space Modeling

    Sep 1, 2026Xin Zhang, Lin Li, Chuanbo Liu +2Neural Audio CodecsVector Quantization

  7. Cluster Assignments in Soft Targets Shape Speech Representations: Evidence from S-JEPA

    Aug 19, 2026Wenxuan He, Yunpeng Li, Zewei Li +4Soft-Label LearningUnsupervised Clustering

  8. ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure

    Aug 8, 2026Zixiang Wan, Xusheng Yang, Zheng Wang +1Neural Audio CodecsAudio Tokenization

  9. Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens

    Jul 28, 2026Daigo Takizawa, Tomohiko Nakamura, Samuele Cornell +3Audio Representation LearningNeural Audio Codecs

  10. Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

    Jul 21, 2026Laurin Wagner, Bernhard Thallinger, Miroslav Stankovic +1Audio TokenizationSelf-Supervised Speech Representation Learning

  11. AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis

    Jul 17, 2026Zhenqi Jia, Yuan Zhao, Aruukhan +2TTS SynthesisEmotional Speech Synthesis

  12. UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling

    Jun 30, 2026Chuanbo Zhu, Wuyou Zhou, Rongxiu Zhong +4Audio EditingSpeech Editing

  13. HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models

    Jun 26, 2026Artem Ploujnikov, Francesco Verdini, Samir Sadok +1Neural Audio CodecsMultimodal Token Compression

  14. wav2tok 2.0: Scalable Audio Tokenization Maintaining Explicit Pairwise Token Alignment for Efficient Audio Retrieval

    Jun 25, 2026Adhiraj Banerjee, Vipul AroraAudio Representation LearningSpeech Processing

  15. On the Effect of Segmentation Width and Cluster Size on Speech Resynthesis and Continuation in Generative Spoken Language Models

    Jun 22, 2026Shunsuke Kando, Wataru Nakata, Shinnosuke Takamichi +1Speech ProcessingSpeech Language Models

  16. Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal

    Jun 18, 2026Syeda Faiza Ahmed Sara, Shammur Absar ChowdhurySpeech ProcessingDiscrete Speech Representation Learning

  17. From Tokens to Faces: Investigating Discrete Speech Representations for 3D Facial Animation

    Jun 11, 2026Pedro Correa, Olivier Perrotin, Samir Sadok +2Audio Representation LearningNeural Audio Codecs

  18. Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation

    Jun 10, 2026Zhen Ye, Xu Tan, Yiming Li +10Spoken Language UnderstandingAudio Representation Learning

  19. ViP-VL: Vietnamese Self-supervised Speech Pretraining Model with Vector-Quantization Learning

    Jun 9, 2026Khanh Le, Kiet Anh Hoang, Bao Nguyen +5Audio Representation LearningSpeech Processing

  20. End-to-End Training for Discrete Token LLM based TTS System

    Jun 8, 2026Changfeng Gao, Yong Ren, Jun Yuan +3TTS SynthesisSpeech Processing

  21. Leveraging Soft Distributions of SSL-Derived Discrete Speech Tokens for Downstream Inference

    Jun 5, 2026Kentaro Onda, Satoru Fukayama, Daisuke Saito +1Automatic Speech RecognitionAudio Tokenization

  22. Multilingual Multi-Speaker Unit Vocoders: A Systematic Analysis of Discrete Speech Representations

    Jun 4, 2026Naman Kothari, Arjun Gangwar, Adarsh Arigala +1TTS SynthesisSpeech Processing

  23. UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception

    May 29, 2026Yuhan Song, Linhao Zhang, Aiwei Liu +6Audio Representation LearningAudio Understanding

  24. MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables

    May 28, 2026Sung-Lin Yeh, Wei Zhou, Gil Keren +6Audio Representation LearningAutoregressive Language Modeling

  25. AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling

    May 11, 2026Jiacheng Shi, Hongfei Du, Xinyuan Song +3Audio Representation LearningNeural Audio Codecs

  26. PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization

    May 7, 2026Adhiraj Banerjee, Vipul AroraAudio TokenizationSelf-Supervised Speech Representation Learning

  27. LLM-Codec: Neural Audio Codec Meets Language Model Objectives

    Apr 20, 2026Ho-Lam Chung, Yiming Chen, Hung-yi LeeNeural Audio CodecsMulti-Token Prediction