cs.CLOct 8, 2026

Prosody-to-Text: Predicting text from low-pass filtered speech

Authors: David Porteš, Aleš Horák

Organizations: Natural Language Processing Centre Faculty of Informatics, Masaryk University Botanicka 68a, 602 00 Brno, Czech Republic

Abstract

While predicting prosody from text is an established task in the field, the opposite direction, predicting text that fits a given prosodic pattern, remains largely overlooked. We find this unfortunate, because this opposite direction could lead to some very interesting use cases. Therefore, in this paper, we make the first steps in the prosody-to-text direction by inves- tigating how much of the original sentence can be recovered from its prosodic pattern. To this end, we fine-tune the Whis- per model using only the 12 lowest Mel bins (low-pass filter with approximately 450Hz cutoff), and obtain surprisingly accurate results (WER 36%), with 10% of utterances be- ing recovered perfectly, and 40% of utterances having Word Error Rate at or below 25%. We also find that, given the correct prefix, the next token was predicted correctly in 79% of cases. Our results suggest that the relationship between low-frequency speech features and lexical content is much stronger than previously thought, and we believe that direct- ing more attention to this topic might open the door to new applications, such as using prosody to guide text generation of modern LLMs

Figures & tables

Explore similar work

CardsList
  1. Dynamic Prosody Prediction in LLM-based TTS for Improving Speaker Similarity

    Jun 13, 2026Zhenwei Mou, Liping Chen, Yajun Hu +3TTS SynthesisSpeech Prosody

  2. Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM

    May 7, 2026Wenqian Cui, Xiao-Hui Li, Daxin Tan +2Audio UnderstandingSpeech Prosody

  3. Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters

    Oct 4, 2026Abdul Rehman, Jian-Jun Zhang, Xiaosong YangTTS SynthesisSpeech Processing