Grapheme-To-Phoneme

Momentum

18 papers in the last four weeks, up 157% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 117

All topics
CardsList
  1. Child-Adapted Structured Phonological Representations for Interpretable Speech Sound Analysis

    Oct 1, 2026Abner Hernandez, Tomás Arias Vergara, Andreas Maier +1Grapheme-To-PhonemeSpeaker

  2. Cross-Linguistic Effects in Bilingual Phoneme BabyLMs

    Sep 29, 2026Nikitas Theodoropoulos, Maria Lymperaiou, Giorgos FilandrianosLanguage AcquisitionCross-Lingual Consistency

  3. Do Audio Language Models Hear and Read Distinctive Features Alike?

    Sep 24, 2026Yuanhao Chen, Peter ChinGrapheme-To-PhonemeAudio Understanding

  4. BranchShine-CR: Compact Multilingual IPA Transcription with Self-Conditioned CTC and Consistency Regularization

    Sep 24, 2026Nikhil Navas, Sergio Chevtchenko, Talisson Damiao +1Grapheme-To-PhonemeRegularization

  5. PTC-Bias: Phoneme-Level Temporal Competition for Bias Retrieval and Post-Decoding Correction in Speech LLMs

    Sep 23, 2026Zhiqi Ai, Han Cheng, Shiyi Mu +2Grapheme-To-PhonemeImplicit Association Test

  6. A Native-Reference Coordinate Geometry for L2 Pronunciation Deviation Using Self-Supervised Speech Models

    Sep 23, 2026Tina Raissi, Nhan Phan, Mikko KurimoPronunciation AssessmentGrapheme-To-Phoneme

  7. Phonemizing User-Generated Text: A Benchmark, Taxonomy, and Compositional Approach

    Sep 23, 2026MinJu Jeon, Younghan Park, Han Sung Park +3Grapheme-To-PhonemeText Analysis and Detection

  8. Structure Before Sampling: Community-Aware Core-Set Selection for Data-Efficient Text-to-Speech

    Sep 21, 2026Mizbaul Haque Maruf, Muhammad Nur YanhaonaGrapheme-To-PhonemeText Corpora

  9. P2Flow: Phoneme-aware Progressive Flow Matching for Extreme Speech Super-Resolution

    Sep 21, 2026Ningyuan Yang, Yize Li, Pu Zhao +5Speech EnhancementReal-World Image Super-Resolution

  10. Per-Aetiology Contrastive Severity Embeddings with Phonological Pseudo-Labelling for Multilingual Dysarthric Speech

    Sep 18, 2026Bernard Muller, Antonio Armando Ortiz Barrañón, LaVonne RobertsDysarthric SpeechGrapheme-To-Phoneme

  11. Multi-Teacher Distillation for Cross-Domain Streaming Electrolaryngeal Speech Encoding

    Sep 16, 2026Benedikt Mayrhofer, Enrique Orozco Olivares, Franz Pernkopf +2Teacher-Student DistillationOpencode

  12. T-SANDHI: Tone Sandhi-aware Adaptive Network with Decoupled Hybrid Injection for Low-resource Taiwanese Hokkien Speech Recognition

    Sep 16, 2026Hung-Yang Sung, Chien-Chun Wang, Tien-Hong Lo +3Grapheme-To-PhonemeMandarin

  13. Towards Stress-Aware Sentence-Level Filipino G2P With Weakly-Supervised ByT5 Fine-Tuning

    Sep 9, 2026Lorenz Bernard Marqueses, Paulo Grane Gabriel Silva, Chastine Cabatay +2Grapheme-To-PhonemeModel Fine-Tuning

  14. Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis

    Sep 9, 2026Hong Nguyen, Sean Foley, Christina Hagedorn +4Vocal Tract ShapeSpeaker

  15. KABURI-TTS: Phoneme-Keyed Activity-conditioned Bi-channel Utterance Rendering for Interaction

    Sep 7, 2026Ryuichiro Higashinaka, Shinnosuke Takamichi, Tetsuji OgawaFull-Duplex Speech ModelsF5-Tts

  16. Opinionated, Hesitant and Stressed: Three Studies of How Politicians Speak in Four Slavic Parliaments

    Aug 31, 2026Ivan Porupski, Nikola LjubešićParliamentary DebatesParalinguistic Cues

  17. Vocal Music under Phoneme-Conditional Analysis

    Aug 31, 2026Hayoon Kim, Kyogu LeeVocalizationsGrapheme-To-Phoneme

  18. Cluster Assignments in Soft Targets Shape Speech Representations: Evidence from S-JEPA

    Aug 19, 2026Wenxuan He, Yunpeng Li, Zewei Li +4Self-Supervised Speech ModelsDiscrete Speech Representations

  19. Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia

    Aug 13, 2026Xiang Guan, Roger D. Newman-Norlund, Yong Yang +8DissociationLinguistics

  20. Edge Phoneme Recognition for Children's Speech through Age-Aware Training

    Aug 10, 2026Matthew Arboleda, Ryan Arboleda, Sophie Haak +6Grapheme-To-PhonemeAutomatic Speech Recognition

  21. Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification

    Aug 10, 2026Abner Hernandez, Tomás Arias Vergara, Daiqi Liu +2ArticulatoryGrapheme-To-Phoneme

  22. ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure

    Aug 8, 2026Zixiang Wan, Xusheng Yang, Zheng Wang +1Neural Speech CodecsSelf-Supervised Speech Models

  23. Pun Intended: Multi-Agent Translation of Wordplay with Contrastive Learning and Phonetic-Semantic Embeddings

    Aug 5, 2026Russell Taylor, Benjamin Herbert, Michael SanaLinguisticsGrapheme-To-Phoneme

  24. Analyzing Speech Condition Effects in Dysarthric ASR: A Layer-wise Probing Study

    Aug 3, 2026Darwin Jelestin Muthu, Navya Gupta, Wei Lin Tay +3Dysarthric SpeechAutomatic Speech Recognition

  25. Latent Softmax for Data-Efficient Phoneme-Based Multilingual ASR Across Tonal and Non-Tonal Languages

    Aug 2, 2026Saierdaer Yusuyin, Nanling Jiang, Hao Huang +1Multilingual Automatic Speech RecognitionGrapheme-To-Phoneme

  26. Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text

    Jul 29, 2026Lucas Zamora Vera, Jose A. Gonzalez-LopezGrapheme-To-PhonemeLanguage Modeling

  27. MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation

    Jul 29, 2026Wei-Jaw Lee, Hsuan-Yu Yeh, Ting-Yi Hu +3Singing Voice ConversionMelody

  28. Evaluation of forced alignment of code-mixed speech: the case of Hindi-English

    Jul 28, 2026Ayushi Pandey, Pamir Gogoi, Kevin TangGrapheme-To-PhonemeIndian Languages

  29. Multi-Phonation Graph Learning with Self-Supervised Speech Embeddings for ALS Detection and Progression Prediction

    Jul 28, 2026Behrad TaghiBeyglou, Fatemeh Bagheri, Ervin SejdicDysarthric SpeechWav2Vec

  30. BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis

    Jul 25, 2026Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran +1Subword TokenizationIndian Languages

  31. Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin

    Jul 23, 2026Zhiheng Qian, Aini Li, Hai Hu +1Motion-Language AlignmentGrapheme-To-Phoneme

  32. Constrained CTC Decoding for Efficient Diacritic Restoration

    Jul 21, 2026Rufael Marew, Amr Keleg, Hanan AldarmakiGrapheme-To-Phoneme

  33. Self-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation Scoring

    Jul 15, 2026Stephen McIntosh, Reuben Smit, Daisuke Saito +2Grapheme-To-PhonemeProsody

  34. Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion

    Jul 14, 2026Ben Maman, Frank Zalkow, Hans-Ulrich Berendes +3Modern Generative Audio ModelsVoice Conversion

  35. FreyaTTS Technical Report

    Jul 10, 2026Ahmet Erdem Pamuk, Ömer Yentür, Ahmet Tunga Bayrak +2F5-TtsGrapheme-To-Phoneme

  36. Generative Testing of Automated Speech Recognition Systems

    Jul 10, 2026Yanis Xabier Wilbrand Peña, Oliver Weißl, Andrea StoccoAutomatic Speech RecognitionGrapheme-To-Phoneme

  37. Phone Segmentation and Recognition through Phonological Activation Mapping

    Jul 10, 2026Shikhar Bharadwaj, Kwanghee Choi, Stephen McIntosh +8Self-Supervised Speech ModelsGrapheme-To-Phoneme

  38. Why Do You Say It Like That? A Phoneme-level Framework for Explainable Speech Deepfake Detection

    Jul 9, 2026Anna Taylor, Michele Panariello, Massimiliano Todisco +3Audio Deepfake DetectionSpoofing

  39. GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech

    Jul 2, 2026Antonis Asonitis, Francesco Verdini, Aref Farhadipour +4Flow-Matching Text-To-SpeechPronunciation Assessment

  40. Towards a Phonology-Informed Evaluation of Multilingual TTS

    Jul 2, 2026Sneha Ray Barman, Neeraj Kumar Sharma, Shakuntala MahantaGrapheme-To-PhonemeMultilingual Automatic Speech Recognition

  41. YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese

    Jul 1, 2026Ryota Mibayashi, Hiroya Takamura, Hitomi YanakaGrapheme-To-Phoneme

  42. UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling

    Jun 30, 2026Chuanbo Zhu, Wuyou Zhou, Rongxiu Zhong +4Audio EditingSpeaker

  43. wav2VOT: Automatic estimation of voice onset time, closure duration, and burst realisation with wav2vec2

    Jun 27, 2026James Tanner, Morgan Sonderegger, Jane Stuart-Smith +2Wav2VecGrapheme-To-Phoneme

  44. BackTranslation2.0 -- A Linguistically Motivated Metric to Assess Sign Language Production

    Jun 27, 2026Oliver Cory, Maksym Ivashechkin, Karahan Sahin +6Sign Language TranslationMultilingual Benchmark

  45. Phonological Perception of Sign Language Models

    Jun 27, 2026Kayo Yin, Jessica Carter, Alex Xijie Lu +1Sign Language TranslationGrapheme-To-Phoneme

  46. Phonetic and semantic analyses of spoken corpora of Beijing and Taiwan Mandarin indicate that the neutral tone is a lexical tone

    Jun 24, 2026Yuxin Lu, Zhexuan Li, R. Harald BaayenMandarinGrapheme-To-Phoneme

  47. Layer-wise Probing of wav2vec 2.0 and Whisper for Consonant Cluster Reduction in African American English

    Jun 22, 2026Hamid Mojarad, Kevin TangWav2VecGrapheme-To-Phoneme

  48. Bridging Self-Supervised Learning and Speech Enhancement: A Wav2Vec2-Conditioned Framework

    Jun 21, 2026Shuubham Ojha, Carol Espy-WilsonSpeech EnhancementWav2Vec

  49. Using Phonological-Level Wav2Vec2 for Mandarin Automatic Mispronunciation Detection and Diagnosis

    Jun 20, 2026Jinghao Chen, Mostafa Shahin, Beena AhmedGrapheme-To-PhonemeWav2Vec