Speech-To-Text Alignment

Momentum

4 papers in the last four weeks, level with the four weeks before. 0.0% of all new papers.

Jul 6Week of Sep 21

Latest papers 40

All topics
CardsList
  1. Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing

    Sep 30, 2026Rui Liu, Bhavin Jawade, Haoqi Li +4Precise Lip SynchronizationSpeech-To-Text Alignment

  2. DualTrack: Synchronized speech-gesture generation via symmetric coupling of pretrained priors

    Sep 29, 2026Yuanzhuo Hu, Zehan Liu, Xiaoyi Qin +1Co-Speech Gesture GenerationSpeech-To-Text Alignment

  3. FuseAlign: Forced Alignment in the Wild

    Sep 27, 2026Mithilesh Vaidya, Stephen Bailey, Sumukh Badam +2Speech-To-Text AlignmentTranscript

  4. ComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios

    Sep 24, 2026Jiaran Cai, Xingpei Ma, Shenneng HuangPrecise Lip SynchronizationAudio-Visual Consistency

  5. Do Audio Representations Compose Additively?

    Sep 23, 2026Chenhao Xue, Zhijin Guo, Joyraj Chakraborty +2Audio UnderstandingSpeech-To-Text Alignment

  6. TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models

    Aug 30, 2026Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh +3Large Audio Language ModelsNeural Audio

  7. The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset

    Aug 11, 2026Rajmund Nagy, Silvia Arellano García, Hendric Voss +4Co-Speech Gesture GenerationSpeech-To-Text Alignment

  8. SubtleTalk: Generating Controllable Weakly-correlated Facial Dynamics for 3D Talking Heads via Residual Flow Matching

    Aug 3, 2026Chenyang Ding, Shuai Tan, Qunfen Lin +3Precise Lip SynchronizationFacial Expression Recognition

  9. Joint Text-Audio Alignment for EEG-to-Text Decoding in Chinese Speech Production and Perception

    Jul 28, 2026Tian Zheng, Xurong Xie, Xinxin Zhu +2Electroencephalography DecodingElectroencephalography

  10. Dense-Sparse Dynamic Time Warping for Customizing Piano Concerto Accompaniments

    Jul 20, 2026TJ Tsai, Kavi Dey, Yigitcan Ozer +1AccompanimentInstrument

  11. Local Multimodal Music Alignment from Global Supervision

    Jul 10, 2026Irmak Bukey, Zachary Novack, Jongmin Jung +2Multimodal AlignmentMultimodal Contrastive Learning

  12. When Synthetic Speech Is All You Have: Better Call GRPO

    Jul 9, 2026Shashi Kumar, Yanis Labrak, Hasindri Watawana +5Flow-GrpoSpeech-To-Text Alignment

  13. Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs

    Jul 7, 2026Albert Zeyer, Ralf Schlüter, Hermann NeySpeech-To-Text AlignmentEnd-To-End Automatic Speech Recognition Models

  14. A Flexible Encoding Model for Non-Unique Note Alignments

    Jun 26, 2026Suhit Chiruthapudi, Adam Štefunko, Silvan Peter +3Symbolic Music GenerationSpeech-To-Text Alignment

  15. MixProLAP: Mixture-Induced Uncertainty Modeling for Probabilistic Language-Audio Pretraining

    Jun 18, 2026Yu Nakagome, Jaesong Lee, Soo-Whan ChungSpeech-To-Text AlignmentLarge Audio Language Models

  16. Montreal Forced Aligner and the state of speech-to-text alignment in 2026

    Jun 16, 2026Michael McAuliffe, Kaylynn Gunter, Michael Wagner +1Brain AlignmentSpeech-To-Text Alignment

  17. AP-GRPO: Anchor-Gated Phonetic Alignment with Policy Optimization for Pathological Speech Reconstruction

    Jun 14, 2026Pengfei Zhang, Hoang H Nguyen, Yutong Song +6Dysarthric SpeechSpeech-To-Text Alignment

  18. Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation

    Jun 10, 2026Zhen Ye, Xu Tan, Yiming Li +10Speech-To-Text AlignmentLarge Audio Language Models

  19. Snapping Matters: Context-Aware Onset Refinement for Automatic Music Transcription

    Jun 10, 2026Abhirup Saha, Hans-Ulrich Berendes, Meinard Müller +1Multi-Instrument Music TranscriptionSpeech-To-Text Alignment

  20. DeRA-MOS: Optimizing Text-to-Music Evaluation via Decoupled Listwise Ranking and Modality Alignment

    Jun 8, 2026Chien-Chun Wang, Hung-Shin Lee, Hsin-Min Wang +1Text-To-MusicSeed-Tts-Eval Benchmark

  21. Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation

    May 31, 2026Zhicheng Zhang, Lei Wang, Yu Zhang +1Speech-To-Text AlignmentSeed-Tts-Eval Benchmark

  22. Native Audio-Visual Alignment for Generation

    May 28, 2026Longbin Ji, Guan Wang, Xuan Wei +6Audio-Video GenerationAudio-Visual Consistency

  23. BEAT: Rhythm-Elastic Alignment for Agentic Music-guided Movie Trailer Generation

    May 26, 2026Yutong Wang, Yunke Wang, Xinyuan Chen +1Audio-Visual ConsistencyMelody

  24. Precise and Simple Audio-to-Score Alignment

    May 19, 2026Silvan Peter, Patricia Hu, Gerhard WidmerSpeech-To-Text AlignmentTranscription

  25. HighSync: High-Quality Lip Synchronization via Latent Diffusion Models

    May 16, 2026Saeid Firouzi Daghigh, Majid Iranpour Mobarakeh, Mostafa Alavi +1Precise Lip SynchronizationSpeech-To-Text Alignment

  26. UMo: Unified Sparse Motion Modeling for Real-Time Co-Speech Avatars

    May 14, 2026Xiaoyu Zhan, Xinyu Fu, Chenghao Yang +9Co-Speech Gesture GenerationSpeech-To-Text Alignment

  27. Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR

    May 14, 2026Ryo Magoshi, Takashi Maekaku, Yusuke ShinoharaSpeech-To-Text AlignmentFlow-Matching Text-To-Speech

  28. AViS-Mamba: Adaptive Visual Steering of Audio State-Space Dynamics for Violence Detection

    Apr 2, 2026Damith Chamalke Senadeera, Dimitrios Kollias, Gregory SlabaughAudio-Visual ReasoningFine-Grained Video Understanding

  29. SyncEdit: Rethinking Lip Synchronization as Editing with Audio-Driven Diffusion Models

    Mar 10, 2026Lixiang Lin, Siyuan Jin, Jinshan ZhangPrecise Lip SynchronizationAudio-Visual Consistency

  30. CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment

    Feb 23, 2026Hanwen Liu, Saierdaer Yusuyin, Hao Huang +1F5-TtsSpeech-To-Text Alignment

  31. TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding

    Jan 11, 2026Mingyue Huo, Yiwen Shao, Yuheng ZhangSpeaker DiarizationSpeech-To-Text Alignment

  32. Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation

    Dec 23, 2025Jingqi Tian, Yiheng Du, Haoji Zhang +6Audio-Visual ConsistencyVideo Object Segmentation

  33. M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR

    Oct 25, 2025Ruixiang Mao, Xiangnan Ma, Qing Yang +7Speech-To-Text AlignmentGrapheme-To-Phoneme

  34. Streaming Translation and Transcription Through Speech-to-Text Causal Alignment

    Date pendingRoman Koshkin, Jeon Haesung, Lianbo Liu +4Speech TranslationTranscription