cs.SDSep 27, 2026

DuraS2ST: Chain-of-Thought and Reinforcement Learning for Duration-Aligned Speech-to-Speech Translation

Authors: Yayue Deng, Dingdong Wang, Yuxuan Hu, Jinyu Li, Yanqing Liu, Yuanyuan Wang, Weidong Chen, Helen M. Meng, +2 more

Organizations: The Chinese University of Hong Kong, China · Microsoft Corporation · Work done during an internship at Microsoft

Abstract

Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target speech without explicit temporal planning, making duration control an unresolved challenge. We introduce DuraS2ST, a duration-aligned reasoning framework that enables a single speech language model to first generate an explicit chain-of-thought (CoT) for planning target wording and phonetic length, and then synthesize the corresponding speech tokens. To support this paradigm, we construct DuraSet-440K, a high-quality duration-aligned CoT corpus for supervised initialization. We further optimize the model with multi-modal multi-dimensional reinforcement learning, using a Duration Margin Reward to balance translation quality and duration consistency, and Modality-Aware Reward Attribution to assign rewards to appropriate token spans. Experiments on CVSS-T show that DuraS2ST achieves a strong balance between translation quality and duration consistency, outperforming competitive open-source and commercial baselines. Project page: https://github.com/Mia11939/DuraS2ST.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision

    Jul 22, 2026Rongshen He, Xinyu Liang, Dekun Chen +3Speech TranslationSpeech Language Models

  2. Transcribe, Translate, and Optimize: Joint Reward Learning for Speech Translation

    Sep 22, 2026Yanghe Dong, Wanting Huang, Weiran WangSpeech TranslationSpeech Language Models

  3. Benchmarking Speech-to-Speech Translation Models

    Jun 2, 2026Alkis Koudounas, Hayato Futami, Quentin Jodelet +3Speech TranslationMultilingual Benchmark