Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target speech without explicit temporal planning, making duration control an unresolved challenge. We introduce DuraS2ST, a duration-aligned reasoning framework that enables a single speech language model to first generate an explicit chain-of-thought (CoT) for planning target wording and phonetic length, and then synthesize the corresponding speech tokens. To support this paradigm, we construct DuraSet-440K, a high-quality duration-aligned CoT corpus for supervised initialization. We further optimize the model with multi-modal multi-dimensional reinforcement learning, using a Duration Margin Reward to balance translation quality and duration consistency, and Modality-Aware Reward Attribution to assign rewards to appropriate token spans. Experiments on CVSS-T show that DuraS2ST achieves a strong balance between translation quality and duration consistency, outperforming competitive open-source and commercial baselines. Project page: https://github.com/Mia11939/DuraS2ST.
Figures & tables
Figure 1: Existing systems suffer from duration mismatch: “你说得对” can be translated as either “You’re right” or “I think you are absolutely right about that,” leading to different speech durations. See Appendix D .
Figure 2: Overview of the DuraS2ST framework. DuraS2ST first generates a CoT rationale ( yCoT ) for explicit duration planning, then synthesizes target speech tokens ( yTA4 ). The model is trained via a two-phase paradigm: SFT on DuraSet-440K, followed by GRPO with a multi-dimensional reward design that integrates translation quality, duration (DMR), and format compliance, dynamically attributed via MARA.
Figure 3: Overview of the DuraSet-440K construction pipeline.
Model
Translation Quality
Voice & Prosody
Duration Consistency
Text-BLEU ↑
Speech-BLEU ↑
COMET ↑
COMET Kiwi ↑
AVG BLEU ↑
AVG COMET ↑
A.PCP ↑
SECS ↑
SLC-0.2 ↑
SLC-0.4 ↑
MADE (s) ↓
MRDE ↓
General Speech Large Language Models
GPT-4o
–
23.13
0.710
0.692
–
0.701
2.49
0.018
0.442
0.836
1.59
0.246
Qwen2.5-Omni
10.86
10.82
0.638
0.611
10.84
0.625
1.75
0.142
0.190
0.355
7.85
1.310
Kimi-Audio
19.87
13.91
0.698
0.700
16.89
0.699
2.10
0.188
0.429
0.746
2.56
0.462
Step-Audio-2-mini
25.42
20.58
0.752
0.743
23.01
0.745
2.49
0.511
0.312
0.681
1.98
0.348
Table 1: Main results on CVSS-T ( ZH-EN ). AVG BLEU and AVG COMET are means of (Text-BLEU, Speech-BLEU) and (COMET, COMET Kiwi ). Best / 2nd–3rd scores are highlighted. “–” indicates unavailable values.
Model
Translation Quality
Voice & Prosody
Duration Consistency
Text-BLEU ↑
Speech-BLEU ↑
COMET ↑
COMET Kiwi ↑
AVG BLEU ↑
AVG COMET ↑
A.PCP ↑
SECS ↑
SLC-0.2 ↑
SLC-0.4 ↑
MADE (s) ↓
MRDE ↓
General Speech Large Language Models
GPT-4o
–
31.46
0.845
0.780
–
0.812
2.59
0.051
0.565
0.858
1.09
0.253
Qwen2.5-Omni
15.72
15.47
0.706
0.693
15.60
0.700
2.17
0.082
0.272
0.414
4.78
1.259
Kimi-Audio
27.65
19.13
0.820
0.768
23.39
0.794
2.28
0.096
0.571
0.830
3.21
0.923
Step-Audio-2-mini
32.47
26.66
0.812
0.762
29.57
0.785
2.75
0.330
0.471
0.852
1.58
0.423
Table 2: Main results on CVSS-T ( EN-ZH ). AVG BLEU and AVG COMET are means of (Text-BLEU, Speech-BLEU) and (COMET, COMET Kiwi ). Best / 2nd–3rd scores are highlighted. “–” indicates unavailable values.
Figure 4: Zero-GRPO training dynamics on CVSS-T. Curves show duration, BLEU, COMET Kiwi , and total rewards over 1,000 GRPO steps without DuraSet-440K SFT.
Model
Trans. ↑
Natur. ↑
SpkSim ↑
Avg ↑
SeamlessM4T-v2-Large
2.87 ∣ 3.07
1.20 ∣ 3.93
1.13 ∣ 1.27
1.73 ∣ 2.76
Qwen2.5-Omni
3.87 ∣ 3.80
4.17 ∣ 4.03
1.00 ∣ 2.40
3.01 ∣ 3.41
Kimi-Audio
2.20 ∣ 2.40
3.80 ∣ 3.32
1.07 ∣ 1.13
2.36 ∣ 2.29
UniSS
3.80 ∣ 3.67
3.87 ∣ 3.33
3.87 ∣ 3.93
3.85 ∣ 3.64
DuraS2ST-Baseline
3.93 ∣ 3.53
4.10 ∣ 4.13
4.00 ∣ 4.33
4.01 ∣ 4.00
DuraS2ST-Instruct
4.00 ∣ 4.06
4.20 ∣ 4.08
4.13 ∣ 4.00
4.11 ∣ 4.05
Table 3: Subjective MOS ( 1 – 5 , higher is better) on CVSS-T. Each cell reports EN-ZH ∣ ZH-EN . Best / 2nd–3rd per direction are highlighted.
Model
SLC-0.2 ↑
SLC-0.4 ↑
MADE (s) ↓
MRDE ↓
Baseline
0.487 ∣ 0.305
0.851 ∣ 0.680
2.04 ∣ 2.22
0.629 ∣ 0.396
+ GRPO
0.455 ∣ 0.302
0.877 ∣ 0.691
1.15 ∣ 1.97
0.259 ∣ 0.346
w/o Rdur
0.448 ∣ 0.319
0.855 ∣ 0.641
1.83 ∣ 1.93
0.565 ∣ 0.385
Table 4: Zero-GRPO duration consistency on CVSS-T (EN-ZH ∣ ZH-EN). Removing Rdur isolates the effect of Duration Margin Reward without DuraSet-440K SFT.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Pre-rollout difficulty filtering improves GRPO stability and efficiency. On ZH → EN, two GRPO runs share an identical reward composition, rollout budget, and hyperparameters, differing only in their training data. Across all four reward components, the pre-rollout-filtered run (green) climbs faster and reaches a higher plateau than the unfiltered baseline (red); the in-legend Δ reports the improvement from the start of training to step 700.
Figure 6: Duration statistics of DuraSet-440K. (a) Source-duration histogram by direction. (b) Hex-binned joint density of source and target durations with the dtgt=dsrc diagonal. (c) Distribution of the target-to-source duration ratio dtgt/dsrc .
Figure 7: Prompt template for translation candidate construction. The prompt encourages stylistically diverse translations with different lexical densities and temporal footprints for the same source utterance.
Figure 8: Prompt template for reasoning rationale construction. The prompt guides the LLM to compare translation choices against deterministic phoneme lengths before outputting the final translation.
Figure 9: Qualitative EN → ZH test samples. For each sample we contrast a representative baseline output against the output of DuraS2ST, aligned to the source utterance. The baseline systems exhibit visible duration drift relative to the source’s temporal envelope, while DuraS2ST tracks the source duration tightly and produces a translation faithful to the reference.
Figure 10: Qualitative ZH → EN test samples. For each sample we contrast a representative baseline output against the output of DuraS2ST, aligned to the source utterance. As in the EN → ZH case (Figure 9 ), DuraS2ST closely matches the source’s duration and produces translations consistent with the reference, whereas the baseline drifts in duration and content.
Figure 11: Subjective evaluation interface used for the MOS study (Section 4.4 ). For every task, the annotator listens to a source clip and a synthesized target clip whose system identity is replaced by an anonymized label, then rates the target on three 1 – 5 Likert scales—translation adequacy, naturalness, and speaker similarity. The presentation order within each direction is randomized per annotator to mitigate position bias.
Long-form streaming speech-to-speech translation (S2ST) requires incremental, unbounded translation under strict latency constraints. Existing methods typically suffer from sentence-bounded supervision or demand massive paired-S2ST supervision. We introduce a training recipe enabling a speech language model for sentence-level and long-form streaming S2ST using only ∼2k hours of paired cross-lingual S2ST data, layered atop auxiliary supervision. Anchored by auxiliary multitask training, our approach remains robust even when the paired-S2ST budget itself is reduced by 90%. Our core contribution, joint text-code trajectory supervision, schedules target text and acoustic semantic codes as a unified commitment path, eliminating the need for separate, unstable speech-side emission controllers. Furthermore, our two-stream Thinker--Talker factorization significantly outperforms unified-decoder baselines by decoupling linguistic reasoning from dense acoustic prediction to mitigate modality interference. Finally, our system achieves highly competitive quality-latency trade-offs on RealSI and ACL60/60-dev, matching state-of-the-art, closed-source S2ST systems such as LiveInterpret~2.0 on ASR-BLEU.
In LLM-based speech translation, transcription-based chain-of-thought (CoT) suffers from a mismatch between reference transcripts used in supervised fine-tuning (SFT) and model-generated transcripts at inference. To address this, we propose joint recognition and translation fine-tuning via group relative policy optimization (GRPO). We score both transcripts and translations, with translation conditioned on model-generated transcripts, and compare three token advantage strategies. Using Qwen2.5-Omni-3B across four languages, we evaluate CoT against direct speech translation (Direct ST) under SFT and GRPO, training on CoVoST 2 and testing on CoVoST 2 and FLEURS. CoT GRPO outperforms Direct ST GRPO by 1.77 and 0.83 average BLEU points on CoVoST 2 and FLEURS. Compared to CoT SFT, GRPO boosts BLEU by 0.82 and 0.67 points and reduces word error rate (WER) by 8.8% and 7.2% relatively. These results highlight reinforcement fine-tuning as an effective method to mitigate the training-inference mismatch, jointly improving recognition and translation.
Yanghe Dong, Wanting Huang, Weiran Wang
Independent Researcher · Department of Computer Science, University of Iowa, USA
Speech-to-speech translation (S2ST) has advanced rapidly, but offline evaluation lacks a unified protocol: studies report non-overlapping metric subsets, preventing direct comparisons. We introduce COMPASS, a unified and reproducible benchmarking framework integrating 46 metrics across eight dimensions, and deploy it on 1,248 model-language configurations from FLEURS and CVSS, spanning cascaded and end-to-end architectures over ten language pairs. Architectures exhibit complementary strengths: best-vs-worst gaps exceed 30% on naturalness and speaker preservation but remain within a few points on translation quality, so single-metric rankings systematically misrepresent system quality. Correlation filtering reduces 46 metrics to 10 per direction, with three axes requiring different metrics across X→EN and EN→X (e.g., TER/UTMOS vs. ChrF++/NISQA-MOS); these subsets preserve rankings (Spearman's ρ>0.80) while cutting evaluation time by ≈2.5×. Human validation across dubbing, podcasts, and medical domains shows standalone MOS predictors fail to predict listener preference, while top domain-specific metrics correlate with human judgment (ρ≥0.90). We release COMPASS as a foundation for domain-aware S2ST evaluation.