cs.CLJun 13, 2026

A Practical Evaluation Method for Long-Form Simultaneous Speech-to-Speech Translation

Authors: Yulin XueSiqi OuyangLei Li

Organizations: Carnegie Mellon University

Abstract

Simultaneous speech-to-speech translation (SimulS2ST) enables real-time cross-lingual communication, but existing evaluation has focused largely on short or pre-segmented speech rather than long-form, continuous input. Prior approaches are difficult to reproduce and make assumptions that do not hold for end-to-end systems. We present a practical evaluation method for long-form SimulS2ST. Given source speech, pre-segmented source transcripts, and reference translations, we run automatic speech recognition (ASR) and forced alignment on the generated target speech to recover token-level timestamps, then apply a sentence-embedding-based aligner to match the target text to its corresponding source sentences. This enables sentence-level computation of latency and quality metrics, including YAAL and xCOMET, which are then aggregated into final system-level scores. Experiments on representative SimulS2ST systems show that the method is effective in practice and reveal that current systems suffer from substantial latency accumulation on long speech.

Explore similar work

CardsList
  1. Benchmarking Speech-to-Speech Translation Models

    Jun 2, 2026Alkis Koudounas, Hayato Futami, Quentin Jodelet +3Multilingual BenchmarkTranslation Quality