cs.CLDate pending

Streaming Translation and Transcription Through Speech-to-Text Causal Alignment

Authors: Roman KoshkinJeon HaesungLianbo LiuHao ShiMengjie ZhaoYusuke FujitaYui Sudo

Abstract

Simultaneous machine translation (SiMT) has traditionally relied on offline machine translation models coupled with human-engineered heuristics or learned policies. We propose Hikari, a policy-free, end-to-end model for simultaneous speech-to-text translation and streaming transcription. We also introduce Decoder Time Dilation, a mechanism that counteracts the overrepresentation of WAIT tokens in training. We present a supervised fine-tuning strategy that trains the model to recover from delays, significantly improving the quality-latency trade-off. Despite its modest size, Hikari delivers competitive translation quality at consistently low latency, comparing favorably with published IWSLT 2026 submissions up to 38x larger and with proprietary API systems across en-ja, en-de, and en-ru. We release our model weights and code to facilitate further research.

Explore similar work

CardsList
  1. Streaming Speech-to-Text Translation with a SpeechLLM

    May 14, 2026Titouan Parcollet, Shucong Zhang, Xianrui Zheng +1Speech TranslationSpeech-Llms