cs.CL · 2606.03948 Copy arXiv ID · Jun 2, 2026 Save A Pocket Offline Model for Simultaneous Speech Translation as CUNI Submission to IWSLT 2026 Authors: Aziz Sharipov Ortega , Dominik Macháček
Organizations: Charles University, MFF, ÚFAL · & University of Edinburgh
Abstract We implement simultaneous translation capability with the offline direct speech-to-text translation model Canary, using the state-of-the-art policy AlignAtt, and submit it to IWSLT 2026 Simultaneous Speech Translation Shared task for Czech to English and English to German and Italian. The strengths of our system are: (1) high translation quality, outperforming similarly sized baselines both in low- and high-latency regimes in computationally unaware simulations; (2) low computational requirements, as the model has only 1B parameters; (3) multilinguality -- support of 25 source and 25 target languages.
Explore similar work Jun 15, 2026 · Jorge Iranzo-Sánchez, Gerard Mas-Mollà, Adrià Giménez +3 Simultaneous Speech Translation Domain-Specific Language
Date pending · Roman Koshkin, Jeon Haesung, Lianbo Liu +4 Simultaneous Speech Translation Speech Translation
Apr 16, 2025 · Biao Fu, Donglei Yu, Minpeng Liao +5 Simultaneous Speech Translation
Jun 15, 2026 · cs.CL J/K move · Enter open · S save
Jorge Iranzo-Sánchez, Gerard Mas-Mollà, Adrià Giménez, Jorge Civera +2
Machine Learning and Language Processing, VRAIN, Universitat Politècnica de València
This work describes the participation of the MLLP-VRAIN research group in the shared task of the IWSLT 2026 Simultaneous Speech Translation track. Our submission utilizes the recently released Parakeet and Qwen 3.5 models to create a robust, cascaded solution for long-form SimulST through the use of adaptive "black-box" policies. We explore relaxations of these policies to achieve better quality-latency trade-offs. Compared to last year, we participate on all language directions. In addition to this, for the En
→ \rightarrow → {De, It, Zh} directions we also participate in this year's new context track employing a combination of ASR word-boosting and a RAG mechanism of offline pre-translated exemplars to guide generation and enrich our system with domain-specific context. Finally, we provide a detailed latency analysis of our system. Compared to last year, results on the MCIF En
→ \rightarrow → De test set shows a substantial quality improvement of +5.82 XCOMET-XL. Our context track processing further improves performance by +1.03.