cs.CLJan 15, 2026

HOMURA: Taming the Sand-Glass for Time-Constrained LLM Translation via Reinforcement Learning

Authors: Ziang CuiMengran YuChenyu ShiYingxuan ShiTianjiao Li

Organizations: Bilibili Inc., Shanghai, China

Abstract

Large Language Models (LLMs) have achieved remarkable strides in multilingual translation but are hindered by a systemic cross-lingual verbosity bias, rendering them unsuitable for strict time-constrained tasks like subtitling and dubbing. Current prompt-engineering approaches struggle to resolve this conflict between semantic fidelity and rigid temporal feasibility. To bridge this gap, we first introduce Sand-Glass, a benchmark specifically designed to evaluate translation under syllable-level duration constraints. Furthermore, we propose Homura, a reinforcement learning framework that explicitly optimizes the trade-off between semantic preservation and temporal compliance. By employing a constrained reinforcement learning objective featuring a novel dynamic syllable-ratio reward, Homura effectively "tames" the output length. Experimental results demonstrate that Homura significantly outperforms strong baselines, achieving precise length control that respects linguistic density hierarchies without compromising semantic adequacy.

Explore similar work

CardsList
  1. Streaming Speech-to-Text Translation with a SpeechLLM

    May 14, 2026Titouan Parcollet, Shucong Zhang, Xianrui Zheng +1Speech TranslationSpeech-Llms