cs.AISep 28, 2026

Dual-Stream Simultaneous Translation via 2D Grid Attention

Authors: Yu Pu, Wei-Qiang Zhang

Organizations: Department of Electronic Engineering, Tsinghua University, Beijing 100084, China · Institute for Embodied Intelligence and Robotics, Tsinghua University, Beijing 100084, China

Abstract

Simultaneous machine translation must generate target tokens before the source input is complete. Existing approaches address this through post-hoc read-write policies, leaving the attention mechanism unaware of bidirectional stream dependencies. We propose a dual-stream attention framework that represents source and target streams as a two-dimensional grid of hidden states and models their interaction through four structurally distinct attention types merged via joint QK Softmax normalization. Two approximations---broadcast and Hadamard---reduce the per-layer complexity from O(X^2Y+XY^2) to O(X^2+Y^2+XY) with provably decaying error. Training uses a self-guided loop: a per-cell loss heatmap drives dynamic-programming path recovery, which generates read/write decision supervision labels without external alignment. An incremental KV cache with anchored rotary position embeddings enables efficient streaming inference. On Chinese-to-English simultaneous translation, the proposed model outperforms the Wait-k baseline by +5.66 BLEURT and +10.36 COMET at comparable latency, and surpasses the non-streaming reference on COMET at a fraction of the response delay.

Figures & tables

Explore similar work

CardsList
  1. Streaming Translation and Transcription Through Speech-to-Text Causal Alignment

    Date pendingRoman Koshkin, Jeon Haesung, Lianbo Liu +4Speech TranslationTranscription

  2. SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision

    Jul 22, 2026Rongshen He, Xinyu Liang, Dekun Chen +3Speech TranslationSpeech Language Models