Dual-Stream Simultaneous Translation via 2D Grid Attention
Authors: Yu Pu, Wei-Qiang Zhang
Organizations: Department of Electronic Engineering, Tsinghua University, Beijing 100084, China · Institute for Embodied Intelligence and Robotics, Tsinghua University, Beijing 100084, China
Simultaneous machine translation must generate target tokens before the source input is complete. Existing approaches address this through post-hoc read-write policies, leaving the attention mechanism unaware of bidirectional stream dependencies. We propose a dual-stream attention framework that represents source and target streams as a two-dimensional grid of hidden states and models their interaction through four structurally distinct attention types merged via joint QK Softmax normalization. Two approximations---broadcast and Hadamard---reduce the per-layer complexity from O(X^2Y+XY^2) to O(X^2+Y^2+XY) with provably decaying error. Training uses a self-guided loop: a per-cell loss heatmap drives dynamic-programming path recovery, which generates read/write decision supervision labels without external alignment. An incremental KV cache with anchored rotary position embeddings enables efficient streaming inference. On Chinese-to-English simultaneous translation, the proposed model outperforms the Wait-k baseline by +5.66 BLEURT and +10.36 COMET at comparable latency, and surpasses the non-streaming reference on COMET at a fraction of the response delay.
Figures & tables
Symbol
Meaning
Typical Value
B
Batch size
—
X
Input sequence length
—
Y
Output sequence length
—
D
Model hidden dimension
896
H
Number of attention heads
14
Hk
Number of Key/Value heads (GQA)
2
TABLE I: Principal Notation
Fig. 1: Illustration of the 2D grid hidden state representation. Each grid cell (x,y) maintains hidden states for both the input stream Ix,y(l) and the output stream Ox,y(l) . The horizontal axis corresponds to input positions and the vertical axis to output positions.
Type
Query
Key/Value
Causal Constraint
I → I
QI
KI,VI
Causal in x ( x′≤x )
O → O
QO
KO,VO
Causal in y ( y′≤y )
I ← O
QI
KO,VO
Prefix in y ( y′≤y )
O ← I
QO
KI,VI
Prefix in x ( x′≤x )
TABLE II: Four Types of Duplex Attention and Their Properties
Fig. 2: Grid illustration of input stream attention. Left: I → I self-attention, where each cell attends causally along x and results are broadcast along y . Right: I ← O cross-attention, where each input cell aggregates from the output prefix y′≤y .
Fig. 3: Grid illustration of output stream attention. Left: O → O self-attention, broadcast along x . Right: O ← I cross-attention, where each output cell aggregates from the input prefix x′≤x .
Attention Type
Exact
Approximated
I → I self-attention
O(X2Y)
O(X2)
O → O self-attention
O(XY2)
O(Y2)
I ← O cross-attention
O(XY2)
O(XY)
O ← I cross-attention
O(X2Y)
O(XY)
Total
O(X2Y+XY2)
O(X2+Y2+XY)
TABLE III: Computational Complexity Per Layer Before and After Approximation
Fig. 4: Illustration of the loss heatmap and the optimal DP path. Color intensity indicates the grid loss L(x,y) (darker = higher loss). The black staircase line is the optimal monotone path: horizontal segments are WAIT steps and vertical segments are EMIT steps.
Cache Entry
Purpose
Update Step
KIy=0
I → I broadcast keys
WAIT
VIy=0
I → I broadcast values
WAIT
KOx=0
O → O broadcast keys
EMIT
VOx=0
O → O broadcast values
EMIT
mii,Zii,Sii
I → I joint Softmax statistics
WAIT
moo,Zoo,Soo
O → O joint Softmax statistics
EMIT
TABLE IV: Per-layer KV cache entries and update rules
Initialize cache with i1:1 and optional seed output o1:Yseed . Set xvis=1 , ygen=Yseed .
While not terminated:
Compute new O-stream column for current ygen using cached I ← O and O → O statistics.
Sample next token oygen+1 .
If EMIT head indicates WAIT and xvis<Xtotal , read next input token and update I-stream row cache.
Otherwise continue EMIT.
Return generated output sequence.
TABLE V: Pseudocode for the duplex inference main loop
Fig. 5: Incremental KV cache update. Left: EMIT step appends a new O-stream column and updates cross-stream statistics. Right: WAIT step appends a new I-stream row. Gray regions are cached; colored regions are newly computed.
TABLE VI: Dual-stream model training hyperparameters
Source
站在那边的男人曾是棒球手。
Reference
The man standing there was a baseball player.
TABLE VII: Chinese-to-English translation case study
Fig. 6: Grid forward propagation visualization. Horizontal axis: input sequence (left to right). Vertical axis: output sequence (bottom to top). Cell text: highest-probability token. Cell color: EMIT probability intensity.
θ
站在
那边
的男人
曾
是
棒
球
手
。
<eos>
0.3
standing
there, the man
had
been
a
pitcher
.
0.4
the
man standing there had
been
a
pitcher
.
0.5
the
man standing there
was
a
baseball
player
.
0.6
the man standing there
was
a
baseball
player
.
0.7
the man standing there
was
a
baseball player
.
0.8
the
man standing there
was
a
baseball
player .
TABLE VIII: Simultaneous inference under different EMIT thresholds
System
Param
BLEURT ↑
COMET ↑
AL ↓
AP ↓
FRL ↓
Base model
—
54.03
63.85
9.39
1.00
18.06
Wait- k
k=3
43.20
52.06
4.98
0.79
3.00
k=5
45.20
54.93
5.99
0.84
4.96
Ours
θ=0.5
47.01
59.59
0.23
0.49
2.76
θ=0.6
48.86
62.42
1.14
0.53
3.31
θ=0.7
49.10
62.79
1.61
0.55
3.79
TABLE IX: Translation quality and latency on the zh → en test set. The baseline reads the complete input before translating ( AP=1.00 ). Bold indicates the best value in each column.
Fig. 7: BLEURT–AL quality–latency trade-off on the zh → en test set.
Simultaneous speech-to-text translation (SimulST) generates translations while speech is still unfolding, requiring a streaming policy that decides when to read and when to write. State-of-the-art approaches rely on attention-based encoder-decoder models where cross-attention provides explicit alignment signals. In contrast, Speech Large Language Models (SpeechLLMs) are decoder-only architectures relying solely on self-attention. This raises a central question: whether decoder self-attention contains sufficiently stable alignment signals to guide the streaming policy. Moreover, existing approaches typically rely on training-based adaptations or heuristic wait-k policies and have not been validated in long-form settings. To fill these gaps, we propose Decoder-Only Attention (DOA), a training-free policy that enables long-form simultaneous translation with off-the-shelf SpeechLLMs by deriving a proxy alignment from self-attention. Experiments on Phi4-Multimodal and Qwen3-Omni show that DOA provides an effective alignment signal for supporting streaming decisions, enabling low-latency long-form SimulST with quality close to offline decoding without retraining.
Simultaneous machine translation (SiMT) has traditionally relied on offline machine translation models coupled with human-engineered heuristics or learned policies. We propose Hikari, a policy-free, end-to-end model for simultaneous speech-to-text translation and streaming transcription. We also introduce Decoder Time Dilation, a mechanism that counteracts the overrepresentation of WAIT tokens in training. We present a supervised fine-tuning strategy that trains the model to recover from delays, significantly improving the quality-latency trade-off. Despite its modest size, Hikari delivers competitive translation quality at consistently low latency, comparing favorably with published IWSLT 2026 submissions up to 38x larger and with proprietary API systems across en-ja, en-de, and en-ru. We release our model weights and code to facilitate further research.
Long-form streaming speech-to-speech translation (S2ST) requires incremental, unbounded translation under strict latency constraints. Existing methods typically suffer from sentence-bounded supervision or demand massive paired-S2ST supervision. We introduce a training recipe enabling a speech language model for sentence-level and long-form streaming S2ST using only ∼2k hours of paired cross-lingual S2ST data, layered atop auxiliary supervision. Anchored by auxiliary multitask training, our approach remains robust even when the paired-S2ST budget itself is reduced by 90%. Our core contribution, joint text-code trajectory supervision, schedules target text and acoustic semantic codes as a unified commitment path, eliminating the need for separate, unstable speech-side emission controllers. Furthermore, our two-stream Thinker--Talker factorization significantly outperforms unified-decoder baselines by decoupling linguistic reasoning from dense acoustic prediction to mitigate modality interference. Finally, our system achieves highly competitive quality-latency trade-offs on RealSI and ACL60/60-dev, matching state-of-the-art, closed-source S2ST systems such as LiveInterpret~2.0 on ASR-BLEU.