Dual-Stream Simultaneous Translation via 2D Grid Attention
Authors: Yu Pu, Wei-Qiang Zhang
Organizations: Department of Electronic Engineering, Tsinghua University, Beijing 100084, China · Institute for Embodied Intelligence and Robotics, Tsinghua University, Beijing 100084, China
Simultaneous machine translation must generate target tokens before the source input is complete. Existing approaches address this through post-hoc read-write policies, leaving the attention mechanism unaware of bidirectional stream dependencies. We propose a dual-stream attention framework that represents source and target streams as a two-dimensional grid of hidden states and models their interaction through four structurally distinct attention types merged via joint QK Softmax normalization. Two approximations---broadcast and Hadamard---reduce the per-layer complexity from O(X^2Y+XY^2) to O(X^2+Y^2+XY) with provably decaying error. Training uses a self-guided loop: a per-cell loss heatmap drives dynamic-programming path recovery, which generates read/write decision supervision labels without external alignment. An incremental KV cache with anchored rotary position embeddings enables efficient streaming inference. On Chinese-to-English simultaneous translation, the proposed model outperforms the Wait-k baseline by +5.66 BLEURT and +10.36 COMET at comparable latency, and surpasses the non-streaming reference on COMET at a fraction of the response delay.
Figures & tables
Symbol
Meaning
Typical Value
B
Batch size
—
X
Input sequence length
—
Y
Output sequence length
—
D
Model hidden dimension
896
H
Number of attention heads
14
Hk
Number of Key/Value heads (GQA)
2
TABLE I: Principal Notation
Fig. 1: Illustration of the 2D grid hidden state representation. Each grid cell (x,y) maintains hidden states for both the input stream Ix,y(l) and the output stream Ox,y(l) . The horizontal axis corresponds to input positions and the vertical axis to output positions.
Type
Query
Key/Value
Causal Constraint
I → I
QI
KI,VI
Causal in x ( x′≤x )
O → O
QO
KO,VO
Causal in y ( y′≤y )
I ← O
QI
KO,VO
Prefix in y ( y′≤y )
O ← I
QO
KI,VI
Prefix in x ( x′≤x )
TABLE II: Four Types of Duplex Attention and Their Properties
Fig. 2: Grid illustration of input stream attention. Left: I → I self-attention, where each cell attends causally along x and results are broadcast along y . Right: I ← O cross-attention, where each input cell aggregates from the output prefix y′≤y .
Fig. 3: Grid illustration of output stream attention. Left: O → O self-attention, broadcast along x . Right: O ← I cross-attention, where each output cell aggregates from the input prefix x′≤x .
Attention Type
Exact
Approximated
I → I self-attention
O(X2Y)
O(X2)
O → O self-attention
O(XY2)
O(Y2)
I ← O cross-attention
O(XY2)
O(XY)
O ← I cross-attention
O(X2Y)
O(XY)
Total
O(X2Y+XY2)
O(X2+Y2+XY)
TABLE III: Computational Complexity Per Layer Before and After Approximation
Fig. 4: Illustration of the loss heatmap and the optimal DP path. Color intensity indicates the grid loss L(x,y) (darker = higher loss). The black staircase line is the optimal monotone path: horizontal segments are WAIT steps and vertical segments are EMIT steps.
Cache Entry
Purpose
Update Step
KIy=0
I → I broadcast keys
WAIT
VIy=0
I → I broadcast values
WAIT
KOx=0
O → O broadcast keys
EMIT
VOx=0
O → O broadcast values
EMIT
mii,Zii,Sii
I → I joint Softmax statistics
WAIT
moo,Zoo,Soo
O → O joint Softmax statistics
EMIT
TABLE IV: Per-layer KV cache entries and update rules
Initialize cache with i1:1 and optional seed output o1:Yseed . Set xvis=1 , ygen=Yseed .
While not terminated:
Compute new O-stream column for current ygen using cached I ← O and O → O statistics.
Sample next token oygen+1 .
If EMIT head indicates WAIT and xvis<Xtotal , read next input token and update I-stream row cache.
Otherwise continue EMIT.
Return generated output sequence.
TABLE V: Pseudocode for the duplex inference main loop
Fig. 5: Incremental KV cache update. Left: EMIT step appends a new O-stream column and updates cross-stream statistics. Right: WAIT step appends a new I-stream row. Gray regions are cached; colored regions are newly computed.
TABLE VI: Dual-stream model training hyperparameters
Source
站在那边的男人曾是棒球手。
Reference
The man standing there was a baseball player.
TABLE VII: Chinese-to-English translation case study
Fig. 6: Grid forward propagation visualization. Horizontal axis: input sequence (left to right). Vertical axis: output sequence (bottom to top). Cell text: highest-probability token. Cell color: EMIT probability intensity.
θ
站在
那边
的男人
曾
是
棒
球
手
。
<eos>
0.3
standing
there, the man
had
been
a
pitcher
.
0.4
the
man standing there had
been
a
pitcher
.
0.5
the
man standing there
was
a
baseball
player
.
0.6
the man standing there
was
a
baseball
player
.
0.7
the man standing there
was
a
baseball player
.
0.8
the
man standing there
was
a
baseball
player .
TABLE VIII: Simultaneous inference under different EMIT thresholds
System
Param
BLEURT ↑
COMET ↑
AL ↓
AP ↓
FRL ↓
Base model
—
54.03
63.85
9.39
1.00
18.06
Wait- k
k=3
43.20
52.06
4.98
0.79
3.00
k=5
45.20
54.93
5.99
0.84
4.96
Ours
θ=0.5
47.01
59.59
0.23
0.49
2.76
θ=0.6
48.86
62.42
1.14
0.53
3.31
θ=0.7
49.10
62.79
1.61
0.55
3.79
TABLE IX: Translation quality and latency on the zh → en test set. The baseline reads the complete input before translating ( AP=1.00 ). Bold indicates the best value in each column.
Fig. 7: BLEURT–AL quality–latency trade-off on the zh → en test set.