Deaf and hard-of-hearing (DHH) signers cannot converse in real time across different sign languages today: existing sign-to-sign translation systems run offline, requiring the full source clip before any target sign is emitted. Live use cases - e.g. broadcast interpretation and two-way video calls - instead demand simultaneous output, while the source signer is still signing. We present, to our knowledge, the first simultaneous sign-to-sign (S2S) translation system, with two wait-k regimes: test-time wait-k inference applied directly to a full-sentence model, and a trained wait-k model via stochastic multi-path supervision. We further introduce ca-Stream-AL, a computation-aware latency metric for streaming output. Averaged across six S2S directions on both a smaller human-verified test set and a larger synthetic S2S corpus, our streaming system achieves a 38% ca-Stream-AL reduction while staying within a 9% DTW-PA-MPJPE increase and a 2.1 BLEU-4 drop compared to the full-sentence baseline. A word-order case study probes how the streaming model handles word order mismatch between different sign languages - a consequence of simultaneous translation.
Figures & tables
Figure 1: Simultaneous sign→sign translation pipeline, shown for CSL → DGS (the same architecture serves all six sign→sign directions). The source video is incrementally fed to SMPL-X and tokenized by a VQ-VAE encoder into the source sign-token stream (purple, per-token duration l ); wait- k decoding (§ 3.1 ) emits target sign tokens (cyan, per-token duration l ) after a k -token lag, and a VQ-VAE decoder maps them back to SMPL-X for rendering as the target ( DGS ) avatar. We use l=160 ms in this work (4 frames at 25 fps). The vertical dashed line marks the end of source tokens at which target signing has already begun — but under offline sign→sign no target frame would be emitted before this point.
Figure 2: From classical AL to Stream-AL . (a) Dropping the τ cutoff captures the per-token theoretical lag of the target tail, but the metric remains target-centric and implicitly assumes the tail can be batch-emitted at the source-end instant t=∣x∣wchunk (visualized by the targets “dropped down” to a single column). (b) We fix the target-tail truncation issue by removing the source-end cutoff, which means extending the averaging horizon from the previous cutoff point τ to the full target sequence length ∣y∣ . (c) Stream-AL lays source and target on a shared wchunk=160 ms time axis and sums the signed gap against the source-side proportional ideal. Every unit is a real wchunk slot of wall-clock time, so the metric converts to seconds via ×wchunk .
Figure 3: In ca-Stream-AL , for source step i , the chunk arrives at i×wchunk , but the model completes generation at ti=max(ti−1,i×wchunk)+ci , where ci includes VQ-VAE encoding( ci(1) )/decoding( ci(3) ), transformer inference( ci(2) ), and SMPL-X rendering( ci(4) ). The wall-clock computation lag Δi=(ti−i×wchunk)/wchunk shifts the wait- k schedule from i−k to i−k−Δi , yielding ca-Stream-AL=Stream-AL+Δˉ .
Figure 4: Quality–latency trade-off on CSL → DGS , a representative direction out of the six sign→sign directions we evaluate. Test-time wait- k (blue) and trained wait- k (orange) traced over k∈{1,3,5,7,9,11,13} ; the full-sentence k=∞ baseline (green star) is connected by a dashed segment. Rows: test set (BT, Strict); columns: quality metric ( DTW-PA-MPJPE , BLEU-4 ); x-axis: ca-Stream-AL (seconds).
Direction
k⋆
AL (s) ↓
DTW ↓
BLEU ↑
ASL → CSL
7
1.97
3.74
7.5
CSL → ASL
9
3.71
4.79
10.4
ASL → DGS
7
1.15
3.85
7.4
DGS → ASL
9
2.41
3.55
9.0
CSL → DGS
7
1.34
2.88
7.5
DGS → CSL
9
3.05
3.27
7.4
Table 1: Best-tradeoff k⋆ per direction on BT under Train+TT, m=5 . AL = ca-Stream-AL (seconds); DTW = DTW-PA-MPJPE ; BLEU = BLEU-4 from the Sign Language Transformer evaluator. AL admits possible negative values for directions with r<1 .
Direction
k⋆
AL (s) ↓
DTW ↓
BLEU ↑
ASL → CSL
7
2.88
2.53
5.5
CSL → ASL
9
1.78
2.09
8.5
ASL → DGS
7
2.11
2.02
5.3
DGS → ASL
9
1.79
1.52
7.6
CSL → DGS
7
1.29
2.39
5.6
DGS → CSL
9
2.94
3.72
6.3
Table 2: Best-tradeoff k⋆ per direction on Strict under Train+TT, m=5 . Columns as in Tab. 1 .
Figure 5: CSL → DGS case study (Train+TT, k=3 ). (a) When the canonical source and target orders align over content words, the model emits the right DGS sequence and drops the CSL topic-marker glossed “color” ( yanse ) that has no DGS counterpart. (b) When the two languages disagree on the position of the negator, the streaming model emits source-order neg-nichts kalt as soon as the source negator ( bu , “not”) arrives, instead of the canonical DGS order kalt neg-nichts .