Deaf and hard-of-hearing (DHH) signers cannot converse in real time across different sign languages today: existing sign-to-sign translation systems run offline, requiring the full source clip before any target sign is emitted. Live use cases - e.g. broadcast interpretation and two-way video calls - instead demand simultaneous output, while the source signer is still signing. We present, to our knowledge, the first simultaneous sign-to-sign (S2S) translation system, with two wait-k regimes: test-time wait-k inference applied directly to a full-sentence model, and a trained wait-k model via stochastic multi-path supervision. We further introduce ca-Stream-AL, a computation-aware latency metric for streaming output. Averaged across six S2S directions on both a smaller human-verified test set and a larger synthetic S2S corpus, our streaming system achieves a 38% ca-Stream-AL reduction while staying within a 9% DTW-PA-MPJPE increase and a 2.1 BLEU-4 drop compared to the full-sentence baseline. A word-order case study probes how the streaming model handles word order mismatch between different sign languages - a consequence of simultaneous translation.
Figures & tables
Figure 1: Simultaneous sign→sign translation pipeline, shown for CSL → DGS (the same architecture serves all six sign→sign directions). The source video is incrementally fed to SMPL-X and tokenized by a VQ-VAE encoder into the source sign-token stream (purple, per-token duration l ); wait- k decoding (§ 3.1 ) emits target sign tokens (cyan, per-token duration l ) after a k -token lag, and a VQ-VAE decoder maps them back to SMPL-X for rendering as the target ( DGS ) avatar. We use l=160 ms in this work (4 frames at 25 fps). The vertical dashed line marks the end of source tokens at which target signing has already begun — but under offline sign→sign no target frame would be emitted before this point.
Figure 2: From classical AL to Stream-AL . (a) Dropping the τ cutoff captures the per-token theoretical lag of the target tail, but the metric remains target-centric and implicitly assumes the tail can be batch-emitted at the source-end instant t=∣x∣wchunk (visualized by the targets “dropped down” to a single column). (b) We fix the target-tail truncation issue by removing the source-end cutoff, which means extending the averaging horizon from the previous cutoff point τ to the full target sequence length ∣y∣ . (c) Stream-AL lays source and target on a shared wchunk=160 ms time axis and sums the signed gap against the source-side proportional ideal. Every unit is a real wchunk slot of wall-clock time, so the metric converts to seconds via ×wchunk .
Figure 3: In ca-Stream-AL , for source step i , the chunk arrives at i×wchunk , but the model completes generation at ti=max(ti−1,i×wchunk)+ci , where ci includes VQ-VAE encoding( ci(1) )/decoding( ci(3) ), transformer inference( ci(2) ), and SMPL-X rendering( ci(4) ). The wall-clock computation lag Δi=(ti−i×wchunk)/wchunk shifts the wait- k schedule from i−k to i−k−Δi , yielding ca-Stream-AL=Stream-AL+Δˉ .
Figure 4: Quality–latency trade-off on CSL → DGS , a representative direction out of the six sign→sign directions we evaluate. Test-time wait- k (blue) and trained wait- k (orange) traced over k∈{1,3,5,7,9,11,13} ; the full-sentence k=∞ baseline (green star) is connected by a dashed segment. Rows: test set (BT, Strict); columns: quality metric ( DTW-PA-MPJPE , BLEU-4 ); x-axis: ca-Stream-AL (seconds).
Direction
k⋆
AL (s) ↓
DTW ↓
BLEU ↑
ASL → CSL
7
1.97
3.74
7.5
CSL → ASL
9
3.71
4.79
10.4
ASL → DGS
7
1.15
3.85
7.4
DGS → ASL
9
2.41
3.55
9.0
CSL → DGS
7
1.34
2.88
7.5
DGS → CSL
9
3.05
3.27
7.4
Table 1: Best-tradeoff k⋆ per direction on BT under Train+TT, m=5 . AL = ca-Stream-AL (seconds); DTW = DTW-PA-MPJPE ; BLEU = BLEU-4 from the Sign Language Transformer evaluator. AL admits possible negative values for directions with r<1 .
Direction
k⋆
AL (s) ↓
DTW ↓
BLEU ↑
ASL → CSL
7
2.88
2.53
5.5
CSL → ASL
9
1.78
2.09
8.5
ASL → DGS
7
2.11
2.02
5.3
DGS → ASL
9
1.79
1.52
7.6
CSL → DGS
7
1.29
2.39
5.6
DGS → CSL
9
2.94
3.72
6.3
Table 2: Best-tradeoff k⋆ per direction on Strict under Train+TT, m=5 . Columns as in Tab. 1 .
Figure 5: CSL → DGS case study (Train+TT, k=3 ). (a) When the canonical source and target orders align over content words, the model emits the right DGS sequence and drops the CSL topic-marker glossed “color” ( yanse ) that has no DGS counterpart. (b) When the two languages disagree on the position of the negator, the streaming model emits source-order neg-nichts kalt as soon as the source negator ( bu , “not”) arrives, instead of the canonical DGS order kalt neg-nichts .
The field of sign language translation has witnessed significant progress in the translation between sign and spoken languages, but the translation between sign languages remains largely unexplored and out of reach. The latter can help 1.5 billion deaf and hard-of-hearing (DHH) people worldwide communicate across language barriers without relying on hearing interpreters or written-language fluency. The cascade approach composing separate sign-to-text, text-to-text, and text-to-sign systems suffers from error propagation and extra latency as well as the loss of information unique in the visual modality. We aim to develop direct sign-to-sign translation. However, a large-scale open-domain parallel corpus has not been curated between sign languages. To enable direct translation between sign language utterances, we use back-translation to produce synthetic sign-sign pairs from unaligned individual language utterance-sign corpora. Using this data, we jointly train a single MBART-based model for both text->sign (T2S) and sign->sign (S2S). On synthetically generated paired sets between American Sign Language (ASL), Chinese Sign Language (CSL), and German Sign Language (DGS), our direct S2S method outperforms the cascaded baseline on geometric sign error metrics (20% lower DTW-aligned MPJPE) and language matching metrics after predicted sign utterances are translated back to sentences (50% high BLEU-4) while achieving a roughly 2.3* speedup. On a small set of pre-existing cross-lingual sign data, we find similar improvements for our proposed method.
Most sign language understanding systems operate at the level of isolated signs, limiting their usefulness in natural communication. We study sentence-level sign language translation (SLT) with the primary goal of real-time deployment rather than proposing a new translation architecture. We fine-tune a SHuBERT-ByT5 translation stack on a uniformly sampled 9,872-example subset of How2Sign, selected because of compute and storage constraints, using QLoRA while keeping SHuBERT frozen. The model obtains a validation BLEU of 16.7 and, on the test split, BLEU 15.9 and BLEURT 44.7. The main contribution is a hardware-aware streaming system: a Raspberry Pi 4B reference client provides camera capture, local text display, and speech output, while compute-intensive perception and translation run on a CPU/GPU backend. The capture protocol remains client-agnostic, so the same backend can serve a browser, phone, or laptop. Chunked ingestion, bounded queues, parallelized perception, temporal reordering, and a sentence-boundary state machine reduce mean post-finalization response latency from 1.873 to 1.354 seconds (27.71%) and P95 latency from 2.919 to 2.130 seconds (27.03%) over the complete 9,872-example working subset.
Thanh-Hoang Nguyen Doan
The University of Danang - University of Science and Technology · The University of Danang – University of Science and Technology
Sign Language Translation (SLT) converts sign language videos into spoken-language text, bridging communication between Deaf and hearing communities. Current gloss-free approaches rely on large encoder-decoder models, limiting deployment. We propose a compact 77M-parameter pipeline that couples MMPose skeletal pose extraction with a single linear projection into T5-small. By varying the input frame rate, we expose a practical efficiency trade-off: at 12 fps the model halves its sequence length, achieving a 75% reduction in encoder quadratic self-attention computational complexity while incurring only a modest BLEU-4 drop (9.53 vs. 10.06 at 24 fps on How2Sign). Our system is roughly 3x smaller than prior T5-base systems, demonstrating that a lightweight architecture can remain competitive without hierarchical encoders or large-scale models.
Kuanwei Chen, Mengfeng Tsai
Computer Science and Information Engineering, National Central University, Zhongli, Taiwan