Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propose SignFLIP, a unified LLM-centered framework for translation and generation. To enable bidirectional mapping between text and sign, SignFLIP adopts a symmetric architecture together with a stage-wise training strategy built on large-scale data. The shared sign--text representation is progressively refined: pre-alignment facilitates subsequent SLT, while the SLT-adapted representation further benefits SLG. Extensive experiments on multiple benchmarks show that SignFLIP shows competitive performance compared with task-specific models on both translation and generation tasks, as well as strong transferability to sign language recognition.
Figures & tables
Figure 1: Framework of SignFLIP, a unified model for SLT and SLG. A shared encoder learns joint representations from sign features and text inputs, which are then used by task-specific decoders to generate spoken text and sign poses. SignFLIP provides a unified framework for SLT and SLG while achieving competitive performance on both tasks.
Figure 2: SignFLIP is trained in three stages. Stage 1 performs pretraining to align sign and text embeddings through reconstruction and contrastive learning. Stages 2 and 3 focus on downstream SLT and SLG tasks, respectively. In particular, Stage 1 unfreezes the pose encoder and pose decoder. Stage 2 trains the pose encoder together with mT5 for text generation. Stage 3 then optimizes the mT5 encoder, projector, and pose decoder for sign generation.
Method
Modality
Dev
Test
Pose
RGB
BLEU-4 ↑
ROUGE ↑
BLEU-4 ↑
ROUGE ↑
Gloss-based
SLRT † Camgoz et al. (2020)
✓
11.88
37.96
11.79
36.74
MMTLB Chen et al. (2022a)
✓
24.42
53.38
23.92
53.25
TS-SLT Chen et al. (2022b)
✓
✓
25.76
55.10
25.79
55.72
SLTUNET Zhang et al. (2023)
✓
23.99
53.58
25.01
54.08
Table 1: SLT results on the CSL-Daily dataset. † and ‡ denote results reproduced by Zhou et al. (2021) and Zhou et al. (2023) , respectively. Underlined results indicate the best performance among gloss-based SLT methods, Blue and Green denote the best results of previous methods and ours, respectively. SignFLIP achieves competitive performance among pose-only methods and remains competitive with RGB-based and multimodal approaches.
Method
Modality
Test
Pose
RGB
BLEU-4 ↑
ROUGE ↑
BLEURT ↑
How2Sign
GloFE-VN Lin et al. (2023)
✓
2.2
12.6
31.7
YouTube-ASL Uthus et al. (2023)
✓
12.4
–
46.6
MSLU Zhou et al. (2025)
✓
2.4
17.2
–
C 2 RL Chen et al. (2024a)
✓
9.4
27.0
–
Table 2: SLT results on How2Sign. SignFLIP continues to outperform existing pose-only methods on How2Sign dataset.
Method
DTW ↓
DTW–MJE ↓
BLEU-4 ↑
How2Sign
Progressive Transformer
Saunders et al. (2020)
0.397
8.98e–4
1.63
Teach Me Sign
An and Kawakami (2025)
0.345
7.20e–4
3.95
Unified
Table 3: SLG results on the How2Sign and CSL-Daily datasets. Metric scales differ across datasets due to variations in generation difficulty and keypoint ranges.
Figure 3: Qualitative results on both How2Sign (left) and CSL-Daily (right) datasets. While the arm movements of the poses generated by our proposed SignFLIP are closer to the Ground Truth, our generated hand motions are also more natural than those from previous methods, by leveraging the feature distillation loss and hand-specific loss.
Method
How2Sign
CSL-Daily
SignFLIP (Ours)
1.69
1.94
GT motions
3.48
3.44
Table 4: Human evaluation scores for generated poses on How2Sign and CSL-Daily datasets.
Method
Modality
WER ↓
Pose
RGB
Dev
Test
SignBT ( Zhou et al., 2021 )
✓
33.2
33.2
SEN ( Hu et al., 2023c )
✓
31.1
30.7
CorrNet ( Hu et al., 2023b )
✓
30.6
30.1
MSLU ( Zhou et al., 2025 )
✓
28.6
27.9
Uni-Sign ( Li et al., 2025 )
✓
28.2
27.4
Table 5: CSLR results on the CSL-Daily dataset with WER scores. SignFLIP demonstrates its strong transferability on CSLR task even with encoders frozen.
Method
Modality
WLASL100
Pose
RGB
P-I
P-C
BEST Zhao et al. (2023)
✓
77.91
77.83
SignBERT+ Hu et al. (2023a)
✓
79.84
80.72
MSLU Zhou et al. (2025)
✓
88.76
89.25
NLA-SLR Zuo et al. (2023)
✓
✓
91.47
92.17
SignFLIP (Frozen)
✓
88.87
89.16
Table 6: ISLR results on the WLASL100 dataset under the frozen setting. Without task-specific fine-tuning of SignFLIP encoders, the model achieves promising performance, demonstrating the transferability of its learned shared representations to downstream tasks.
Figure 4: We visualize sign and text embeddings before and after SigFLIP, showing that initially separated feature distributions become closely aligned after training. Black-circled points and connecting lines indicate randomly selected sign–text pairs.
Setting
CSL-News
YouTube-ASL
BLEU-4
BLEU-4
w/o Pretraining
18.24
4.17
SignFLIP
21.61
8.43
Table 7: Ablation study on stage design for SLT task.
Setting
DTW
DTW-MJE
w/o Pretraining & SLT
0.034
8.44e-5
w/o SLT
0.030
7.37e-5
SignFLIP
0.025
6.15e-5
Table 8: Ablation study on stage design for SLG task.
Setting
CSL-News
CSL-News
BLEU-4
ROUGE
Qwen3 Yang et al. (2025)
5.19
20.07
mT5 (dec only)
3.10
19.61
mT5 (enc-dec, w/o stage1)
18.24
39.13
Table 9: Ablation study of LLM choices for SLT on CSL-News dataset.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Config
Stage 1
Stage 2
Stage 3
batch size
256
128
64
optimizer
AdamW
learning rate
3×10−4
weight decay
1×10−4
optimizer momentum
β1,β2=0.9,0.999
learning rate schedule
cosine decay
Appendix
Table 10: Training recipe of each stage.
Name
Language
Samples
Vocab.
Hours
Source
YouTube-ASL Uthus et al. (2023)
ASL
∼ 610K
60K
984
Web
How2Sign Duarte et al. (2021)
ASL
33K
16K
79
Lab
OpenASL Shi et al. (2022)
ASL
97K
33K
288
Web
WLASL Li et al. (2020)
ASL
2,038
100
Web
CSL-News Li et al. (2025)
CSL
∼ 700K
5K
1,985
TV
CSL-Daily Zhou et al. (2021)
CSL
20K
2K
23
Lab
Appendix
Table 11: Selected ASL and CSL datasets.
Method
Gloss
Visual
BLEU-4
ROUGE-L
BLEURT
OpenASL
GloFE-VN Lin et al. (2023)
✓
7.06
21.75
36.35
Conv-GRU † Camgoz et al. (2018)
✓
4.58
16.10
25.65
I3D-transformer Shi et al. (2022)
✓
5.66
18.64
28.82
OpenASL Shi et al. (2022)
✓
8.59
21.02
31.09
C 2 RL Chen et al. (2024a)
✓
13.21
31.36
–
Appendix
Table 12: SignFLIP SLT performance on OpenASL after Stage-2 fine-tuning.
Setting
DTW
DTW-MJE
Lalign
16-frame
0.30
6.1e–4
0.17
24-frame
0.12
2.7e–4
0.10
32-frame
0.31
7.2e–4
0.15
Appendix
Table 13: Ablation study with MPM window size on CSL-News dataset.
Setting
DTW
DTW-MJE
w/o Lhand and Ldistill
0.044
1.19e–4
w/o Lhand
0.029
7.84e–5
w/o Ldistill
0.037
1.02e–4
Ours
0.025
6.15e–5
Appendix
Table 14: Ablation study with SLG losses on CSL-Daily dataset.
Figure 5: Qualitative results on both How2Sign (left) and CSL-Daily (right) datasets.
Systems
Translation Output
Examples from CSL-Daily
Reference
你 和 小 张 什 么 时 候 认 识 的 ?
(When did you and Zhang meet?)
SignFLIP
你 的 小 张 什 么 时 候 认 识 的 ?
(When did you meet your Zhang?)
Reference
我 不 去 爬 山 , 我 有 事 。
Appendix
Table 15: Case study of translation outputs on CSL-Daily, How2Sign, and OpenASL. Examples are taken from the test sets. Sentences in brackets are our approximate English translations.
State Key Laboratory of Novel Software Technology, Nanjing University, Suzhou, Jiangsu, China · State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing, Jiangsu, China