SignRAG: Unified Retrieval-Augmented Gloss-Free Sign Language Translation
Authors: Zhi Rao, Yucheng Zhou, Qianran Sun, Yiqing Huang, Longcan Yuan, Jiayi Hou, Chengwen Yao, Lin Cheng, +3 more
Organizations: Faculty of Innovation Engineering, Macau University of Science and Technology, Macau, China · SKL-IOTSC, CIS, University of Macau, Macau, China · MAIS, Institute of Automation, Chinese Academy of Sciences, Beijing, China · Yale University, New Haven, CT, USA · VIVO AI Lab, China
Contemporary decoder-only large language models (LLMs) have demonstrated strong capabilities across a wide range of domains. However, existing pretraining paradigms for gloss-free sign language translation (SLT) are largely designed around conventional encoder-decoder pretrained language models, which limits their direct applicability to decoder-only LLMs. To address this limitation, we propose SignRAG, a unified framework combining hierarchical pretraining, target-domain retrieval augmentation, and retrieval-aware reinforcement fine-tuning. Hierarchical pretraining first learns linguistically grounded sign representations and then jointly aligns the sign encoder with an LLM, mitigating cross-modal optimization imbalance. For downstream adaptation, SignRAG complements parameter-based fine-tuning with a target-domain retrieval gallery that provides instance-specific translation cues. To ensure that retrieved contexts are used appropriately, we further introduce Retrieval Utility-Guided Reinforcement Fine-Tuning (RUG-RFT), which combines translation-quality and retrieval-utility rewards to encourage beneficial retrieval use while suppressing harmful reliance. Experiments on multiple SLT benchmarks establish new state-of-the-art performance. In particular, to the best of our knowledge, SignRAG is the first gloss-free approach to outperform gloss-supervised methods across all reported metrics on CSL-Daily. Our code has been released at GitHub, together with models of different sizes to support future academic research.
Figures & tables
Figure 1: Motivation. Hierarchical pretraining and retrieval augmentation improve sign-to-language learning. (a) , (b) Comparison of different pretraining strategies. The solid and dashed lines denote the previous state-of-the-art performance under the gloss-free and gloss-based settings, respectively, corresponding to Geo-Sign ( Fish and Bowden, 2026 ) and CV-SLT ( Zhao et al., 2024 ) . (c) Comparison of the semantic distance between generated and ground-truth (GT) texts with and without RAG. Semantic distance is measured using the independent sentence encoder paraphrase-multilingual-MiniLM-L12-v2 . Bars report the mean cosine distance over all samples, with error bars denoting ± SEM. Samples are drawn from the CSL-Daily dev set.
Figure 2: Overview of SignRAG. (1) Hierarchical pretraining on large-scale sign corpora: Stage I learns language-aware sign representations with a contrastive objective Lcon and a skeleton-to-text objective Lslt ; Stage II couples the pretrained sign encoder with the LLM (Qwen3) via a multimodal projector and performs large-scale joint pretraining. (2) Target-domain supervised fine-tuning: a sign retrieval gallery provides semantically related reference translations as auxiliary cues, and the model is fine-tuned with sign features and retrieved texts, with random retrieval dropout. (3) Retrieval utility-guided reinforcement fine-tuning: a quality reward and a retrieval-utility reward are combined and optimized with GRPO.
Dataset
Gloss
Method
Venue
Ext.
R-L ↑
B@1 ↑
B@2 ↑
B@3 ↑
B@4 ↑
CSL-Daily
w/ gloss
SignBT ( Zhou et al., 2021 )
CVPR’21
✗
49.31
51.42
37.26
27.76
21.34
SLTUNET ( Zhang et al., 2023 )
ICLR’23
✗
54.08
54.98
41.44
31.84
25.01
TS-SLT ( Chen et al., 2022c )
NeurIPS’22
✗
55.72
55.44
42.59
32.87
25.79
CV-SLT ( Zhao et al., 2024 )
AAAI’24
✗
57.06
58.29
45.15
35.77
28.94
w/o gloss
GFSLT-VLP ( Zhou et al., 2023 )
ICCV’23
✗
36.44
39.37
24.93
16.26
11.00
Sign2GPT ( Wong et al., 2024 )
ICLR’24
✗
42.36
41.75
28.73
20.60
15.40
Table 1: Main results on four SLT benchmarks. “w/ gloss” and “w/o gloss” denote methods with and without gloss supervision, respectively. “Ext.” indicates whether external sign-language datasets are used for pre-training. For the “w/o gloss” setting, methods with and without external pre-training data are separated and compared independently. Bold scores indicate the best results within each dataset, gloss setting, and external-data setting. “Venue” denotes the publication venue and year.
Configuration
R-L ↑
B@1 ↑
B@2 ↑
B@3 ↑
B@4 ↑
(a) Framework components
SFT baseline
55.03
51.16
37.99
29.40
23.46
+ Pretraining
58.78
56.81
44.85
36.18
29.72
+ Retrieval
60.20
59.39
46.97
37.95
31.22
+ RUG-RFT
62.34
61.12
48.74
39.73
32.91
(b) Reward design
Table 2: Ablation studies on CSL-Daily. Components in (a) are added cumulatively. Reward variants in (b) start from the same retrieval-augmented SFT model.
Figure 3: Main analysis. (a) Relationship between retrieval quality and final translation performance under different numbers of retrieved contexts. (b) BLEU-4 scaling with LLM size. We evaluate Qwen3 backbones with different parameter scales (0.6B, 4B, and 8B) under the same training and evaluation protocol.
Source
Sentence
Reference
红包在传统意义上也叫压岁钱。 (Traditionally, a red envelope is also called “lucky money.”)
Retrieved
[1] 你打算怎么花这笔压岁钱呢? (How do you plan to spend this lucky money?) [2] 带着祝福的红包比攀比中的计较更有意义。 (A red envelope carrying good wishes is more meaningful than one given for comparison or competition.) [3] 你打算怎么花这笔压岁钱呢? (How do you plan to spend this lucky money?)
Geo-Sign
红包在历史上的原意是成年礼。 (Historically, the original meaning of a red envelope was a coming-of-age gift.)
SignRAG
红包在传统意义上是压岁钱。 (Traditionally, a red envelope means “lucky money.”)
Stage-II LoRA weights are merged into the base LLM, followed by a new set of LoRA modules. LoRA targets: q_proj , v_proj LoRA rank: 16 Scaling factor: 32 Dropout: 0.3
Training data: downstream SLT datasets Retrieved texts and sign features are jointly used as multimodal input. Prompt: Fig. 4
SFT LoRA weights are merged with the pretrained LLM weights, followed by a new set of LoRA modules. LoRA targets: q_proj , v_proj LoRA rank: 64 Scaling factor: 128 Dropout: 0.05
Optimization: GRPO Rewards: sentence-level quality reward and retrieval-utility reward The same prompt template as SFT is used.
Appendix
Table 4: Implementation details of the four training stages. LR denotes learning rate, FFN denotes feed-forward network, and RFT denotes reinforcement fine-tuning.
Figure 4: System prompt and instruction template. Here, with RAG denotes the setting where retrieved contexts are provided to the LLM, while without RAG denotes the setting where no retrieved contexts are provided.
Setting
BLEU-4
SignRAG-SFT w/o identical-text filtering
11.21
SignRAG-SFT w/ identical-text filtering
31.22
Appendix
Table 5: Effect of identical-text filtering during training-set self-retrieval on CSL-Daily.
Pretraining Strategy
R-L ↑
B@1 ↑
B@2 ↑
B@3 ↑
B@4 ↑
Pretraining evaluation on CSL-News
Sign Encoder + Shallow Transformer
36.31
41.58
30.24
23.51
19.02
Downstream SFT evaluation on CSL-Daily
Encoder-only Pretraining
22.51
18.22
8.76
5.02
3.20
End-to-End Pretraining
38.70
32.74
19.87
13.30
9.53
Hierarchical Pretraining (Ours)
58.78
56.81
44.85
36.18
29.72
Appendix
Table 6: Comparison of pretraining strategies. Results on CSL-News and CSL-Daily correspond to pretraining and downstream SFT evaluation, respectively.
Figure 5: Performance comparison between directly fine-tuning VideoLLaMA2 and the SignRAG-SFT model.
Stage
Dataset
Sign language
# Samples
Pretraining
CSL-News ( Li et al., 2025b )
Chinese
722,715
YouTube-ASL ( Uthus et al., 2023 ) (original)
American
530,161
YouTube-ASL (available)
American
434,953
Fine-tuning
How2Sign ( Duarte et al., 2021 )
American
35,172
OpenASL ( Shi et al., 2022 )
American
98,417
CSL-Daily ( Zhou et al., 2021 )
Chinese
20,654
Appendix
Table 7: Statistics of the pretraining and fine-tuning datasets. For YouTube-ASL, we report both the original dataset size and the number of samples available to us.
R-L ↑
B@4 ↑
chrF ↑
BSim ↑
SFT
60.20
31.22
26.93
82.10
RFT
62.34
32.91
28.42
82.96
Appendix
Table 8: Additional automatic metrics on CSL-Daily for checking metric overfitting. SFT and RFT denote SignRAG-SFT and SignRAG-RFT; R-L, B@4, and BSim denote ROUGE-L, BLEU-4, and BERTSim.
Method
B@1 ↑
B@2 ↑
B@3 ↑
B@4 ↑
R-L ↑
Uni-Sign
51.23
37.14
27.89
21.66
52.01
Geo-Sign
51.81
37.22
28.31
22.07
50.24
Pretrain → SFT (Ours)
53.18
39.14
30.27
24.44
54.55
Appendix
Table 9: Comparison on the CE-CSL benchmark. We further evaluate the adaptability of different pretrained SLT models on CE-CSL. Uni-Sign and Geo-Sign are included as representative baselines with publicly available pretrained weights. The results demonstrate the capability of our Pretrain → SFT to transfer and adapt to different downstream SLT benchmarks.
Source
Text
Ground-Truth
The weather is nice today.
Retrieved Text
The supermarket is closed today.
Pretrain → SFT
The weather is nice today.
SignRAG-SFT
The supermarket is holding an event .
SignRAG-RFT
The weather is nice today.
Appendix
Table 10: Qualitative example illustrating harmful reliance on retrieved contexts and the effect of RFT. Red text highlights misleading or incorrectly generated content.
Method
Latency
Throughput
Peak GPU Mem.
(ms/sample)
(samples/s)
(GiB/GPU)
Non-RAG ( Pretrain → SFT )
42.69
23.42
23.28
SignRAG-SFT
61.98
16.13
29.03
Appendix
Table 11: Online inference efficiency on CSL-Daily. Both methods use Qwen3-8B and four NVIDIA RTX 4090D GPUs under the same inference configuration. Retrieval itself is performed offline and is therefore excluded from the deployed generation path.
Gallery Size
Latency
Throughput
(ms/query)
(queries/s)
1K
3.49
286.67
5K
7.10
140.80
10K
14.65
69.32
18.4K
23.88
42.02
Appendix
Table 12: Offline retrieval efficiency on CSL-Daily with varying gallery sizes. All settings use the same 256 queries.
Gallery Size
Latency
Throughput
(ms/query)
(queries/s)
1,000
11.20
89.38
5,000
39.98
25.02
10,000
77.28
12.94
95,888
705.85
1.42
Appendix
Table 13: Offline retrieval efficiency at different gallery sizes on OpenASL. The same 256 queries are used for all gallery-size experiments.
Component
Cost
Pre-training Stage I
106.67 GPU-hours
Pre-training Stage II
277.50 GPU-hours
SignRAG-SFT
24.04 GPU-hours
SignRAG-RFT
5.78 GPU-hours
Total training compute
413.99 GPU-hours
Appendix
Table 14: Training cost of SignRAG.
Source
Sentence
Reference
我想要找借口。 (I want to find an excuse.)
Retrieved
[1] 他随口编了一个迟到的借口。 (He casually made up an excuse for being late.) [2] 为自己找借口的人,永远不会进步。 (People who make excuses for themselves will never improve.) [3] 我不会用出差作为任何事情的借口。 (I would not use a business trip as an excuse for anything.)
Geo-Sign
你找我有什么事吗? (Why are you looking for me?)
SignRAG
找借口。 (Find an excuse.)
Reference
封面设计是一种艺术设计。 (Cover design is a form of artistic design.)
Retrieved
[1] 一个好的封面设计会让你脱颖而出。 (A good cover design will make you stand out.) [2] 我姐姐是从事封面设计工作。 (My sister works in cover design.) [3] 我姐姐是从事封面设计工作。 (My sister works in cover design.)
Appendix
Table 15: Additional qualitative comparisons on CSL-Daily. Each example presents the reference translation, top-3 retrieved texts, and outputs of Geo-Sign and SignRAG. English translations are shown in parentheses below the original Chinese sentences.
Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks. However, fine-tuning LLMs for Gloss-Free Sign Language Translation (GFSLT) remains a challenge. In this paper, we investigate how to effectively adapt LLMs to the GFSLT task. We show that there are two key issues that need to be solved: (1) the inherent distributional gap between visual feature inputs and text feature inputs makes it difficult for LLMs to interpret visual inputs; and (2) existing approaches typically concatenate visual and textual features in an autoregressive framework, which leads to the model overemphasizing textual inputs and deprioritizing visual cues, as LLMs are pretrained predominantly on text-centric data. To address the first challenge, we propose a simple yet effective method named Filtered Pseudo-Gloss CTC Pretraining, which leverages filtered pseudo-gloss sequences generated from text sequences to supervise the training of the visual backbone. To tackle the second issue, we introduce a Visual-Prioritized Distillation training strategy. Specifically, we define a visual-only prediction path in which text inputs are masked, and the model is required to generate the target sequence relying solely on visual inputs. To guide this path, the outputs from the standard visual-textual prediction are then distilled into the visual-only prediction path, encouraging the model to prioritize visual features. Comprehensive experiments and qualitative analyses demonstrate the effectiveness of the proposed model. The proposed SignLlama achieves very competitive performance on multiple datasets for GFSLT tasks, without using any extra modalities or external sign language datasets for pretraining.
Many SLT systems quietly assume that brief chunks of signing map directly to spoken-language words. That assumption breaks down because signers often create meaning on the fly using context, space, and movement. We revisit SLT and argue that it is mainly a cross-modal reasoning task, not just a straightforward video-to-text conversion. We thus introduce a reasoning-driven SLT framework that uses an ordered sequence of latent thoughts as an explicit middle layer between the video and the generated text. These latent thoughts gradually extract and organize meaning over time. On top of this, we use a plan-then-ground decoding method: the model first decides what it wants to say, and then looks back at the video to find the evidence. This separation improves coherence and faithfulness. We also built and released a new large-scale gloss-free SLT dataset with stronger context dependencies and more realistic meanings. Experiments across several benchmarks show consistent gains over existing gloss-free methods. Our code and data are available at https://github.com/fletcherjiang/SignThought.
Yiyang Jiang, Li Zhang, Xiao-Yong Wei +1
The Hong Kong Polytechnic University · Sichuan University
Recent advances in large language models (LLMs) have led sign language translation (SLT), the task of converting sign-language videos into spoken-language text, to increasingly adopt LLMs as textual backbones. However, despite their strong language modeling capabilities, existing LLM-based SLT methods often undermine rather than exploit this language prior, producing disfluent translations, a failure we term language-prior degradation. Meanwhile, existing methods typically align videos and text at the sentence level, which does not ensure accurate lexical details and creates a lexical fidelity gap. To address both issues, we propose DualAnchor, a gloss-free LLM-based SLT training framework that couples two complementary anchors for linguistically fluent and visually faithful generation. Token-level Prior Anchoring (TPA) preserves the LLM's language prior by regularizing the multimodal decoder at each decoding step toward the next-token distribution of a frozen LLM conditioned on the same autoregressive prefix. Optimal Transport Alignment (OTA) improves lexical fidelity by formulating visual-textual matching as entropy-regularized partial optimal transport, with Sinkhorn optimization inducing a soft alignment between visual tokens and textual content tokens under a cosine cost. DualAnchor achieves strong overall performance on both PHOENIX-2014T and CSL-Daily. Targeted analyses attribute these gains to the complementary effects of the two anchors: TPA improves fluency, whereas OTA reduces fine-grained lexical errors.
Hongbin Zhang, Junhao Liu, Xuefeng Bai +3
Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen, China · Peng Cheng Laboratory, Shenzhen, China