Sign language-text alignment remains a fundamental challenge for text-driven sign language understanding. Existing methods predominantly rely on appearance-heavy RGB representations, which entangle motion semantics with visual variations and lead to ambiguous motion-language grounding. In this paper, we reformulate sign language-text alignment in a structured kinematic space and propose a kinematics-centric framework that adopts 3D SMPL-X motion as the primary representation. By explicitly modeling the kinematic dynamics of signing in a unified motion space, our approach reduces reliance on appearance signals and yields more semantically consistent representations. To capture the compositional nature of sign language, we introduce a gloss-guided local alignment mechanism that leverages gloss temporal spans as weak supervision to decompose continuous motion into coherent segments and establish fine-grained motion-text correspondences, thereby reducing ambiguity in localizing word-level semantics in continuous signing. Furthermore, we develop a visual distillation strategy, where RGB signals serve as privileged supervision during training to provide complementary contextual cues, while being completely removed at inference time. Extensive experiments on standard benchmarks demonstrate that our method achieves state-of-the-art bidirectional retrieval performance on CSL-Daily and competitive results on PHOENIX-2014T. These results highlight the effectiveness of kinematic representations and explicit local grounding for sign language-text alignment.
Figures & tables
Figure 1. Comparison of RGB-based sign language-text alignment and our gloss-guided kinematic alignment.
Figure 2. Overview of the proposed gloss-guided kinematic sign language–text alignment framework. During training, the model jointly exploits text, RGB, motion, and gloss supervision to learn cross-modal alignment, where gloss spans provide local motion-language grounding and RGB serves as privileged supervision. Retrieval uses only the text and motion branches.
Figure 3. Illustration of the gloss-guided local alignment loss. Motion-token spans are sliced by gloss boundaries and padded for batching. Similarities between gloss and motion-span embeddings form a matrix, where row-wise and column-wise InfoNCE losses enforce motion-to-gloss and gloss-to-motion alignment. Their sum yields the gloss-guided loss Lg .
Figure 4. Representative Cases. Given an input text query, we show the GT frames of the matched sample (left), the video retrieved by SEDS with its paired text (middle), and the top-1 SMPL-X retrieved by our method with its paired text (right).
DataSet
Methods
T2V
V2T
R@1
R@5
R@10
R@1
R@5
R@10
CSL-Daily
UPRet ( Wu et al., 2024 )
78.4
89.1
92.0
77.0
89.2
92.7
CiCo ( Cheng et al., 2023 )
56.6
69.9
74.7
51.6
64.8
70.1
SEDS ( Jiang et al., 2024 )
85.8
94.4
95.6
85.4
93.8
95.8
Spot-ALIGN ( Duarte et al., 2022 )
34.2
48.0
52.6
23.6
47.0
53.0
CMCM ( Yang et al., 2026 )
84.7
93.3
95.1
79.3
90.2
92.7
Table 1. Comparison with prior methods on CSL-Daily and PHOENIX-2014T (R@K, %, ↑ ). Best results are in bold.
Model
Same-Sentence Diff-Signer ↑
Diff-Sentence Same-Signer ↓
Diff-Sentence Diff-Signer ↓
RGB
0.7798
0.4137
0.3651
SMPL-X
0.8975
0.0448
0.0041
Table 2. Signer-invariance analysis on CSL-Daily. Higher is better for Same-Sentence Diff-Signer; lower is better otherwise.
Figure 5. Intra-modal cosine similarity matrices for our SMPL-X embeddings (left) and SEDS RGB embeddings (right) on CSL-Daily. Samples are ordered by sentence identity, forming 2×2 diagonal blocks. SMPL-X exhibits clearer diagonal blocks and lower off-diagonal similarities, indicating stronger signer invariance.
Setting
Text to Motion
Motion to Text
R@1
R@5
R@10
R@1
R@5
R@10
Ours
87.09
95.36
96.87
86.99
94.64
96.85
No MOVE
80.58
93.86
95.86
81.38
93.20
95.92
No I3D
83.71
95.24
96.74
85.63
94.81
96.77
No Gloss
68.55
86.47
90.23
67.68
85.29
90.48
Table 3. Component ablation on CSL-Daily (R@K, %, ↑ ).
Continuous sign language recognition (CSLR) aims to recognize gloss sequences from unsegmented sign videos under weak sequence-level supervision. However, existing methods rely on sentence-level gloss annotations, providing limited temporal and semantic guidance for fine-grained representation learning. Conventional video-text alignment also requires large batch sizes, making it inefficient for memory-intensive sign language video training. In this work, we propose SMART, an MLLM-guided temporal alignment framework for joint sign recognition and spotting. SMART uses MLLMgenerated motion descriptions as auxiliary semantic cues and performs stable videotext alignment under small-batch training. To improve temporal representation learning, we introduce a Multi-Scale Temporal Adapter that models temporal interactions during transformer encoding. For dense temporal localization, SMART incorporates CSFormer, a CSLR-guided spotting module that injects recognition-derived gloss evidence into a boundary-aware spotting network. This unified framework enables CSLR features to benefit spotting, while spotting supervision complements weak CTC-based recognition. Experiments on four sign language benchmarks, including PHOENIX14-T, CSL-Daily, Large-scale KSL, and Disaster and Safety KSL datasets, demonstrate the effectiveness of SMART across both recognition and spotting tasks.
Sign language translation (SLT) converts continuous sign videos into spoken language text. Gloss-free approaches leverage pre-trained visual encoders and language models but rely on implicit cross-modal alignment from translation supervision alone. We present VTaMo, a framework that introduces explicit multi-granularity alignment at three levels: (1) local alignment via entropy-regularized optimal transport with a learnable null token for fine-grained frame-to-token correspondences; (2) global alignment via a learnable orthogonal transformation that calibrates embedding space geometry through Earth Mover's Distance; and (3) position-aligned contrastive learning for discriminative token-level representations. Experiments on Phoenix-2014T, CSL-Daily, How2Sign, and OpenASL demonstrate consistent state-of-the-art performance, with ablations confirming the complementary contributions of each component. Code is available at https://github.com/junyi2005/vtamo.
Junyi Hu, Zhewen He, Haomian Huang +2
New York University Abu Dhabi, UAE · ChatSign Technology
Sign Language Production (SLP) faces a fundamental trade-off: direct text-to-pose models suffer from regression-to-the-mean effects, while dictionary-retrieval methods produce disjointed transitions. To resolve this, we propose a novel training paradigm that leverages sparse keyframes to capture the underlying kinematic distribution of human signing. By generating dense motion from discrete anchors, our approach mitigates regression-to-the-mean while ensuring fluid articulation. To achieve this at scale, we introduce FAST, an ultra-efficient sign segmentation model that automatically mines precise temporal boundaries. We then present SignSparK, a Conditional Flow Matching (CFM) framework that utilizes these temporal anchors to synthesize 3D signing sequences. This keyframe-driven formulation also unlocks Keyframe-to-Pose (KF2P) generation, making precise spatiotemporal editing of signing sequences possible. Furthermore, SignSparK scales across four distinct sign languages, constituting the largest multilingual SLP framework to date, and integrates 3D Gaussian Splatting for photorealistic rendering. Extensive evaluations demonstrate that SignSparK achieves state-of-the-art across diverse SLP tasks and multilingual benchmarks. Our code is available at https://github.com/JianHe0628/SignSparK.