cs.CVOct 4, 2026

Kinematics-Centric Continuous Sign Language Retrieval with Gloss-Guided Boundary-Aware Alignment

Authors: Chang Liu, Ke Han, Davide Talon, Elisa Ricci, Nicu Sebe

Organizations: University of Trento Trento, Italy · Fondazione Bruno Kessler Trento, Italy

Abstract

Sign language-text alignment remains a fundamental challenge for text-driven sign language understanding. Existing methods predominantly rely on appearance-heavy RGB representations, which entangle motion semantics with visual variations and lead to ambiguous motion-language grounding. In this paper, we reformulate sign language-text alignment in a structured kinematic space and propose a kinematics-centric framework that adopts 3D SMPL-X motion as the primary representation. By explicitly modeling the kinematic dynamics of signing in a unified motion space, our approach reduces reliance on appearance signals and yields more semantically consistent representations. To capture the compositional nature of sign language, we introduce a gloss-guided local alignment mechanism that leverages gloss temporal spans as weak supervision to decompose continuous motion into coherent segments and establish fine-grained motion-text correspondences, thereby reducing ambiguity in localizing word-level semantics in continuous signing. Furthermore, we develop a visual distillation strategy, where RGB signals serve as privileged supervision during training to provide complementary contextual cues, while being completely removed at inference time. Extensive experiments on standard benchmarks demonstrate that our method achieves state-of-the-art bidirectional retrieval performance on CSL-Daily and competitive results on PHOENIX-2014T. These results highlight the effectiveness of kinematic representations and explicit local grounding for sign language-text alignment.

Figures & tables

Explore similar work

CardsList
  1. SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting

    Aug 26, 2026Eunjee Choi, JungHoon Sung, Seongwhan Cho +2Sign Language TranslationGuidance

  2. VTaMo: Video-Text Alignment Model for Sign Language Translation

    Jul 10, 2026Junyi Hu, Zhewen He, Haomian Huang +2Gloss-Free Sign Language TranslationSign Language Translation

  3. SignSparK: Efficient Multilingual Sign Language Production via Sparse Keyframe Learning

    Mar 11, 2026Jianhe Low, Alexandre Symeonidis-Herzig, Maksym Ivashechkin +2Sign Language TranslationKeyframe Selection