cs.CLSep 28, 2026

SignFLIP: A Unified Model for Sign Language Translation and Generation via Stage-wise Alignment at Scale

Authors: Zhaoyi An, Sihan Tan, Youngbae Hwang, Kazuhiro Nakadai, Rei Kawakami

Organizations: Institute of Science Tokyo · Chungbuk National University

Abstract

Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propose SignFLIP, a unified LLM-centered framework for translation and generation. To enable bidirectional mapping between text and sign, SignFLIP adopts a symmetric architecture together with a stage-wise training strategy built on large-scale data. The shared sign--text representation is progressively refined: pre-alignment facilitates subsequent SLT, while the SLT-adapted representation further benefits SLG. Extensive experiments on multiple benchmarks show that SignFLIP shows competitive performance compared with task-specific models on both translation and generation tasks, as well as strong transferability to sign language recognition.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Direct Translation between Sign Languages

    May 20, 2026Zetian Wu, Bowen Xie, Wuyang Meng +3Gloss-Free Sign Language TranslationSign Language Translation

  2. Bridging the Gap Between Semantics and Reconstruction:Unifying Sign Language Translation and Production

    Aug 10, 2026Xiao Liu, Shiwei Gan, Yafeng Yin +5Sign Language TranslationMotion-Language Alignment

  3. SIGNET: Motion-Level Knowledge Transfer for Cross-Language Sign Language Translation

    Jun 26, 2026Sobhan Asasi, Ozge Mercanoglu Sincan, Richard BowdenGloss-Free Sign Language TranslationSign Language Translation