cs.CVApr 5, 2022

A Transformer-Based Contrastive Learning Approach for Few-Shot Sign Language Recognition

Authors: Silvan FerreiraEsdras CostaMarcio DahiaJampierre Rocha

Organizations: CESAR · Universidade Federal do Rio Grande do Norte (UFRN), Natal, Brazil · Lenovo

Abstract

Sign language recognition from monocular video or 2D pose sequences is challenging, both because 3D information must be inferred from 2D observations and because the signal is inherently spatiotemporal. Moreover, the large and continually growing vocabulary of signs in production settings makes conventional closed-set classification impractical: adding a class requires new labeled data and retraining. We propose a contrastive Transformer-based model that learns rich representations of body key-point sequences, enabling direct comparison between embedding vectors. These representations support one-shot and few-shot tasks such as classification of signs never seen during training. On the LSA64 dataset, using only 48 classes for representation learning, the model reaches 88.4% accuracy on 16 held-out classes with as few as eight reference examples per class, and its accuracy improves consistently with the number of training classes and support examples.

Explore similar work

CardsList
  1. TransSLR: A Lightweight Transformer for Sign Language Recognition

    Aug 3, 2026Lucia Yen Wanchi, Samuel Johnny, Victor Tolulope Olufemi +2Sign Language Recognition3D Encoder

  2. Zero-Shot Cross-Lingual Recognition of Sign Language Handshapes

    Sep 16, 2026Marcel Granero-Moya, Carolina del Corral Farrarós, Gloria Haro +2Sign Language RecognitionSign Language