cs.CLApr 28, 2026

Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation Standards

Authors: Yiming NiZhi-Qi ChengJiayu LiWei Cheng

Organizations: Tacoma School of Engineering & Technology, University of Washington

Abstract

Sign languages are expressive visual languages used by Deaf and Hard-of-Hearing (DHH) communities. Despite substantial progress in sign-language recognition, translation, and production, advances remain constrained by fragmented datasets, inconsistent annotations, and limited linguistic coverage. Existing benchmarks often fail to reflect real-world communication needs, and systematic analyses of these limitations remain limited. In this survey, we present a comprehensive index of sign-language datasets, covering 120 resources across 35 sign languages. We analyze key challenges such as modality imbalance, annotation granularity, and signer bias, and outline considerations for future dataset design. We also introduce a 24-field Sign-Language Datasheet and release a public GitHub repository (https://github.com/Ginqwerty/Open-Sign-Language) to support standardized documentation and reproducible evaluation. Overall, our work provides a unified and practical foundation for developing inclusive, robust, and scalable sign-language technologies in real-world applications.

Explore similar work

CardsList
  1. Isolated Sign Language Recognition for Icelandic Sign Language: Experiments in a Low-resource Setting

    Sep 22, 2026Finnur Ágúst Ingimundarson, Guðný Björk Þorvaldsdóttir, Mathias Müller +1Sign Language Recognition