cs.CLFeb 28, 2026

QQ: A Language Metadata Toolkit for Multilingual NLP

Authors: Wessel PoelmanYiyi ChenMiryam de Lhoneux

Organizations: LAGOM·NLP, Department of Computer Science, KU Leuven · AAU-NLP, Department of Computer Science, Aalborg University

Abstract

Multilingual NLP research increasingly involves hundreds or thousands of languages across different datasets. Managing, discovering, and reporting language metadata becomes a common hurdle at these scales. We present QQ, a metadata toolkit and browser explorer. QQ compiles language metadata sources into a graph of language varieties, scripts, regions, identifiers, names, and relations, and exposes it through a Python API, a command-line interface, and a browser-based explorer. Users can normalize identifiers, retrieve metadata, traverse relations, and discover which external resources contain a language. We demonstrate QQ on three workflows: an audit of the HuggingFace Hub, linking resources that use different identifier systems, and generating reproducible language-reporting tables. QQ supports FAIR-oriented metadata practices through versioning, open formats, and reusable interfaces.

Explore similar work

CardsList