cs.CLJul 22, 2026

emb-diversity: A Tool for Embedding-Based Measurement of Data Diversity

Authors: Cantao SuMenan VelayuthanEsther PloegerDong NguyenAnna Wegmann

Organizations: Department of Information & Computing Sciences, Utrecht University · Utrecht, The Netherlands

Abstract

There is growing evidence that data diversity is crucial for developing fair and robust NLP models. However, current approaches to measure diversity remain inconsistent and fragmented: While there exist a number of tools for measuring the lexical diversity of texts, researchers lack standardized tools for quantifying diversity based on embeddings. Embedding-based diversity measures are highly flexible: They work with any embedding model and any data that can be embedded, and are thus applicable to many notions of diversity. With emb-diversity, we provide a comprehensive embedding-based diversity measurement tool, spanning a broad range of measures. We demonstrate its potential for several use cases: measuring the stylistic, semantic, language and speaker diversity of datasets. https://github.com/nlpsoc/emb-diversity/

Explore similar work

CardsList
  1. STEB: Style Text Embedding Benchmark

    Jun 30, 2026Rafael Rivera Soto, Anna Wegmann, Cristina AggazzottiDomain-Invariant Text EmbeddingsMachine-Generated Text Detection