cs.AISep 21, 2026

Towards participatory speech dataset curation: A queer case study and conceptual framework

Authors: Brooklyn SheppardAnaelia OvalleAdina WilliamsLevent Sagun

Abstract

In this paper, we motivate the need for a participatory speech dataset creation framework through a case study of the LGBTQIA+, or queer, community - a community with documented concerns about AI and reported harms, including attempts to develop 'gaydar' technologies that purportedly identify individuals as queer. We review common speech data collection practices, why these methods may be unsuitable for engaging with queer speakers, and discuss previous efforts in participatory AI with queer community engagement, as well as participatory endeavours specific to speech data collection for other marginalized communities. From this review, we develop a conceptual framework for participatory speech data curation by, for, and with marginalized communities drawing on insights from co-design and knowledge sharing. We propose a framework comprising overlapping and two-way processes of defining a community, project formulation, modes of participation, and personal autonomy.

Explore similar work

CardsList
  1. Queer inclusion in speech datasets: An audit and taxonomy of practical tensions

    Sep 21, 2026Brooklyn Sheppard, Anaelia Ovalle, Adina Williams +1Hate Speech