Research
Towards participatory speech dataset curation: A queer case study and conceptual framework
Overview Research area: Responsible and ethical AI, specifically participatory approaches to speech dataset curation, with a case study on the LGBTQIA+ ("queer") community. Technical level: Beginner-F

- arXiv
- 2609.25496
- Published
- 2026-09-21
- Authors
- Brooklyn Sheppard, Anaelia Ovalle, Adina Williams, Levent Sagun
AI summary
Overview
Research area: Responsible and ethical AI, specifically participatory approaches to speech dataset curation, with a case study on the LGBTQIA+ ("queer") community.
Technical level: Beginner-Friendly. The paper is conceptual and normative. It contains no experiments, no model training, no evaluation metrics, and no mathematical formalism.
Scope (one sentence): Drawing on a review of existing speech data collection practices, prior participatory AI work with queer communities, and participatory methods in field linguistics, the paper proposes a four-phase conceptual framework for curating speech datasets by, for, and with marginalized communities.
Publication details: arXiv:2609.25496v1 [cs.AI], 21 Sep 2026. Authors Brooklyn Sheppard, Anaelia Ovalle, Adina Williams, and Levent Sagun. Affiliations listed: University of Calgary (1) and Meta FAIR (2), with correspondence at the University of Calgary.
What This Paper Is About
Speech datasets that are both high-quality and diverse are difficult to collect, and common practices such as crowdsourcing tend to engage communities only briefly. This creates risks of "participation washing," where engagement is temporary rather than a sustained commitment, and it leaves marginalized speakers such as those outside the gender binary largely absent from speech data. The paper argues that participation should be considered as early as dataset creation itself, and develops a conceptual framework to guide speech dataset curation that is led by, intended for, and built with marginalized communities, using queer speakers as the motivating case study.
Key Contributions
-
Motivating a participatory framework through the lens of queer speakers. The authors assemble evidence on why the queer community is a distinctive and instructive case for inclusive speech technology, covering dataset representation gaps, reported AI harms, and the practical challenges of direct engagement.
-
Reviewing the landscape of prior participatory and community-engaged work. The paper surveys participatory AI efforts with queer communities (notably work in AI-assisted mental health), participatory speech data collection for other marginalized communities (such as the StammerTalk stuttered-speech initiative), and reciprocal methods in field linguistics.
-
Proposing a four-phase conceptual framework for community-led speech data curation. The framework comprises overlapping, two-way processes of (1) community, (2) project formulation, (3) modes of participation, and (4) personal autonomy.
-
Supplying guiding questions and a schematic comparison of approaches. For each of the four phases, the paper offers concrete guiding questions for community contributors, and Figure 1 contrasts a top-down approach (a one-directional flow from researcher to output) with a co-design approach (nested circles with bidirectional arrows indicating iteration, collaboration, and co-design).
Main Findings
-
Mainstream speech datasets underrepresent gender-diverse speakers. Although datasets such as Fairspeech and Casual Conversations aim to represent a variety of speaker backgrounds and include speaker gender labels, the majority do not include any speakers outside the gender binary, and those that do contain only a small number of speakers from this group.
-
Evidence of model bias against non-binary speakers exists but is explicitly preliminary. One prior study examined three speaker gender categories, "male", "female", and "other", in the Common Voice 16.0 dataset. All speech models tested performed worse for speakers of a particular gender, but models differed in whether they favored "male" or "female" identifying speakers. The "other" gender category was never favored by any model tested. The authors of that study note the sample ranged from only 8 to 86 speakers depending on the language, and describe the investigation as preliminary with more data needed for generalizable conclusions.
-
Binary and perceived-gender annotation practices risk misgendering and erasure. Gender annotation has historically been based on perceived gender and limited to binary labels, although self-reported gender identity is becoming more common in speech datasets.
-
Gender annotation carries unusual sensitivity for queer speakers. While gender annotations are not commonly treated as sensitive compared to demographic factors such as sexual orientation or disability status, volunteering a gender-diverse identity in effect "outs" a speaker as a member of a marginalized community, in a way distinct from cisgender binary-identifying individuals.
-
Queer speech characteristics can change over time and context. The paper notes that queer speakers' speech may shift as a result of voice coaching or hormone therapy, or when communicating with speakers inside versus outside the community.
-
The queer community has a documented history of AI and technology harms. The paper cites evidence including claims of an AI "gaydar," use of geolocation data to target trans healthcare facilities for harm, disproportionate flagging of queer dialects as toxic by content-moderation models, and qualitative interviews in which trans and non-binary users of voice AI assistants conveyed a lack of trust in AI developers, privacy concerns, discomfort at being visible to developers, and doubts about the ethics of AI practices.
-
Crowdsourcing is efficient but structurally limited for sustained engagement. Crowdsourcing has produced datasets for specific languages, specific dialects, and specific demographic groups within a language community, but the paper identifies a gap when these efforts lack meaningful, long-term collaboration.
-
Community preferences can conflict with researcher assumptions about data sharing. Prior work on sensitive speech data for neurological disorders proposes desiderata including informed consent, stringent security and privacy measures, and transparency about which communities are represented. By contrast, the StammerTalk initiative saw community members advocate for public release of their own co-designed stuttered-speech dataset, with a post-data collection survey indicating that the process contributed to feelings of empowerment for both researchers and speakers. StammerTalk invited anyone who is interested and self-identifies as a stutterer to participate, with no restrictions on age, gender, or other demographic factors.
-
Even strong participatory projects can retain a collector/contributor divide. In StammerTalk, the core research team formulated the project design independently from the data contributors, or speakers themselves.
-
Field linguistics is shifting from one-way extraction to two-way knowledge sharing. Historically, typically white, European descendant scientists engaged in a one-way relationship, extracting data from speakers to advance the discipline. More recent participatory efforts call for a cyclical approach in which speakers and linguists transfer knowledge in both directions, often aiming to revitalize a language through documentation and curriculum development.
-
The proposed framework is deliberately normative, not prescriptive. The authors present it as a form of normative infrastructure for speech dataset governance rather than a set of requirements, and describe the four phases as potentially cyclical.
-
Representation within a community is itself a design decision. The paper gives the example of targeting Queer in AI, a network primarily composed of individuals from countries of the Global North and predominantly English speaking, which could skew representation toward those with historical privilege; it suggests broadening outreach to queer community-based initiatives worldwide, and argues that projects should be explicit about who is not represented.
Methodology in Plain English
This is a position and framework paper rather than an empirical study. The authors do three things. First, they review common speech data collection and annotation practices, including crowdsourcing, and explain why those practices may be unsuitable for engaging with queer speakers. Second, they review prior participatory and community-engaged efforts relevant to their goal: participatory AI work that consults queer communities, participatory speech data collection for other marginalized groups, and reciprocal data-collection methods from field linguistics. Third, they synthesize insights from co-design and knowledge sharing into a conceptual framework, organized into four phases, each accompanied by guiding questions intended to walk a project through the speech data curation process. They also present a schematic figure contrasting top-down and co-design dataset creation. The paper reports no dataset collection, no model training, and no quantitative evaluation of its own.
Why This Matters
Impact on research. The paper argues that scaling up data collection to be more inclusive is not sufficient, and that participation should extend to the earliest stages of dataset creation. For speech technology, the practical consequence is that without more data from gender-diverse speakers, it is difficult to advance research that is appropriate and useful to queer users, and difficult to evaluate existing models on this group. The framework also reframes dataset governance questions, such as consent, storage, revocation, and allowable uses, as things to negotiate with communities rather than assume on their behalf.
Real-world applications:
- Speech synthesis and voice cloning, including diversifying voice options for end users and supporting those who may soon or have already lost their ability to speak.
- Voice assistant customization, where prior work cited in the paper finds that users offered voice customization options are more likely to have a positive experience and increased trust in their assistant, with that increased trust potentially mitigating prior exposure to misinformation.
- Speech recognition and speech categorization, including speaker verification and emotion classification, where evaluation and mitigation of demographic bias depend on having adequate data from the affected speakers.
- Content moderation, given reported disproportionate flagging of queer dialects as toxic.
- Sensitive data governance, with examples such as the Databrary project, which shares child behavioural video data and requires researchers to have completed ethics training and be affiliated with an institution governed by an ethics review board.
Industry relevance. The paper addresses how companies and research labs should structure engagement with marginalized speaker communities, including compensation, authorship, data storage, revocation, and how datasets are released. One author affiliation is Meta FAIR, and the framework speaks directly to organizations building speech datasets and voice systems.
Future Directions
-
Iteratively refining the framework for different use cases and speech communities. The authors state that future work would benefit from refining the framework across a variety of use cases, and that if it is used to develop a dataset of queer speech, the framework could be improved through reflections on the process and feedback from community members.
-
Applying the framework beyond the queer community. The authors believe the framework can be applied to participatory speech dataset development projects outside the queer community, so testing it in other settings is an open opportunity.
-
Resolving the tension between group consensus and individual autonomy. The paper leaves open how a community should decide when consensus is expected, such as on target speech genres or overall project goals, versus when individual members act on personal preferences regarding participation, acknowledgment, and data revocation, especially given the downstream effects one individual's participation may have on the community.
-
Handling fluctuating identities and representation gaps over time. Open questions include how community members can revoke their data, whether all members have access to the data indefinitely, how a project updates contributions when gender identities fluctuate, and how a project should be explicit about who is not represented.
Target Audience
Speech and language technology researchers who collect or curate datasets; responsible AI and ethics researchers interested in participatory methods; researchers in field linguistics and community-based data collection; community organizers and queer community members considering or being invited to contribute to speech data projects; and policy, governance, or product teams at organizations that build speech datasets and voice systems, including those weighing consent, compensation, storage, and public-release decisions. Readers wanting quantitative benchmarks or evaluations will not find them here, because the paper is conceptual and reports no such experiments of its own.
Authors’ abstract
In this paper, we motivate the need for a participatory speech dataset creation framework through a case study of the LGBTQIA+, or queer, community - a community with documented concerns about AI and reported harms, including attempts to develop 'gaydar' technologies that purportedly identify individuals as queer. We review common speech data collection practices, why these methods may be unsuitable for engaging with queer speakers, and discuss previous efforts in participatory AI with queer community engagement, as well as participatory endeavours specific to speech data collection for other marginalized communities. From this review, we develop a conceptual framework for participatory speech data curation by, for, and with marginalized communities drawing on insights from co-design and knowledge sharing. We propose a framework comprising overlapping and two-way processes of defining a community, project formulation, modes of participation, and personal autonomy.