Skip to content
AI.info

Research

Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI

Overview Research area: Natural Language Processing / speech AI, specifically automatic speech recognition (ASR) and ASR-mediated voice interfaces, analyzed through sociolinguistics, language policy r

arXiv
2608.06141
Published
2026-08-06
Authors
Jay L. Cunningham, Mark Atta Mensah, Richard Martinez, Joao Vieira da Silva Neto, Efi Dawodu

AI summary

Overview

Research area: Natural Language Processing / speech AI, specifically automatic speech recognition (ASR) and ASR-mediated voice interfaces, analyzed through sociolinguistics, language policy research, and decolonial computing.

Technical level: Beginner-Friendly to Intermediate. The paper is conceptual and framework-oriented rather than a report of a new model or experiment; its technical content is mostly vocabulary-level discussion of ASR pipeline components (training data, metrics, model priors) and evaluation protocols.

Scope in one sentence: The paper argues that ASR failures for low-resource, Indigenous, and non-standard language varieties are implicit linguistic policies that reproduce colonial language hierarchies, and it proposes a harm taxonomy, a situatedness model, an audit protocol, and a participatory framework to address them.

What This Paper Is About

ASR and voice interfaces now mediate access to public services, healthcare, education, and legal processes, but they routinely fail speakers of low-resource, Indigenous, and non-standard language varieties. The authors argue these failures are not isolated bugs but predictable outcomes of design rules, assumptions, and evaluation practices — which they call linguistic policies — that determine whose voices become machine-legible. The goal is to supply a theoretical account of how these policies operate, a taxonomy of resulting harms, and a participatory framework and minimum audit protocol that position affected communities as co-designers, evaluators, and governance partners.

Key Contributions

  1. A theoretical account of linguistic policies in ASR, synthesizing linguistic capital theory (Bourdieu), raciolinguistic ideology (Rosa and Flores; Lippi-Green; Lawrence), language policy research (Spolsky; Markl), and decolonial computing (Irani et al.; Dourish and Mainwaring; Mohamed et al.). This account includes a seven-layer situatedness model for linguistic diversity in speech AI.

  2. The Three Harms (3M) taxonomy — Misrecognition, Misalignment, and Mistrust — operationalized with definitions, minimum evidence to collect, and likely downstream harms, plus a five-step minimum 3M audit protocol covering evaluator roles, sampling expectations, metrics, annotation procedures, and adjudication.

  3. A participatory framework for culturally competent ASR comprising four interconnected pillars: Participatory Auditing, Community Co-Design, Equitable Deployment, and Feedback Integration, modeled as a continuous loop of repair and redress rather than a one-time alignment exercise.

  4. A survey of community-led speech and language infrastructure classified into three categories (speech-specific corpora and benchmarks, NLP/machine-translation initiatives that shape language infrastructure, and community research networks that build local capacity), used as evidence that alternatives to extractive development already exist.

The authors explicitly state they do not propose a new ASR architecture or benchmark result; the contribution is a human-centered evaluation framework for identifying culturally situated harms that WER-centered evaluation regimes under-specify.

Main Findings

  • ASR performance gaps track social power, not linguistic complexity. The paper cites a system achieving 5% WER for Standard American English versus 35% WER for African American Language, and states that Blasi et al. show language technology performance correlates with socioeconomic power rather than linguistic complexity. Koenecke et al. are cited as finding large WER gaps between African American and White speakers across five major ASR systems.

  • "Low-resource" is a political-economic condition, not an inherent property. The paper argues that languages labeled low-resource have been made resource-poor through centuries of institutional marginalization, and that training corpora derive from the documentary legacy of colonial language dominance (English imposed in Ghana, Nigeria, Kenya, South Africa; Spanish in the Caribbean; Portuguese on Indigenous and Afro-Brazilian populations; French across West African nations including Senegal and Ivory Coast, and in Haiti).

  • WER encodes a linguistic policy. WER treats word-level errors as equivalent regardless of communicative consequence, cannot capture meaning-changing errors, and presupposes a single correct transcription — a contested notion for varieties without standardized written forms or for speakers who code-switch. For tonal languages such as Yoruba, which uses high, mid, and low tones to encode lexical contrasts, the paper cites Chen et al.'s Tone Error Rate (TER) as a tone-aware extension revealing errors WER collapses. For click-consonant languages (Xhosa, Zulu, Yeyi, Hadza, Sandawe, and Khoisan-family languages), the paper recommends reporting click-specific deletion, substitution, insertion, and misclassification rates alongside WER, CER, phoneme error rate, and community-weighted harm scores.

  • Misalignment is invisible to surface-form metrics. A Twi utterance such as "wo maame awo wo" may function as praise in context but appear insulting under a literal rendering; Hindi "aap" versus "tum" marks social distance and respect that normalization can erase; code-switching within a single utterance is a communicative strategy that single-language interpretation can distort. Detection requires pragmatic adequacy ratings, culturally grounded evaluation sets, and implicature/intent checks.

  • Mistrust is structural, not just perceptual. The paper cites Mengesha et al. on speakers of stigmatized English dialects framing voice assistants as unintelligent (a coping strategy signaling eroded trust) and Harrington et al. on Black older adults' experiences with a health-information voice assistant being shaped by code-switching, privacy concerns, and doubts about cultural fit. Mistrust creates an adoption and repair loop: disengagement means fewer corrections and less representative data, reducing performance and deepening mistrust. It also has a surveillance edge case, where communities with histories of monitoring may reject accurate systems.

  • Community-led resources already exist and vary in their relation to speech. UGSpeechData reports about 5,000 hours of validated speech across five Ghanaian languages — Akan, Ewe, Dagbani, Dagaare, and Ikposo — with distinct roles for speakers, validators, and transcribers. Mensah et al. evaluated seven Akan ASR models across four speech domains (culturally relevant image descriptions, informal conversations, biblical scripture readings, and spontaneous financial dialogues), finding domain dependence and accuracy degradation under dataset mismatch. IndicSUPERB, connected with AI4Bharat, introduced Kathbath, a labeled speech dataset with 1,684 hours across 12 Indian languages, with benchmarks for ASR, speaker verification, speaker identification, language identification, query-by-example, and keyword spotting. Mozilla Data Collective (formerly Mozilla Common Voice) is cited for open contribution while showing that contribution alone does not resolve governance, consent, and downstream reuse questions. Masakhane and AmericasNLP are cited as redistributing research authority and building resources ASR may later depend on.

  • Cultural competence cannot be inferred from language-level coverage. The paper uses Afro-diaspora English varieties (African American English, Jamaican Patwa, Caribbean Creoles) as a within-English comparison, showing that even a well-resourced language produces the same 3M pattern.

  • Participation is not inherently emancipatory. Without safeguards, it can produce representational capture, community-washing, and extractive data collection without withdrawal rights or benefit-sharing governance, so the framework treats participation as governance.

Methodology in Plain English

This is a conceptual and synthesis paper, not an experimental one. The authors combine four intellectual traditions — linguistic capital theory, raciolinguistic ideology, language policy research, and decolonial computing — to explain how social judgments about language become operational technical decisions in ASR pipelines.

They then work through three specific places in the pipeline where policy is enacted: training data curation (which varieties are collected and how they are transcribed), evaluation metrics (chiefly WER, and its alternatives), and model priors (language models' probability assumptions about "probable" speech). From this analysis they derive two tools: a seven-layer situatedness model (Language; Nationality/Nation-of-use; Geography and Region; Ethno-linguistics; Sociolinguistic Ideologies and Linguistic Markets; Social Justice Correlations; Socio-technical Implications) as a diagnostic map, and the 3M taxonomy as an operational diagnostic of harm.

The authors also include a positionality statement, describing their standpoints across African American English and other Afro-diasporic English varieties of North America and the Caribbean, Brazilian Portuguese, Nigerian English alongside Yorùbá and Ìgbò, Mexican and Latin American Spanish varieties, and Ghanaian English alongside Twi and Akan. They frame this as an accountability mechanism rather than a resolution, and acknowledge writing from North American universities.

The audit protocol they propose has five documented steps: define context and risk; construct a culturally grounded test set; specify evaluator roles (fluent community evaluators per target variety, at least one local domain expert for high-stakes settings, and one technical evaluator computing ASR metrics); annotate 3M failure modes using a shared codebook; and adjudicate and define repair, with disagreements resolved by community evaluators with documented rationales rather than external majority vote.

Why This Matters

Impact on research: The paper reframes ASR accuracy gaps as policy outcomes rather than data-availability accidents, and argues that exclusion is not merely the absence of inclusion but "the presence of a normative order that constitutes some speech as unintelligible by design." It offers reusable instruments — the seven-layer model, the 3M taxonomy, the five-step audit protocol, and the four-pillar loop — for research that evaluation should be situated by domain, register, and interactional purpose rather than inferred from a single benchmark score.

Real-world applications:

  • Healthcare and legal voice interfaces, where the paper argues WER cannot capture whether a transcription error changes meaning or has differential consequences.
  • Public service access, where refusal or English-only fallback behavior for unsupported varieties functions as a policy of categorical exclusion.
  • Language documentation and normalization decisions, where audits must decide when to preserve a Quechua, Nahuatl, K'iche', or Mixtec term, when to translate, and when translation would erase cultural or legal meaning.
  • Deployment governance, where the framework treats participation as governance — with withdrawal rights, consent, and benefit-sharing — rather than a methodological add-on.

Industry relevance: The paper gives teams a diagnostic table of policy sites and observable signals: benchmark composition skewing toward prestige languages; crowdsourcing that excludes rural or low-connectivity speakers; transcription guidelines enforcing standardized orthography over local spellings or code-switching; primary evaluation using WER without domain- or community-specific weighting; reference transcripts assuming a single correct version; language models penalizing code-switching; normalization rules removing honorifics or politeness markers as noise; and fallback defaulting to English-only responses. It also identifies measurable indicators for each harm — WER, CER, TER, refusal rate, intent error, subgroup gaps, community-rated pragmatic adequacy, code-switching and implicature checks, abandonment and opt-out rates, correction counts, trust and agency scales, complaint logs, and short qualitative interviews.

Future Directions

  • Extending the audit protocol's sampling requirements. The paper notes that a pilot audit may begin with a modest diagnostic set but that authors must report the number of speakers, utterances, minutes of audio, devices, recording conditions, and demographic or regional strata, and that deployment claims require larger samples and confidence intervals. What those thresholds should be is left open.

  • Developing community-specific audit designs. The Yoruba example shows that the relevant unit of harm depends on the linguistic structure of the evaluated community, and the authors explicitly reject the assumption that all tonal languages require the same audit design. Parallel designs for click-consonant languages, creoles, and Indigenous-contact varieties remain to be worked out.

  • Fleshing out the participatory framework's later pillars. The provided content details Participatory Auditing (Pillar 1) and begins Community Co-Design (Pillar 2), which draws on Harrington et al.'s deconstructed community-based design methodology; Equitable Deployment and Feedback Integration are named in Figure 2 but their details are not included in the material available here.

  • Turning conceptual bridges into concrete artifacts. The paper suggests community-led language work can yield test utterances, error labels, normalization rules, refusal conditions, and adjudication records usable for auditing ASR-mediated Spanish interfaces, rather than merely describing harms. Whether such artifacts get built and adopted is an open question, as is the paper's stated aim of remaining accountable to Global South knowledge-making.

Target Audience

Speech and language technology researchers and engineers who build or evaluate ASR systems and want a structured vocabulary for situated harms beyond WER; product and policy teams deploying voice interfaces in multilingual or high-stakes service contexts; computational sociolinguists, language policy scholars, and decolonial computing researchers; and community organizations, language activists, and Indigenous or minoritized language communities seeking a framework for demanding governance authority over speech AI that affects them. Readers looking for a new model architecture, training recipe, or benchmark result will not find one here — the paper states that its contribution is a human-centered evaluation framework, not an ASR architecture or benchmark result.

Authors’ abstract

This paper focuses on automatic speech recognition (ASR) and ASR-mediated voice interfaces that shape access to public services, healthcare, and education. We argue that persistent failures for low-resource, Indigenous, and non-standard language varieties are not only technical errors, but also implicit linguistic policies that reproduce colonial language hierarchies. Drawing on linguistic capital, raciolinguistic ideology, language policy research, and decolonial computing, we show how data, metrics, and model priors determine whose voices become machine-legible. We introduce the Three Harms (3M) taxonomy---Misrecognition, Misalignment, and Mistrust---and a seven-layer situatedness model for linguistic diversity in ASR and ASR-mediated voice interfaces. We then propose a participatory framework and minimum audit protocol for culturally competent ASR, positioning affected communities as co-designers, evaluators, and governance partners.

Read the original paper