Research
Shaping the Future of Generative AI for Black Communities: A Frame Analysis of Public Discourse and Empirical Scholarly Research
Overview Research area: Human-Computer Interaction / AI ethics, with a focus on generative AI (genAI), Black communities, and the framing of technology-related harm. Technical level: Intermediate. No
- arXiv
- 2608.24767
- Published
- 2026-08-25
- Authors
- Angela D. R. Smith, Gabriella Thompson, Christopher L. Dancy, Mark Díaz, Seyi Olojo, Christina N. Harrington
AI summary
Overview
- Research area: Human-Computer Interaction / AI ethics, with a focus on generative AI (genAI), Black communities, and the framing of technology-related harm.
- Technical level: Intermediate. No model building or programming is required, but the paper assumes familiarity with framing theory, algorithmic harm taxonomies, and the AI ethics literature.
- Scope (one sentence): The paper pairs a systematic literature review (SLR) of 91 empirical papers with a frame analysis of 28 public media resources to compare how scholarly research and public discourse define, explain, and propose fixes for genAI-related harm to Black communities.
What This Paper Is About
As genAI spreads through education, employment, healthcare, and creative industries, both journalists and researchers are debating what it means for Black communities — but they may be talking past each other. The authors ask whether the concerns raised in public discourse align with what empirical computing research actually studies and recommends. Using Entman's framing theory to organize both bodies of text, they find a shared focus on representational harm alongside a shared absence of Black epistemic agency, plus a sharp divergence in how each side explains where harm comes from.
Key Contributions
- Characterizing public discourse frames: The authors identify and describe the dominant frames shaping public narratives about genAI's anticipated impacts on Black communities, coding 28 media resources across Entman's four frame functions (problem definition, causal interpretation, moral judgment, treatment recommendation).
- Demonstrating a structurally produced misalignment: They show that public discourse attributes genAI harm to historical and systemic forces while the empirical literature stops its causal accounts at the dataset and its treatment recommendations at technical reform — and argue this gap is not incidental but a product of anti-Blackness operating in both registers.
- Advancing frame analysis as an AI ethics methodology: The paper argues that frame analysis can surface the structural and anticipatory dimensions of harm that individual-focused technical evaluation forecloses and make visible what research epistemology is structured to exclude.
- Mapping the empirical research record: The SLR characterizes 91 papers by publication year, venue, contribution type, genAI modality, domain, and the degree to which Black communities were actually engaged as participants or knowledge sources.
Main Findings
- Equal attention to representational harm: Representational harm was the only problem definition appearing with comparable weight in both corpora — 12 of 28 media resources and 62 of 91 research papers. The authors treat this convergence as a symptom of anti-Blackness structuring both registers rather than as confirmation of an adequate shared problem definition.
- Divergent causal accounts: 25 of 28 media sources identified systemic bias as a causal force and 21 of 28 identified historical bias, with over half citing both simultaneously. In the SLR corpus, 30 of 91 papers included no causal interpretation at all, and among those that did, 55 of 91 attributed bias to embedded bias in training datasets.
- Structural causal framing crossed ideological lines: Structural causes appeared not only in critical outlets but also in economic and business media (the paper names Forbes and McKinsey), suggesting the structural account is broadly legible rather than a specialist claim.
- Treatment recommendations follow causal accounts: Regulation was the most frequent media treatment (13 of 28), followed by reform of AI development (12 of 28). In the SLR corpus, reform was by far the most common treatment (78 of 91), with devising strategies to mitigate harm at 24 of 91, expand access at 5, resist at 3, and regulate at 5.
- Media surfaced harms the research corpus largely did not: Social system harm appeared in 21 of 28 media resources versus 36 of 91 SLR papers; interpersonal harm in 12 of 28 media resources versus 40 of 91 papers; allocative harm in 8 of 28 versus 24 of 91; quality of service harm in 11 of 28 versus 36 of 91. Media examples included findings that Black students were more than twice as likely as white or Latino students to be falsely accused of using AI to write their work.
- A shared evacuation of Black epistemic agency: Among 28 media resources, no frame centered on Black-led AI innovation, community-driven design, or Black technological agency as a structural response to representational harm. Among 91 research papers (14 of which engaged Black communities directly), only 3 engaged Black community members as sources of knowledge and validation rather than subjects of study, and two of those positioned community epistemology as constitutive of research design.
- Research concentrates on technical bias detection: 54 of 91 papers focused on addressing racial and gender bias in text-to-text and text-to-image models; 52 were model evaluations, 25 were benchmark or dataset development, and 12 were artifact or system development.
- Blackness as categorical variable: 28 papers featured Black and/or African American communities as the sole focus, 38 treated them as a subset of or comparison to other racial groups, and 4 grouped them under "People of Color." Many papers reduced Blackness to skin tone or dialect features. Only two papers framed Blackness in non-categorical terms — both noting variation across the diaspora within the United States and internationally.
- Minimal community engagement: Only four papers engaged Black communities in needs finding for genAI systems, and four papers used human annotation by Black community members as African American Vernacular English (AAVE) experts to verify dataset evaluation accuracy. Six papers pushed for greater participation of Black communities in AI development and research.
- Publication timeline: Although the search spanned 2010 to 2025, relevant papers appeared only between 2021 and 2025, with 2023 (n=9), 2024 (n=16), and 2025 (n=62).
- Moral judgment split the media corpus: Two resources — both from Black Enterprise, covering Amazon's AI initiatives for Black entrepreneurs and Robert Smith's investment in HBCUs — framed genAI as purely beneficial and were the only sources oriented primarily toward opportunity rather than harm. The rest split between "harmful only" (concentrated in cultural representation coverage) and "both harmful and beneficial" (concentrated in economic and educational coverage).
- Environmental harm coverage is diagnostic only: The two most recent media sources, on data center infrastructure and environmental harms to Black communities, offered no treatment recommendations at all — a pattern the authors read as an early diagnostic phase lacking a remedial vocabulary.
- Corpus construction: The database search returned 223 papers; 71 were removed at title and abstract screening, 61 duplicates were removed, and 18 were removed for focusing on marginalized groups broadly rather than Black communities, leaving 91 papers for coding (19 IEEE, 41 ArXiv, 13 ACM, 18 ACL). AAAI was initially included but dropped because repeated identical queries returned inconsistent results; Scopus was excluded for overlap with ACM and IEEE, and Google Scholar was excluded as unsuitable for systematic review.
- Note on completeness: The provided paper content is truncated mid-sentence in the treatment recommendation discussion of the SLR corpus, so that section's full conclusions are not reported here.
Methodology in Plain English
The authors ran two parallel studies and then read the results side by side.
First, a systematic literature review. They searched major computing and AI databases (ACM venues including CHI, CSCW, DIS, FAccT, and TOCHI; IEEE; ACL; EAAMO; NeurIPS; and arXiv) for full peer-reviewed empirical papers published between January 2010 and December 2025 that focused on genAI and Black populations. They combined "solo" racial identity terms such as "Black communities" with genAI terms such as "large language models" and "text-to-image." A coding instrument was iterated by all authors, tested on five papers, and then used to split and code the remaining corpus. Coding covered publication year, venue, search terms, study type, genAI type, intended contribution, domain, whether an artifact or dataset was involved, whether it was evaluated, and forms of engagement — plus whether authors defined Blackness or other identity terms. After noticing the corpus centered on harms, they recoded using Shelby et al.'s taxonomy of sociotechnical harms.
Second, a frame analysis of public discourse. They collected 28 articles from the popular press, op-eds, investigative journalism, blogs, and white papers from research organizations and think tanks, using keywords such as "Generative AI," "African American," "Black," "Black American," and "genAI." Two authors conducted an initial thematic analysis, then coded each source against Entman's four frame functions. Frame attributes were built following Matthes and Kohring's approach of treating each frame element as composed of analytical devices: problem definition used Shelby et al.'s five sociotechnical harms (social system, representational, allocative, quality of service, interpersonal); causal interpretation split into historical bias and systemic bias; moral judgment into "GenAI is harmful" and "GenAI can be both harmful and beneficial" (coded holistically, not double-coded); treatment recommendation into expand access, resist, regulate, devise strategies to mitigate harm, and reform. Findings from both corpora were then interpreted through Dancy and Saucier's onto-epistemological framework.
Why This Matters
Impact on research. The paper argues that the research field's causal accounts stop at the dataset and its remedies stop at technical reform, while public discourse reaches further back into historical and systemic forces. It also documents that only a small fraction of empirical work engages Black communities as knowledge producers rather than as measurement targets, which the authors frame as a methodological limit rather than merely a demographic gap. For AI ethics, the paper positions frame analysis as a complement to system-level evaluation.
Real-world applications:
- Research agenda setting: Funding bodies and labs can use the frame comparison to identify questions the technical literature is not asking, such as structural causes of harm and community-led design.
- Policy and regulation: Because regulation was the most common proposed treatment in media discourse (13 of 28) yet appeared in only 5 of 91 research papers, the analysis highlights where policy debate and empirical evidence diverge.
- Media and journalism practice: The corpus itself — including coverage of environmental harms and data center infrastructure with no proposed remedies — illustrates how framing choices shape whether audiences see a problem as actionable.
- Community organizing and advocacy: The documented absence of frames centered on Black technological agency identifies a gap advocates could target in public messaging and in research partnerships.
Industry relevance. The SLR shows that industry-adjacent and academic work overwhelmingly produces model evaluations, benchmarks, and datasets (52, 25, and 12 papers respectively) rather than participatory studies, so companies building genAI will find little in the current empirical record on how Black communities want to be engaged. The findings that darker-skinned subjects were progressively lightened across successive prompts, that models produced more homogeneous stories about darker skin tones, and that Black students faced elevated rates of false AI-use accusations point to concrete product and deployment risks.
Future Directions
- Test whether the misalignment is changing. The corpus is dominated by 2025 publications (n=62 of 91), so the field is moving quickly; a follow-up review could check whether causal accounts and treatment recommendations have shifted toward structural framing.
- Build out frame analysis as a method. The paper argues frame analysis can surface what technical evaluation forecloses but does not specify a repeatable protocol beyond the Entman and Matthes-and-Kohring scaffolding described here.
- Increase community epistemic engagement. Only 3 of 91 papers engaged Black community members as sources of knowledge and validation, and only 4 engaged communities in needs finding — an open question is what research designs would make community epistemology constitutive rather than consultative.
- Study the environmental harm discourse as it matures. The two most recent sources on data center infrastructure and environmental harms named the problem without proposing interventions, raising the question of what treatments emerge as that discourse develops.
- Broaden venue and discipline coverage. The authors flag the exclusion of AAAI as a coverage limitation and note they prioritized computing venues, leaving social science and humanities work on genAI reception outside the corpus.
Target Audience
AI ethics and fairness researchers, HCI scholars studying bias and community engagement, and systematic review or frame analysis methodologists. The paper is also relevant to policy analysts and advocates working on technology and racial equity, to journalists covering genAI's effects on Black communities, and to product and trust-and-safety teams at genAI companies who want to understand where current empirical research does and does not engage affected communities. Readers without a background in framing theory or algorithmic harm taxonomies will need to read the methods section carefully.
Authors’ abstract
As generative AI (genAI) systems become embedded in education, employment, healthcare, and creative industries, the impact and engagement among marginalized groups have become both a widespread discourse and a focus in scholarly research. As a starting point, we examine public discourse and empirical research to explore the impact of genAI systems on Black communities. We conducted a systematic literature review (SLR) of 91 empirical papers alongside a media discourse frame analysis of 28 public resources, applying Entman's framing theory to map how each corpus defines problems, attributes causes, and proposes treatments. Our SLR reveals that scholarly research concentrates heavily on technical bias detection, reducing Blackness to measurable variables rather than engaging with cultural practices, structural conditions, or Black knowledge systems. Our frame analysis reveals that public discourse attributes genAI-related harm to historical and systemic forces, while scholarly research stops its causal accounts at the dataset and its treatment recommendations at technical reform. We demonstrate that this misalignment is structurally produced: anti-Blackness operates simultaneously across both registers, generating a shared evacuation of Black epistemic agency. We argue for frame analysis as an AI ethics methodology capable of surfacing what technical evaluation forecloses.