Research
Missing the Margins: A Systematic Literature Review on the Demographic Representativeness of LLMs
Overview Research area: AI safety and ethics, specifically demographic representativeness and bias in Large Language Models used for social applications (simulation, personalization, advice). Technica
- arXiv
- 2511.01864
- Published
- 2025-10-15
- Authors
- Indira Sen, Marlene Lutz, Elisa Rogers, David Garcia, Markus Strohmaier
AI summary
Overview
- Research area: AI safety and ethics, specifically demographic representativeness and bias in Large Language Models used for social applications (simulation, personalization, advice).
- Technical level: Beginner-Friendly. The paper is a systematic literature review; it does not introduce new models, training methods, or benchmarks, and its content is primarily conceptual, categorical, and statistical.
- One-sentence scope: A systematic review of 211 papers published before December 1st, 2024, examining which demographic categories LLM studies probe, how they operationalize them, and whether the field's conclusions about LLM representativeness hold up given how those studies are designed and reported.
What This Paper Is About
Researchers disagree sharply about whether LLMs accurately represent different demographic groups. Some studies claim LLMs can simulate the opinions and behaviors of subgroups of a population, while others find LLMs only reflect certain groups, flatten differences within groups, or caricature people. The authors ask why these findings conflict by systematically reviewing how 211 studies select their target populations, report demographic subcategories, and evaluate models across and within groups.
Key Contributions
- A systematic literature review of 211 papers at the intersection of LLMs and demographics, drawn from arXiv, the ACL Anthology, Semantic Scholar, ACM Digital Library, OpenAlex, and community resources, covering work available before December 1st, 2024.
- A three-part annotation codebook that captures the context of LLM use (advice, simulation, content analysis, writing, generic), persona type (impersonation vs. personalization), response format, the LLMs and steering methods used, demographic categories and subcategories, target population, and each paper's conclusion on representativeness.
- A quantitative diagnosis of why conclusions about representativeness diverge, showing that papers with positive conclusions are more likely to skip demographically disaggregated evaluation, omit demographic subcategories, and exclude marginalized groups.
- A set of recommendations for reporting and evaluation, including tailored benchmarks with explicit target populations, demographically disaggregated analyses combining open and closed-form evaluation, and the addition of population and demographic fields to reproducibility checklists and model/data documentation.
Main Findings
- Conclusion split: Of the 211 papers, 29% conclude "yes" on representativeness, 34% conclude "partial," and 32% conclude "no." Eleven papers (5%) provide no conclusion and were excluded from that part of the analysis.
- Underreporting in positive papers: Among papers claiming LLMs are representative, 30% do not evaluate representativeness across multiple demographic categories nor within subcategories of a demographic; they report an overall assessment instead. This compares to 19% of partially representative papers and 10% of papers concluding no representativeness.
- Missing subcategories: 35% of positively concluding papers make claims about gender representativeness without reporting the gender subcategories studied. For race, the equivalent figure is 47%.
- Marginalized groups excluded: Of papers that do report subcategories, fewer than half include marginalized groups. Overall, 22% of all papers include gender subcategories beyond the binary, and 35% include diverse racial subgroups.
- Unstated populations: 36% of papers do not specify a target population at all. Among those that do define it explicitly or implicitly, most study the U.S. Specifically, 26% explicitly focus on the U.S. and 16% do so implicitly through U.S.-specific racial or political subcategories.
- Representativeness by target population: Studies targeting the global population report the lowest rate of positive representativeness at 11%. That compares with 23% for "Other" populations, 22% for explicitly U.S. populations, 14% for implicitly U.S. populations, and 45% for studies with an undefined target population.
- Context distribution: Advice is the most studied context at 43%, followed by simulation at 23%, generic at 13%, and content analysis at 11%. Advice and generic studies show strongly diverging conclusions, while simulation, writing, and content analysis mostly report partial representativeness.
- Steering method matters: 24% of papers using prompting to induce personas conclude positively about representativeness, compared to 63% of papers using fine-tuning. Prompt-based techniques appear in 64% of papers and fine-tuning in 13%.
- Models used: 62% of studies include more than one model, and 80% include at least one OpenAI model. LLaMa appears in 39% of papers and Mistral or Mixtral in 21%. The authors note that 11% of papers say only "ChatGPT."
- Demographic categories: Gender and race are the most studied categories, though the distribution is less skewed than in prior bias research. Simulation studies have a comparatively more balanced distribution of demographics. Political leaning, disability, and sexuality have comparatively fewer studies claiming LLMs are representative.
- Subcategory reporting gap: On average across all demographic categories, 38% of papers do not explicitly report the demographic subcategories and descriptors used. Nationality and class studies underreport subcategories more than other categories.
- Time trend: Over the studied period, articles with a positive outlook grew relatively slower while papers claiming partial representativeness increased, suggesting a move toward more nuanced evaluations.
- Other factors behind disagreement: Qualitative analysis found that many positively concluding papers rely more on closed-form response formats or do not account for the variance of LLM responses.
Methodology in Plain English
The authors searched five scholarly sources (arXiv, the ACL Anthology, Semantic Scholar, ACM Digital Library, and OpenAlex) plus community paper lists for work containing "Large Language Models" or "LLM" together with "demographic*," published before December 1st, 2024. They started from 13 papers assessed in a prior scoping review by Agnew and colleagues.
This produced 1,076 longlisted papers, reduced to 615 after semi-automatic deduplication, and finally to 211 included papers. Papers were included only if they were empirical research (not reviews, perspectives, or theoretical articles), touched on demographics, and studied text-only generative LLMs. Vision, speech, and multimodal work was excluded.
Three annotators, who are also authors, independently coded all 211 papers across three rounds. Their codebook covered the LLM use context, persona type (impersonation or personalization), response format, models and steering approaches, demographic categories with their subcategories and descriptors, the target population, and each paper's conclusion on representativeness. Debiasing and alignment papers were coded as positively concluding if their technique was found effective, since the authors treat bias reduction as improving representativeness. After each round, three papers per annotator were re-annotated by the other two as a reliability check; disagreement ranged from 3% to 8% across rounds and was resolved through discussion. The annotated paper list and analysis code are publicly available at https://github.com/Indiiigo/LLM_rep_review.
Why This Matters
Impact on research. The review argues that the field's perception of LLM representativeness is inflated. Positive claims are systematically concentrated in studies that evaluate less rigorously, report fewer demographic details, and include marginalized groups less often. This means the apparent disagreement in the literature is not simply a matter of different models or populations but is closely tied to methodological and reporting choices. The authors show that design freedoms such as persona induction method, response format, and model selection affect not just reproducibility but the assessment of representativeness itself.
Real-world applications.
- Personalized healthcare and educational recommendations, where an LLM should be equally helpful across groups rather than lacking background knowledge about some of them.
- Simulation of survey respondents and behavioral experiments in the computational social sciences, where misrepresentation of subgroups produces misleading social data.
- Hiring and decision-support advice, an advice subcontext the review identifies as frequently studied.
- Content analysis and subjective annotation tasks such as sentiment analysis and hate speech labeling, where the paper notes prior limits of demographic dimensions in human annotation appear to extend to LLM annotators.
Industry relevance. Most studies in the sample rely on commercial LLMs, with 80% including at least one OpenAI model. Because the review ties positive conclusions partly to evaluation practices rather than to demonstrated model capability, it suggests that product claims about a model serving diverse populations may rest on thin evidence. The recommendation to add explicit population and demographic categories to reproducibility checklists and model/data documentation sheets is directly actionable for model developers and auditors.
Future Directions
- Build tailored benchmarks that explicitly define target populations and demographic subcategories, and that combine open- and closed-form evaluations with demographically disaggregated analysis.
- Conduct a deeper meta-analysis of reports on individual demographic factors, which the authors say their codebook was not granular enough to support because many papers conduct no disaggregated analysis.
- Extend the analysis of marginalized groups to other widely studied categories such as age and political leaning, for example the elderly and political fringe groups.
- Explore under-examined representativeness interventions, including model editing and RLHF, and move beyond repurposing algorithmic bias toward intentionally designed representative LLMs through approaches such as detailed personas or pluralistic alignment.
- Track model versions and parameter sizes in a meaningful way, which the authors identify as currently impractical given multi-model studies, imprecise reporting, and heterogeneous evaluation metrics.
Target Audience
This paper is most useful to researchers and practitioners who use LLMs to simulate people or to serve specific populations: computational social scientists, NLP and AI ethics researchers, and evaluation and policy teams at organizations deploying LLMs for personalization. It is written at a level accessible to newcomers, so students and interdisciplinary researchers entering the field can read it without specialized technical background. The paper also speaks to reviewers, editors, and benchmark designers, since its recommendations target reporting standards, reproducibility checklists, and documentation practices.
Authors’ abstract
Many applications of Large Language Models (LLMs) require them to either simulate people or offer personalized functionality, making the demographic representativeness of LLMs crucial for equitable utility. At the same time, we know little about the extent to which these models actually reflect the demographic attributes and behaviors of certain groups or populations, with conflicting findings in empirical research. To shed light on this debate, we review 211 papers on the demographic representativeness of LLMs. We find that while 29% of the studies report positive conclusions on the representativeness of LLMs, 30% of these do not evaluate LLMs across multiple demographic categories or within demographic subcategories. Another 35% and 47% of the papers concluding positively fail to specify these subcategories altogether for gender and race, respectively. Of the articles that do report subcategories, fewer than half include marginalized groups in their study. Finally, more than a third of the papers do not define the target population to whom their findings apply; of those that do define it either implicitly or explicitly, a large majority study only the U.S. Taken together, our findings suggest an inflated perception of LLM representativeness in the broader community. We recommend more precise evaluation methods and comprehensive documentation of demographic attributes to ensure the responsible use of LLMs for social applications. Our annotated list of papers and analysis code is publicly available.