Research
Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Research in AI Safety & Ethics
Overview Research area: AI Safety & Ethics (AISE) — specifically the sociology/epistemology of the AISE research community, examining why human-subjects research is underused relative to technical met
- arXiv
- 2608.05656
- Published
- 2026-08-06
- Authors
- Jessica Y. Bo, Paula Akemi Aoyagui, Shalaleh Rismani, Dipto Das, Syed Ishtiaque Ahmed, Ashton Anderson
AI summary
Overview
Research area: AI Safety & Ethics (AISE) — specifically the sociology/epistemology of the AISE research community, examining why human-subjects research is underused relative to technical methods. Labelled on arXiv as cs.CY with the paper's category given as AI Safety & Ethics.
Technical level: Intermediate. The paper is qualitative-and-survey-driven rather than technical, but it uses statistical tests, mixed-effects modelling, and field-specific vocabulary (construct validity, epistemic fit, RLHF, red-teaming) that assumes some familiarity with the AI research landscape.
One-sentence scope: A survey (n=93) and interview study (n=17) of AI Safety & Ethics researchers from Technical, Sociotechnical, Governance, and Normative backgrounds, examining how they value human-subjects research and what stops them from doing it.
What This Paper Is About
AI safety and ethics research is largely about preventing harm to people, yet the field's dominant evaluation methods are technical: benchmarks, LLM simulations of human reactions, and normative assumptions baked into alignment rules. The paper asks why empirical research with actual human subjects remains rare and marginal in AISE despite this apparent mismatch. To answer, the authors go directly to AISE researchers — rather than analysing published literature — to capture what experts see as the epistemic fit of human research and what barriers they face in performing it.
Key Contributions
-
A direct study of researchers rather than publications. Instead of a literature review, the authors surveyed (n=93) and interviewed (n=17) AISE experts to capture motivations, deterrents, and lived experience behind how research is pursued or why it is not performed or valued.
-
Documentation of an "aspiration-practice gap." Using gap scores (aspiration minus practice) across four methodological dimensions — quantitative vs. qualitative, lab vs. field, experimental vs. observational, and representative vs. specialized sampling — the authors show that across all disciplines and dimensions, human research is valued more than it is practised.
-
Identification of a disciplinary epistemic tension. Technical researchers value human methods significantly less and collaborate across disciplines significantly less than other groups, suggesting epistemic incompatibility with human-centred methods rather than mere disagreement about priorities.
-
A framework of barriers spanning resources, infrastructure, and sector boundaries, plus the concept of "human-washing" — the superficial inclusion of human subjects that creates an appearance of validity without genuine epistemic contribution.
Main Findings
-
Consensus exists that human evidence is needed, but there is a human evidence gap. Participants agreed AISE is shifting toward human-centred problems, and many — including technically oriented participants — critiqued benchmarks and computational proxies for construct validity failures. As one Technical participant put it, a 0-to-100 scale is "totally arbitrary" with "no construct validity."
-
Sociotechnical researchers practise and aspire to human research more than Technical and Governance. Practice: Sociotechnical vs. Technical (Z=4.19, p_adj<.001); Sociotechnical vs. Governance (Z=2.67, p_adj=.007). Aspiration: Sociotechnical vs. Technical (Z=2.69, p_adj=.007); Sociotechnical vs. Governance (Z=2.79, p_adj=.005). Normative was excluded from these analyses due to small sample size.
-
Gap scores were consistently positive across all disciplines and dimensions, meaning human research is valued beyond current practice. The only significant dimension-level difference was that researchers aspire to conduct field studies more than lab studies — a signal of interest in ecological validity over experimental control.
-
Technical researchers showed a significantly smaller preference for human over non-human methods. In the mixed-effects model (reference: Sociotechnical), the positive intercept (β=0.74, p<.001) confirms Sociotechnical preference for human methods; Technical showed a smaller preference (β=-0.32, p=.017); Governance showed a similar but non-significant trend (β=-0.24, p=.084). Preference for human methods also weakened for low-risk (β=-0.38, p<.001) and future scenarios (β=-0.33, p<.001).
-
Governance researchers rated risk scenarios as more imminent than Sociotechnical (Z=2.62, p=.04) and Technical (Z=3.25, p=.007) researchers — the only disciplinary divergence in scenario classification. Otherwise, disciplines agreed on severity; Skill Erosion and Decision Biases were rated more imminent, and Decision Biases and Human Extinction higher risk.
-
Technical researchers collaborate across disciplines the least. Technical collaborated significantly less than both Sociotechnical (Z=-3.57, p_adj=.001) and Governance (Z=-3.45, p_adj=.001) by Holm-corrected Dunn post-hoc tests following Kruskal-Wallis comparisons. Notably, 77 of 93 survey participants identified with more than one research area.
-
The barrier is often access, not disagreement. Participants reported that methods differ between AIS and AIE communities but goals align — "the methods are different, but the mental structure is the same" — and that finding entry points into human-centred AI expertise is difficult.
-
Qualitative methods are simultaneously undervalued and defended. Technically oriented researchers reported less familiarity with qualitative methods and some viewed qualitative evidence as potentially easier to manipulate. Others argued that small-n work has value: "it's only 10 people, but you can still learn a lot from 10 people."
-
Resource barriers were felt broadly, with no significant disciplinary differences. Lack of time ranked most impactful, followed by funding, then participant access. One researcher could not run a neuroscience-based human-AI interaction study after losing a major grant, with an estimated cost of 300k USD — contrasted against an estimate of "less than a thousand dollars per paper" for computational work.
-
Infrastructure and mentorship constrain junior researchers. Normative participants rated lack of collaborators as their highest-impact barrier, the highest of any discipline. Interviews captured advisors discouraging human-subjects work, institutions not recognizing empirical methods, and one participant describing a plan to "shoehorn in" their interests while making a PhD "still look like an AI research PhD."
-
Survey sample demographics. Participants were located primarily in North America (47.3%) and Europe (44.1%); 67.8% hold or are completing a PhD; 70.9% work in academia; 88.2% are substantially or fully engaged in research; years of AISE experience ranged from 1 to 10+ (M=4.44, SD=2.70). Interview participants averaged 4.7 years of experience, spanning academia, non-profits, and industry.
Methodology in Plain English
The authors used a mixed-methods design with two components. First, an expert survey (n=93) with four sections: professional background, current practice and aspirations regarding human research, scenario-based perceptions of human vs. non-human methods, and barriers to human research. Inclusion criteria were deliberately broad — only self-identification as working on AISE-relevant research was required, so independent researchers, governance practitioners who may not publish, and those excluded by narrower definitions of AI safety could participate. Two attention checks were used, and responses failing both were excluded.
The scenario section had participants rate four AISE risk scenarios varying in risk level (low vs. high) and imminence (future vs. current), then rate the usefulness of human methods (e.g., testing an intervention with users) versus non-human methods (e.g., developing an automated benchmark). Participants rated the field generally, not their own discipline's practice.
Second, semi-structured interviews (n=17) provided richer accounts covering research area, relationship to human research, cross-disciplinary collaboration, perceived AI harms, and barriers. Coding was deductive from the research questions, refined inductively using interview themes; the codebook was developed by the first author and refined through independent coding by two additional authors, with themes agreed collaboratively.
Recruitment ran primarily through the annual conference of the International Association for Safe & Ethical AI (IASEAI 2026), chosen partly because it convenes a substantial technical AI Safety community, unlike ethics-focused venues such as FAccT and AIES. The authors supplemented this with open recruitment via team social media and targeted Google Scholar searches, prioritizing underrepresented areas (for example, contacting recent authors from the journal Minds and Machines to target Normative researchers). Participation was not compensated. The interview sample included 6 Technical, 5 Sociotechnical, 4 Governance, and 2 Normative participants; participants who opted into follow-up interviews were likely already engaged with human research, which the authors acknowledge introduces selection bias.
Why This Matters
The paper argues that if human research is excluded from AISE's epistemic boundaries, the evidence base for making AI systems safe will be fundamentally incomplete — and that if human research is adopted as a performative checklist item instead, the result is "human-washing": superficial inclusion of human subjects that manufactures the appearance of validity. It draws an explicit parallel to explainable AI (XAI), where technical methods built on the assumption that transparency improves human-AI collaboration were later undercut by empirical studies showing explanations can increase over-reliance and worsen decisions. The authors warn AISE risks repeating this sequence at larger scale.
Real-world applications:
-
AI policy and governance: Participants argued policy should be supported by empirical evidence such as large-cohort A/B experiments, with one Governance participant stating "the studies we need for AI policy are non-existent yet" and another noting "we have such little evidence on the effects on people."
-
AI risk evaluation and benchmarking: The critique of construct validity in benchmarks and LLM simulations points toward evaluation practices that validate what "harm" actually means in practice rather than reducing it to a single optimizable number.
-
Organizational and institutional design: The mentorship and infrastructure findings apply to how labs, departments, and funders structure incentives, since advisors discouraging human-subjects work and institutions not recognizing empirical methods directly shape which research gets done.
-
Participant recruitment practices: Documented difficulties — fraud in specialized recruitment, and reliance on crowdworkers or undergraduates as a lower-validity substitute — speak to practical research-operations problems in human-subjects AI studies.
Industry relevance: Recruitment spanned academia, industry, non-profits, and government, and the paper reports that community dynamics at cross-sector venues still reflect the dominance of safety groups that "have money and presence." The finding that evaluation methods acceptable in industry versus academia may diverge — and the career-cost framing for researchers choosing human methods — is directly relevant to how AI labs staff and resource safety and ethics evaluation.
Future Directions
-
Establishing the epistemic fit of human research in AISE, which the authors frame as a recommendation rather than a solved problem — including how to legitimize human methods without performative inclusion.
-
Bridging prohibitive limitations such as time, funding, and participant access, potentially through resource models that make human studies competitive with the low marginal cost of computational work.
-
Making cross-disciplinary bridging bidirectional. The paper notes HCI and social science researchers should also actively engage with AIS communities rather than waiting to be contacted, and raises the open question of how to approach disciplinary differences in a structured way "rather than immediately going into this reflex of 'I hate this,' or 'I love this.'"
-
Whether and how top-down unification efforts can work without epistemic collapse or academic gatekeeping. Participants reportedly felt cross-disciplinary venues remained superficial and "didn't quite work this year," leaving unclear what institutional design would genuinely support integration.
Target Audience
This paper is most useful for AISE researchers and research leaders deciding how to allocate evaluation effort; human-computer interaction and social science researchers seeking entry points into AI safety work; program officers, funders, and institutional leaders designing incentives and interdisciplinary structures; and graduate students and junior researchers navigating the tension between human-centred interests and technically oriented advisors or publication norms. The paper's framing also speaks to policy and governance practitioners who need empirical evidence about AI's effects on people and currently find it lacking.
Note: The provided paper content is truncated mid-sentence during a participant quotation in the Results section, so any findings appearing after that point in the full paper are not reflected here.
Authors’ abstract
Safety risks of AI are becoming increasingly evident in human interactions with AI technologies. The prominent approaches to evaluating these risks favor technical methods, such as model benchmarks and LLM simulations, often sidelining empirical research with human subjects. To examine this apparent gap in the acceptance of human research, we conduct an expert survey (n=93) and expert interviews (n=17) with AI Safety & Ethics (AISE) researchers from Technical, Sociotechnical, Governance, and Normative backgrounds. Our findings suggest that although there is a consensus that human research is valuable for generating evidence for AISE, its adoption and acceptance are constrained by perceived validity issues, tangible resource barriers, epistemic and personal preferences in methods, and infrastructural constraints from the broader research community. In particular, Technical researchers tend to value human research less and collaborate across disciplines less, suggesting an epistemic tension towards human methods. We propose recommendations for establishing the epistemic fit of human research within AISE and bridging the prohibitive limitations that researchers face, while avoiding performative 'human-washing'.