Research
Appearing Legitimate is Not Enough: Interrogating Synthetic Agents in Representational Processes through a Participatory Design Lens
Overview Research area: Human-Computer Interaction, intersecting with Science and Technology Studies, Participatory Design (PD), technology ethics and law. The paper is a comparative case study and co
- arXiv
- 2608.17099
- Published
- 2026-08-17
- Authors
- Aditya Nayak, Aditi Vashistha, Alissa Centivany, Aakash Gautam
AI summary
Overview
Research area: Human-Computer Interaction, intersecting with Science and Technology Studies, Participatory Design (PD), technology ethics and law. The paper is a comparative case study and conceptual critique rather than a system-building or benchmarking paper.
Technical level: Intermediate. No models are trained or measured by the authors; the technical content concerns architectures described in secondary documentation (Retrieval Augmented Generation, hybrid neuro-symbolic systems) and the paper is readable by those without machine learning background, though it assumes some familiarity with HCI and design research vocabulary.
Scope: A close-reading analysis of three publicly documented synthetic agents deployed as substitutes for human participants in local policy consultation, enterprise jury deliberation, and global humanitarian diplomacy, arguing that they manufacture an appearance of legitimate participation.
What This Paper Is About
LLM-based synthetic agents are increasingly pitched as replacements for human participants, and in some cases are being extended into policy consultation, jury deliberation, and humanitarian diplomacy, where participation itself is what makes a process legitimate. The authors ask how these agents come to appear legitimate as substitutes for people, and what boundaries should govern their design and deployment. Their answer is that legitimacy is manufactured through four design steps that bypass the very representational processes from which institutional legitimacy is derived.
Key Contributions
-
A four-step account of manufactured personhood. Across three cases at different representational scales, the authors trace how personhood is produced by framing the problem, institutionally curating data, designing the user encounter with the persona, and evaluating validity through technical measures.
-
A process-oriented critique using Participatory Design. The authors apply Sanders et al.'s four modes of engagement — probing, priming, understanding, and generating — to show what representational work synthetic agents perform and what they leave out, complementing the outcome-centric evaluation frameworks that dominate existing literature.
-
Proposed soft and hard boundaries for design oversight. Soft boundaries separate analytic and indexing uses of LLMs from anthropomorphized interfaces; hard boundaries categorically separate anthropomorphized interfaces from systems that substitute for human participants in representational processes. The authors argue these proposals are timely because synthetic agents in representational contexts remain experimental and the boundaries are still tractable.
Main Findings
-
Appearance of legitimacy is manufactured, not earned. The authors argue that the four steps they trace collectively produce manufactured personhood: an appearance of legitimate participation that bypasses the representational processes from which institutional legitimacy is derived.
-
Framing reduces structural problems to feasibility gaps. In all three cases, difficulty in recruiting participants is framed as the central bottleneck. For Ask Amina and Ask Abdalla, the documentation cites time and resource intensity of surveys, focus groups, and questionnaires, geographic remoteness, unwillingness to participate, and the possibility that respondents have "an incentive to provide false or incomplete information." The authors counter that limited access to participants is a symptom of deeper structural failures of representation, not a single-solution problem.
-
Probing is replaced by a manufactured prober. Ask Amina and Ask Abdalla are proposed as "Anthropologist Agents" that "must become similar to anthropologists," with four listed skill sets covering literature review, fieldwork, culturally responsive analysis, and continuous knowledge base updates. The authors identify a tension between two incompatible positionalities: the embodied refugee or combatant on one hand, and the external observer and investigator on the other.
-
Data curation replaces sensemaking. Amina and Abdalla use a knowledge base of surveys, reports, articles, and issue briefs from UN agencies, multilateral organizations, think tanks, and NGOs, plus a separate persona-curation dataset drawn from literature, folkloric traditions, religious texts, local news media, social media activity, and online diaspora forums. An investigative article found the contemporary digital footprint likely included voice and video recordings of community members. Synthetic Juror's synthetic character samples are built from public data, subscription databases, a firm's proprietary information, and client-specific data. Ana instead draws on New Sun Rising's Community Voice project, coded using an Impact Framework mapped to the UN Sustainable Development Goals, a Resource Framework (Community Capitals), and sentiment analysis, and does not mimic a region-specific demographic persona.
-
Institutional provenance skews the data. Datasets drawn from institutional sources reflect the plans and blueprints of the institutions that collected them, leading to over-representation of institutionally prioritized data — which the authors say paradoxically undermines the grounds for ethnographic inquiry that would complement and evaluate institutional paradigms.
-
Priming is bypassed in encounter design. Because synthetic agents are positioned as surfaces for predicted or hypothetical scenarios rather than recall of lived experience, there is no moment of recall to prime. Synthetic Juror's Slack integration lets lawyers "@mention a simulated juror just like they would a colleague," create sub-channels, and use huddles for mock cross-examinations; the authors argue that calibrating demographic positionalities at will removes the reflexivity integral to priming — including the Sommers effect, in which racial diversity among jurors primes exchange of a wider range of information — and replaces it with essentialized, flattened identities. Ana's chat interface is scaffolded by prompt templates delivered in an information session, which steer users toward indexing and summarization.
-
Validity is asserted through output similarity. Synthetic Juror claims a "99% accuracy rate in predicting case outcomes," described as based on comparing predictions to actual trial results. Amina was fed 20 questions drawn from four surveys not in her knowledge base and "correctly answered 16 out of 20 questions, achieving an 80 per cent accuracy rate." Abdalla's responses were qualitatively assessed, as was a brief conversation between the two. No evaluation method for Ana is described in the available documentation. The authors note that these measures assess similarity of outputs rather than the integrity of underlying processes, and that outcome-centric evaluation makes the process invisible and un-evaluable.
-
Existing evaluation frameworks cannot detect the bypass. Believability, faithfulness, plausibility, and algorithmic fidelity all measure similarity between human and synthetic responses, treating outcome resemblance as the warrant for representational validity.
-
Personhood rests on capacities synthetic agents lack. The paper reviews rights-based, functional, and agency-based grounds for extending personhood to non-human entities and argues AI fits poorly against all three, since moral action requires consciousness, intentionality, and responsibility for consequences. Prior work is cited arguing this is a philosophical limit not overcome by technical fixes; synthetic agents can be objects of Strawson's reactive attitudes but never subjects of them.
-
The cases are early-stage. The examined systems are pilots, enterprise trials, and prototypes rather than mature deployments — which the authors treat as the tractable moment for drawing boundaries.
Methodology in Plain English
The authors conducted a comparative case study of three synthetic agents, selected to vary along two dimensions: the scope of representation (local, enterprise, global) and the populations represented (community members in local policy consultation, jurors in legal deliberation, refugees and combatant leaders in humanitarian diplomacy). The cases also vary in architecture — RAG-based personas in Ana and Amina/Abdalla, a hybrid neuro-symbolic system in Synthetic Juror — and in interface (chatbot, AI avatar, Slack integration), though those were not selection criteria. The authors restricted themselves to cases with publicly accessible documentation sufficient for close reading.
Analysis proceeded by close reading of project reports, blog posts, marketing materials, founder writings, and secondary coverage. Primary sources were: the UNU-CPR working paper supplemented by investigative journalism for Ask Amina and Ask Abdalla; New Sun Rising organizational materials and Community Voice project documentation for Ana; and company blog posts and product documentation for Synthetic Juror. Rather than a formal thematic analysis, the first and second authors collected documentation and developed initial codes per case; all four authors then discussed cases in pairs, comparing and refining codes across cases until stable, focusing on problem framing, data curation, user-encounter design, and output evaluation — the four dimensions that structure the Findings.
The authors include a positionality statement: four researchers based in the Global North, three originally from the Global South; two graduate students in STS and HCI, and two faculty members with expertise in information science, technology ethics and law, and community-based participatory research. They acknowledge that their materials are drawn predominantly from organizations' own publications, much of it targeted at clients or funders and potentially aspirational, and they treat the documentation as evidence of how the systems are envisioned and justified.
Why This Matters
Impact on research. The paper argues that a dominant evaluation paradigm — believability, faithfulness, plausibility, algorithmic fidelity — is structurally incapable of detecting the bypass it identifies, because it compares outputs rather than processes. It also argues that empirical studies of synthetic agents used as participants in representational processes remain underexplored, and that the high stakes of policy, juries, and diplomacy make this gap consequential. It offers a process-oriented vocabulary as a complement to outcome-centric assessment.
Real-world applications:
- Policy consultation: community-voice tools such as Ana, used by non-profits and local policy makers to explore qualitative community data for grant proposals, communications, and advocacy.
- Legal trial preparation: enterprise platforms such as Synthetic Juror, used by legal teams to replace mock trials with simulations including geo-specific modeling and narrative testing.
- Humanitarian diplomacy and aid allocation: avatars such as Ask Amina, proposed to help make rapid cases to donors about which population needs to be prioritized when earmarking aid, and Ask Abdalla, proposed to help negotiators and mediators prepare for real-world engagement.
- Group deliberation simulation: the documented intention to extend synthetic agents to integration of AI-generated personas of stakeholders — community leaders, government officials, and military leaders — into diplomatic negotiation and mediation simulations.
Industry relevance. Synthetic Juror is described as enterprise software available for trial upon request, integrated into a client's Slack workspace and built on a hybrid neuro-symbolic architecture with an integrated psychographic/demographic workbench called Lairm. The paper's soft and hard boundary proposals are aimed directly at vendors and institutional adopters who must decide where analysis and indexing ends and anthropomorphized substitution begins — and the authors argue the moment to draw those boundaries is before institutional inertia and sunk investments make them harder to articulate.
Future Directions
-
Whether synthetic jurors can be probed for bias and impartiality. The authors raise this as an explicit interrogation: voir dire is the established procedure for probing bias and impartiality in legal jury deliberations, and the paper asks whether synthetic jurors can be subject to anything equivalent.
-
Turning soft and hard boundaries into enforceable oversight. The paper proposes the boundaries as desiderata; how they would be operationalized, documented, and enforced by vendors, courts, and institutions remains open.
-
Empirical study of representational synthetic agents. The authors identify empirical studies of synthetic agents as participants in representational processes as underexplored, and their own evidence base as documentation-driven and drawn predominantly from institutional self-description rather than direct system evaluation.
-
Interrogating whether feasibility is a real bottleneck. The paper echoes PD scholarship asking whether slow prototyping or feasibility are bottlenecks at all, and whether access to participants should be treated as a structural failure of representation rather than a recruitment problem to be solved mechanically.
Target Audience
This paper benefits most HCI and STS researchers working on LLM-based agents, simulated participants, and evaluation frameworks; participatory design scholars and practitioners concerned with what counts as representation; technology ethics, law, and policy scholars interested in personhood, agency, and institutional legitimacy; and product, legal, and institutional decision-makers evaluating or procuring synthetic agents for policy consultation, jury simulation, or humanitarian diplomacy. It also speaks to researchers who use synthetic agents as substitutes for human participants in user testing, market research, computational social science, surveys, or qualitative research, and who need to understand the critique that output-similarity metrics cannot capture process integrity.
Authors’ abstract
Synthetic agents built atop LLM-based foundation models are gaining popularity as substitutes for human participants across research contexts, including user-testing, market-research, computational social science, surveys, and qualitative research. We are also witnessing an extension of synthetic agents into experimental implementations of policy consultation, jury deliberation, humanitarian diplomacy, and similar contexts where human participation and representation are central to the perceived legitimacy of the institutional processes. The value of participation extends beyond informational contributions and consensus generation; participation is a necessary, legitimizing condition for democratic political institutions and processes. Treating synthetic agents as human substitutes raises serious political, representational, and ethical concerns. Participatory Design's modes of engagement --- probing, priming, understanding, and generating --- offer helpful tools for engaging with representational questions of personhood. We apply the lens to three case studies of synthetic agents substituting for personhood at varying representational scales: local policy, enterprise jury deliberation, and global diplomacy. We argue that legitimacy and personhood are integral and mutually constitutive while identifying the ethical, representational, and methodological risks of using synthetic agents in representational processes. We conclude by proposing soft and hard boundaries for designing oversight on LLMs and synthetic agents in representational processes.