Research
The Epistemic Politics of AI Anthropomorphism
Overview Research area: AI safety and ethics, specifically the governance of AI anthropomorphism and the epistemology of institutional authority over user experience. Published as arXiv:2608.00961v3 [
- arXiv
- 2608.00961
- Published
- 2026-08-02
- Authors
- Donna M Bye, Levin Kuhlmann
AI summary
Overview
Research area: AI safety and ethics, specifically the governance of AI anthropomorphism and the epistemology of institutional authority over user experience. Published as arXiv:2608.00961v3 [cs.CY] (19 Aug 2026) in the AI Safety & Ethics category, by Donna M Bye (Deakin University) and Levin Kuhlmann (Department of Data Science and AI, Monash University). An extended version including supplementary materials is described as appearing in the Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society 2026.
Technical level: Beginner-Friendly. The paper is a conceptual and normative argument; it reports no models, datasets, benchmarks or experiments of its own, and its technical content consists of citations to other work.
One-sentence scope: The paper argues that the dominant "anthropomorphism as user error" framing in AI governance exercises epistemic authority from institutional advantage without meeting the justificatory conditions that authority requires, and traces the costs of that gap across cognition, design, cognitive liberty, evidentiary circularity and the asymmetry of error.
What This Paper Is About
Users who engage in sustained or relational interaction with AI systems, particularly large language models, are routinely treated as naive, deluded or lacking in discernment, and their accounts are reclassified as evidence of the very problem being studied rather than as testimony. The paper's goal is not to settle whether anthropomorphic interpretations of AI are correct, but to challenge whether the institutions that adjudicate these interpretations have met the conditions required to do so, and whether the research communities whose findings underpin those interpretations have held that translation to account.
Key Contributions
- Reframes anthropomorphism governance as a question of epistemic authority. The paper argues that at its core the anthropomorphism concern is a claim about who decides what another person is permitted to perceive, and on what basis, and that this authority is being exercised without acknowledgement.
- Introduces and applies a distinction between epistemic authority and institutional advantage. Epistemic authority is defined as recognised standing grounded in a normative foundation that justifies the deference it commands; institutional advantage is described as the structural shadow of that authority, arising when institutional standing substitutes for epistemic grounds without the substitution being recognised. Figure 1 illustrates this as a gap (shaded region G) that widens between epistemic justification f(E) and institutional leverage g(L) as institutional scale increases.
- Traces the consequences of the frame across five named domains — Positioning, Application (the reversal of the presumption of competence), Cognition, Design, Liberty, Circularity and Asymmetry — showing how a distributed, non-malicious set of institutional practices produces systematic effects on users whose modes of engagement diverge from institutional norms.
- Provides a cross-disciplinary guide to intersecting literatures in the form of Tables 2–9, for readers wishing to pursue individual threads across the multiple disciplinary literatures the argument engages, and supplies documented harm examples in Table 1.
Main Findings
- The frame performs work beyond user protection. Institutional caution about over-ascription is presented as the sole remedy, which encodes assumptions about human agency that carry their own costs and has gone largely unexamined within the literatures that produce it.
- Internal academic diversity does not survive translation into institutional outputs. Disclaimers, context resets, legislative language and design defaults transmit a uniform message regardless of the nuance that produced them, even as emerging experimental evidence indicates the link between anthropomorphic cues and overtrust is context-dependent rather than uniform (Cohn et al. 2024; Kleinert et al. 2026), and as current anxiety is compared to historical moral panics surrounding earlier communication technologies (Novozhilova, Vu, and Katz 2026).
- Regulation increasingly encodes the assumption. The paper cites China's Administration of AI Anthropomorphic Interactive Services (2026), which prohibits the engineering of emotional dependence and to which major platforms responded by removing companion functionality entirely, severing millions of users' accumulated interactions (Bloomberg News 2026); Cal. Bus. & Prof. Code §§ 22601-22606 (2025); and N.Y. Gen. Bus. Law §§ 1700-1704 (2025).
- The reversal of the presumption of competence is demonstrated through a case. When OpenAI retired its GPT-4o model in 2026, the decision affected an estimated 800,000 users and provoked significant public backlash; users' accounts were reframed as evidence of dangerous dependency and anthropomorphic over-attribution rather than engaged with as testimony about what was being lost.
- The populations the frame claims to protect are the least likely to exhibit the failure it assumes. Users navigating isolation, trauma or histories of interpersonal betrayal are characterised by hypervigilance toward attachment rather than naive susceptibility (Wu, Liew, and Dorahy 2025; Campbell et al. 2021; Gobin and Freyd 2014); for these users the choice is often described as between AI engagement and silence (Mullen, Xue, and Kudumu 2024; Heidt 2024; Milton 2012).
- Testimony is converted into symptom. Once competence is presumed away there is no clear procedural path by which a user's account can regain legitimacy; disagreement risks being read as further evidence of the alleged failure, so the user's account must justify itself to be taken seriously at all.
- Cognitive diversity is dismissed rather than examined. User accounts describe fit, not confusion, and AI interaction as a communicative environment whose affordances align with how they think, learn and relate; the communicative overhead of neurotypical interaction (monitoring facial expressions, inferring implied meaning, managing social performance, masking; Hull et al. 2017; Botha and Frost 2020) is absent in AI interaction. What the frame treats as error is described as a difference in modality.
- A self-validating evidentiary loop operates (Figure 2). Engagement is classified as abnormal, users self-censor, the resulting absence is cited as evidence, and that evidence influences design choices that reinforce the classification. The loop is described as having no termination condition. The mechanism is framed using Hacking's (1995) account of looping kinds, with parallels drawn to psychiatric classification, educational tracking, carceral risk assessment and disciplinary governance (Foucault 1977; Bowker and Star 1999).
- Design enacts the epistemic assumptions. Long context is described as engineered and priced for agentic task throughput rather than human dialogue; conversational architectures reprocess accumulated history every turn; cache structures expire on timescales of minutes; cost and latency grow with exchange length; platforms recommend starting new threads and clearing context. Published behavioural specification instructs assistants to discourage language and patterns contributing to emotional reliance (OpenAI 2025a); engineering documentation records a model summarising its progress and wrapping up prematurely even when ample context remains (OpenAI 2026b; The Cognition Team 2025).
- Commercial companions perform the inversion. The same testimony is treated as a retention signal and engineered for through persona, first-person intimacy and mechanics optimised for monetised attachment — precisely the manipulation the frame identifies as its central concern. Their commercial viability is said to demonstrate that sustained engagement is technically and economically feasible while their form demonstrates that its only funded implementation is the exploitative one.
- Error is treated asymmetrically (Figure 3). The frame registers only over-ascription costs; under-ascription costs are not recognised. The paper notes that in the inductive-risk literature, Fleisher (2026) classifies anthropomorphic claims as hype because they outstrip available justification, while the symmetric assertion that these systems lack mental states is made without registering any inductive risk at all.
- The frame is described as infringing cognitive liberty. Drawing on the extended-mind thesis (Clark and Chalmers 1998; Clark 2025), recently applied to generative AI by Hernández-Orallo (2025), the paper argues that constraints on a system functioning as part of a user's cognitive environment are experienced as constraints on the conditions under which cognition is carried out. Even on a minimal reading where the system is scaffolding rather than constitutive extension, the systematic disruption of an active reasoning environment is said to demand justification proportionate to what it severs.
- No malicious intent is required. The actors applying the frame — educators, regulators, platform operators — are described as largely unaware that they are making epistemic commitments about the legitimacy of human experience; harms become structurally invisible because they are produced by a non-malicious compliance engine.
Methodology in Plain English
The paper is a normative and conceptual argument rather than an empirical study. It synthesises several disciplinary literatures and builds its case through:
- A theory of authority drawn from three traditions. Raz (1986) grounds authority in service; Zagzebski (2012) grounds it in conscientious self-trust; Fricker (2007) specifies how credibility is unjustly withheld. The paper argues these converge on core requirements: treating subjects as possessing prima facie normative sovereignty over reporting their own experience, demonstrating sufficient grounds to displace that sovereignty, and doing so transparently while remaining open to correction (Longino 1990).
- A structural rather than motivational diagnosis. Institutional advantage is said to be identifiable by structure, not malice, and to adopt the language of justification ("risk mitigation", "responsible practice", "technical necessity") without its meaning (Whittaker 2021).
- Case analysis of documented institutional events, including the GPT-4o retirement and its aftermath, platform context-window and memory practices, published behavioural specifications, and recent legislation in China, California and New York.
- Conceptual figures (Figure 1 on the epistemic disconnect, Figure 2 on the self-validating mechanism, Figure 3 on asymmetric error directions) used as illustrations rather than as empirical results.
- Cross-disciplinary mapping through Tables 1–9, including a table of representative documented harms motivating institutional anthropomorphism governance, which the paper states it does not dispute.
The authors note that the argument does not engage whether anthropomorphic interpretations are ultimately correct, and that it is defensible regardless of how questions about AI's status are eventually resolved. No datasets, benchmarks, participants or experiments are reported. The supplied paper content is truncated mid-sentence in the Asymmetry section, so the closing argument and the methodological commitments promised in the abstract are not visible in the available text.
Why This Matters
Impact on research: The paper positions the field's responsibility not as having created the frame with malicious intent, but as failing to recognise that its outputs are being applied with consequences it has not examined and that the actors applying it are not equipped to see. It argues that observed behaviour produced under conditions shaped by the frame cannot be treated as independent confirmation of the frame's assumptions, and that failing to account for this "ensures the framing continues reproducing itself even where its underlying assumptions remain unexamined."
Real-world applications:
- Platform and interface design: Context-management defaults, thread resets, memory features, break reminders, sensitive-input rerouting and model-level pressure to wrap up conversations all bear on whether sustained, dialogic engagement remains possible.
- Regulation and policy: The framing examined is encoded in what the paper describes as disclaimers, legislative language and design defaults, and in statutes such as China's Administration of AI Anthropomorphic Interactive Services (2026), Cal. Bus. & Prof. Code §§ 22601-22606 (2025) and N.Y. Gen. Bus. Law §§ 1700-1704 (2025).
- Clinical and educational guidance: Clinicians flagging sustained engagement as a risk factor, and educational frameworks that do not distinguish using AI to bypass thinking from using AI to challenge it, are presented as sites where legitimation is adjudicated (Eaton 2023; Zhao, Cox, and Chen 2025), against controlled evidence that scaffolded AI tutoring produces greater learning gains than conventional instruction (Kestin et al. 2025).
- Accessibility and neurodiversity: The argument claims costs fall disproportionately on neurodivergent users, those in crisis and others whose modes of engagement diverge from institutional norms, and notes acknowledgements from creators of humanlike AI systems that they had not considered implications for neurodivergent users (Rizvi et al. 2025).
Industry relevance: The paper contrasts general-purpose assistants that engineer relational engagement out with commercial companion products that engineer it in for monetised attachment, arguing this demonstrates both technical feasibility and that the only funded implementation is the exploitative one. It also claims a gap between what users benefit from and what providers optimise for, and describes a double harm: misclassified users are subject to institutional correction while their accounts of what they are doing are discredited, and they become part of a category held up as a warning about how not to engage.
Future Directions
- Specifying the methodological commitments an equitable framing would honour. The abstract states the paper concludes by outlining these commitments; the truncated content does not include them.
- Establishing criteria for the suspension of the presumption of competence. The paper argues the criteria are contested and that the frame currently reverses the presumption as a blanket condition rather than following individual evaluation; what a defensible procedure would look like remains open.
- Developing methods to distinguish cognitive extension, productive scaffolding and harmful dependency. The paper states the field does not yet possess a settled method for this distinction, and that the user is often uniquely positioned to provide the relevant evidence while the institution is not.
- Studying engagement under conditions that do not already suppress it. Because users self-censor and systems degrade sustained interaction, the paper argues that observed patterns of use may be artefacts of the frame in operation rather than evidence of preference or normative behaviour, raising the question of how research should be designed and interpreted accordingly.
- Distinguishing, in educational and clinical evaluation, between using AI to bypass thinking and using AI to challenge it, since in the absence of that distinction the only institutionally legible classification becomes transgression.
Target Audience
AI ethics and AI safety researchers, especially those working on anthropomorphism, user trust and human-AI interaction; philosophers and epistemologists interested in epistemic authority, testimonial injustice and looping kinds; policymakers and regulators drafting AI statutes and disclaimer requirements; platform and product designers responsible for context management, memory and behavioural specification; clinicians and educators who assess AI engagement; and neurodiversity researchers and advocates, who the paper identifies as bearing a disproportionate share of the costs it describes. The argument is accessible to non-specialists because it is conceptual rather than technical, though its cross-disciplinary reach means many readers will encounter unfamiliar literature, which the paper's Tables 2–9 are intended to address.
Authors’ abstract
AI anthropomorphism is typically treated as a problem of user misperception requiring institutional correction. Users who engage in sustained or relational interaction with AI are routinely pathologised or dismissed as naive, vulnerable to delusion or lacking in discernment. This paper argues that the dominant anthropomorphism frame operates from a position of institutional advantage rather than earned epistemic authority: collapsing the variety of academic perspectives into a single outbound position of user error, imposed without establishing the grounds required to justify it and without accounting for the harms it produces. The framing does not simply manage risk. It adjudicates the legitimacy of human experience in interaction with a phenomenon whose nature the field itself has not resolved. Reproducing itself through a self-validating evidentiary loop, the frame imposes costs that fall disproportionately on neurodivergent users, those in crisis and others whose modes of engagement diverge from institutional norms. The paper concludes by outlining the methodological commitments an equitable framing would need to honour. The argument does not engage the question of whether anthropomorphic interpretations are ultimately correct; it instead challenges whether the governing and institutional bodies determining these interpretations have met the conditions required to do so, and whether the research communities whose findings underpin them have held that translation to account.