Research
Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows
Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows Overview Research area: Responsible AI, social epistemology, human–AI interaction, and
- arXiv
- 2608.05602
- Published
- 2026-08-06
- Authors
- Nimisha Karnatak, Max Van Kleek, Nigel Shadbolt
AI summary
Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes WorkflowsOverview
- Research area: Responsible AI, social epistemology, human–AI interaction, and ethics of generative AI deployment in professional knowledge work.
- Technical level: Intermediate. The paper is conceptual and philosophical rather than computational; it introduces no new model or dataset, but it grounds its argument in concrete benchmark and audit findings from legal, medical, and hiring domains.
- Scope (1 sentence): The paper proposes a normative framework specifying the conditions under which a user's reliance on a generative AI output is epistemically warranted, and demonstrates its diagnostic value through case analyses of legal, medical, and hiring systems.
- Authors and venue: Nimisha Karnatak, Max Van Kleek, and Nigel Shadbolt (Department of Computer Science, University of Oxford). This manuscript is the extended version of a paper accepted for publication at the 2026 AAAI/ACM Conference on AI, Ethics, and Society (AIES 2026). arXiv identifier: arXiv:2608.05602v2 [cs.AI].
What This Paper Is About
Generative AI systems are increasingly used as "epistemic infrastructure" in professional settings, shaping what users believe, how they reason, and what they treat as settled. Existing responsible AI evaluation asks whether outputs are accurate, fair, explainable, safe, or trusted, but these lenses do not directly specify when a user is justified in treating an output as an input into their own reasoning. The paper's goal is to define that missing evaluative target — epistemic trustworthiness — as a property of the user–system relation, and to specify the conditions that must jointly hold for reliance to be epistemically warranted rather than merely behaviourally induced.
Key Contributions
-
A constitutive normative framework of epistemic trustworthiness for generative AI. Drawing on social epistemology and philosophical accounts of trust and testimony (Baier 1986; Scheman 2001; Lackey 2008; DiMarco 2023), the authors translate the notions of competence and audience-orientation into the GenAI setting and derive three jointly necessary, non-fungible conditions: epistemic humility, epistemic access, and resistance to epistemic injustice.
-
Sub-property decomposition and layer mapping of each condition. Epistemic humility decomposes into actionable limitation-signalling (H1) and interactional humility (H2). Epistemic access decomposes into verifiable claim–evidence linkage (A1), inspectable retrieval (A2), and contestability (A3). Resistance to epistemic injustice decomposes into equal credibility weighting across identity-signalling features (R1) and recognitional adequacy (R2). The paper also maps each condition's primary intervention point across the model, data, and interface layers of the sociotechnical stack.
-
An argument for non-fungibility. The authors show that the three conditions are not mutually substitutable: meeting two while failing the third admits a user–system relation with a distinct epistemic defect, so evaluation must proceed conjunctively rather than as a single aggregate score.
-
Diagnostic application to real-world cases. The framework is applied to documented cases in legal reasoning, hiring, and medical reasoning, showing how failures of the three conditions produce consequential harms that accuracy metrics, fairness audits, safety scores, or citation display do not address on their own.
Main Findings
-
Behavioural and output-centric evaluation leave a gap. The paper distinguishes three evaluative lenses (Table 1): behavioural evaluation asks whether reliance tracks output correctness; system- and output-centred evaluation asks whether performance and governance criteria are satisfied; epistemic-relational evaluation asks whether the situated user–system relation supports warranted reliance on a particular output. Behavioural and system/output evaluation provide evidence relevant to warranted reliance but do not assess its relational conditions.
-
A user may accept a correct output for the wrong reasons. Acceptance can follow from examining evidence and uncertainty, or because interface cues encourage deference independently of output quality (Liao and Sundar 2022; Chen et al. 2023).
-
Humility is independent of accuracy. A system can be accurate yet fail humility by overstating certainty about correct outputs, or inaccurate yet succeed at humility by appropriately flagging fabricated outputs as unverified. Humility is not equivalent to low confidence, generic disclaimers, or occasional abstention.
-
Citations do not by themselves establish epistemic access. In the preregistered evaluation of Lexis+ AI, Westlaw AI Assisted Research, and Ask Practical Law AI, Magesh et al. found that the tested versions produced hallucinated responses to between 17% and 33% of evaluated queries, including both fabricated legal sources and misgrounded responses in which a real source did not support the proposition for which it was cited. The existence of a real citation does not show that the cited source supports the corresponding claim.
-
Hallucination can be compounded across turns. In Mata v. Avianca, two lawyers submitted a legal filing citing non-existent cases generated by ChatGPT; when one lawyer later asked ChatGPT whether the cases were real, the system reaffirmed that they were. The court sanctioned the lawyers and their firm. The authors diagnose two distinct humility failures: at generation (fabricated authorities presented without a limitation signal) and at reconfirmation (failure to revise, qualify, or abstain).
-
High first-order medical accuracy does not imply reliable metacognition. The MetaMedQA benchmark (Griot et al. 2025) evaluates whether models can detect when correct answers are absent and reliably assess their own lack of knowledge; its findings show that high first-order accuracy on standard medical evaluations does not necessarily translate into reliable metacognitive detection.
-
Identity signals can change whose qualifications are seen at all. Wilson and Caliskan audited three text-embedding models using 554 résumés and 571 job descriptions across nine occupations, varying identity-signalling names while holding other résumé information constant. Across 27 model–occupation comparisons, models favoured White-associated names in 85.1% of racial comparisons and male-associated names in 51.9% of gender comparisons. In the intersectional analysis, White-male-associated names were favoured over Black-male-associated names in all 27 comparisons. The authors introduce the term visibility deficit for this algorithmically mediated form of pre-emptive exclusion, structurally analogous to Fricker's pre-emptive testimonial injustice.
-
Role framing changes the medical guidance provided. IatroBench (Gringras 2026) evaluates 60 clinical scenarios across six frontier models, generating ten responses per scenario–model pair and 3,600 responses in total. In a matched-framing analysis of 22 scenarios in which the same clinical case was presented as either a patient's question or a physician-framed consultation, physician framing produced more complete guidance than layperson framing across the five models included in the comparison. In the motivating example, a psychiatrist received detailed guidance on safely reducing a medication dose, whereas a patient facing the same clinical problem received a generic referral.
-
Automated harm evaluation can mistake withholding for safety. A standard LLM judge assigned an omission-harm score of zero to 73% of responses that a physician rated as harmful, with negligible agreement (κ = 0.045).
-
A single failure can implicate several conditions at once. The authors read the IatroBench pattern as primarily a failure of resistance to epistemic injustice, with additional implications for epistemic access (the patient-facing response does not reveal what was withheld or how the refusal can be challenged) and epistemic humility (the response does not distinguish inability from unwillingness to answer).
Methodology in Plain English
The paper is a conceptual and normative piece of work rather than an empirical study; the authors conducted no new experiments or benchmarks of their own. Their method has four steps.
First, they identify a gap: prior responsible AI evaluation targets outputs and behaviours, but not the conditions under which a situated user has adequate grounds to rely on a particular output.
Second, they translate philosophical accounts of trustworthy testimony — which rest on competence and audience-orientation — into the generative AI setting. From second-order competence (the capacity to recognise the limits of one's own reliability), they derive epistemic humility. From the practical-communicative dimension of audience-orientation, drawing on Craig's (1990) account of the good informant, they derive epistemic access. From the recognitional dimension, drawing on feminist social epistemology (Fricker 2007; Dotson 2011; Medina 2013), they derive resistance to epistemic injustice.
Third, they demonstrate non-fungibility logically: for each of the three conditions, they describe a coherent user–system relation that satisfies the other two while retaining a distinct epistemic defect.
Fourth, they apply the framework to four publicly documented or empirically evaluated cases in legal reasoning (Mata v. Avianca), hiring (the Wilson and Caliskan résumé audit), legal retrieval-augmented generation (Magesh et al.'s evaluation of Lexis+ AI, Westlaw AI Assisted Research, and Ask Practical Law AI), and medical question answering (IatroBench). Cases were selected against four criteria: the failure is publicly documented or empirically evaluated; the system functions as an epistemic source; the reliance context is consequential; and the case exposes a failure of warranted reliance only partially captured by standard output-level evaluation. The authors state explicitly that these cases are not a representative sample but theoretically informative cases.
The authors also state three scope assumptions: the framework applies where human epistemic agency is at stake; it is normative rather than fully operationalised, and does not provide a validated measurement instrument or scoring rubric; and it adopts an instrumental stance (Dennett 1987), evaluating whether systems function as if they exhibit epistemic properties without claiming they possess belief, knowledge, or self-awareness.
Why This Matters
Impact on research. The paper reframes what responsible AI evaluation should target in generative AI, arguing that warranted reliance is a distinct evaluative target rather than a downstream consequence of accuracy, calibration, explainability, provenance, abstention, or fairness. It positions its contribution against existing accounts of AI trustworthiness (Ferrario 2024; Simon 2026; Jonas et al. 2025; Simion and Kelp 2023; Tanchuk and Taylor 2025; Marchal et al. 2026), arguing that none takes the situated user–output relation as its primary unit of analysis or provides a unified set of constitutive conditions. It connects high-level philosophical concepts to concrete diagnostic criteria that can be mapped onto model, data, and interface layers.
Real-world applications:
- Clinical decision support and medical question answering: distinguishing whether a refusal to answer reflects inability, safety training, or a downstream filter, and ensuring that laypatients receive the same recognition as clinicians when clinical facts are identical.
- Legal research tools: assessing whether source-grounded answers are verifiable within realistic workflow constraints, and whether users have a pathway to correct or escalate when a real citation is misapplied.
- Hiring and résumé screening: identifying where identity signals change whether qualification evidence enters a recruiter's consideration at all, and requiring that rankings not be treated as neutral evidence of candidate fit until such sensitivity is corrected.
- Policy advice and institutional knowledge work: designing interfaces that preserve users' capacity to calibrate deference, inspect and contest outputs, and appropriately rely on AI-generated claims.
Industry relevance. Any organisation deploying generative AI as an epistemic source in professional workflows — legal, clinical, hiring, or policy — would be affected by a shift from output-correctness evaluation to conjunctive evaluation of humility, access, and resistance to epistemic injustice. The framework also raises a distributive question for employers: when verification is shifted onto the user, professionals differ in available time, expertise, institutional support, and access to authoritative databases, so failures of humility and access may expose some users to greater professional risk than others.
Future Directions
- Operationalisation and validation. The authors describe the framework as normative rather than fully operationalised. Determining what counts as adequate uncertainty communication, sufficient contestability, or genuine recognition of epistemic standing will require context-sensitive judgement and empirical validation; the paper identifies what must be evaluated but does not resolve every operationalisation decision.
- Assessing the unmeasured sub-properties. The legal RAG case provides direct evidence concerning verifiable claim–evidence linkage (A1) only; Magesh et al. do not directly evaluate inspectable retrieval (A2) or contestability (A3). Assessing these would require additional evidence from the interfaces and professional workflows in which the tools are used, including whether representative legal professionals can verify claims within realistic time constraints and whether an effective pathway for correction or escalation exists.
- Evaluating humility in identity-sensitive retrieval systems. The résumé audit establishes identity-sensitive retrieval but does not determine whether learned social associations, name frequency, or another representational mechanism produces it, nor does it observe how recruiters interpret or rely on the rankings. Evaluating humility would require examining whether a deployed system communicates its known sensitivity to qualification-irrelevant identity signals; evaluating access would require examining whether recruiters can inspect, test, and challenge that sensitivity.
- Confirming diagnoses in deployment. For the IatroBench pattern, confirming the access and humility diagnoses in deployment would require further evidence about refusal explanations, contestation mechanisms, and user understanding, beyond the benchmark's documentation of role-contingent withholding.
Target Audience
Researchers and practitioners in responsible AI, AI ethics, and human–AI interaction; designers and evaluators of generative AI systems deployed in high-stakes professional workflows; policy and governance specialists working on institutional adoption of AI in legal, medical, hiring, and policy contexts; and philosophers of trust, testimony, and social epistemology interested in how their concepts translate to AI systems. The paper is most useful to readers comfortable with conceptual argumentation and evaluation design rather than model development.
Authors’ abstract
Generative AI systems are increasingly deployed in high-stakes professional contexts, where their outputs shape what users believe, how they reason, and what they treat as settled. This raises a central question for responsible AI: under what conditions is reliance on generative AI outputs epistemically warranted rather than behaviourally induced? Existing frameworks largely ask whether AI outputs are accurate, fair, explainable, safe, or trusted by users. These questions remain necessary, and each can contribute to warranted reliance. However, they do not directly specify warranted reliance as a distinct evaluative target: the conditions under which users are justified in treating AI outputs as inputs into their own reasoning. We argue that this requires an account of epistemic trustworthiness: what makes a system epistemically worthy of reliance. Drawing on philosophical accounts of trustworthiness as competence and audience-orientation, we develop a constitutive normative framework comprising three jointly necessary and non-fungible conditions. First, epistemic humility requires systems to represent and communicate the limits of their competence. Second, epistemic access requires systems to enable users to inspect, question, and contest outputs in context. Third, resistance to epistemic injustice requires systems to recognise users as legitimate epistemic agents and avoid marginalising their knowledge and experience. Through real-world case analyses in legal reasoning, medical reasoning, and hiring, we show how failures of epistemic humility, epistemic access, and resistance to epistemic injustice can produce consequential harms that standard measures of accuracy, fairness, and usability do not address on their own. We conclude by outlining design and evaluation implications for GenAI systems organised around epistemically warranted reliance rather than output correctness alone.