Skip to content
AI.info

Research

What Can Artificial Intelligence Learn from Medicine? Generative Analogies and Reliable Machine Learning Systems

Overview Research area: Philosophy of AI and epistemology of machine learning, drawing on philosophy of science (analogy and modeling) and on the epistemology of clinical translation in medicine. Tech

arXiv
2608.18186
Published
2026-08-18
Authors
Emanuele Ratti, Lena Zuchowski

AI summary

Overview

  • Research area: Philosophy of AI and epistemology of machine learning, drawing on philosophy of science (analogy and modeling) and on the epistemology of clinical translation in medicine.
  • Technical level: Intermediate. No mathematics or empirical experiments; the paper works with conceptual distinctions from epistemology (warrant, reliabilism) and from the philosophy of medicine. Familiarity with basic ideas about justification and with debates about trustworthy AI will help, but the argument is readable without a technical ML background.
  • Scope: The paper argues that the widely invoked parallel between clinical translation and ML development should be treated as a generative analogy, and uses that treatment to articulate a reliabilist account of what makes ML systems epistemically and methodologically warranted.

What This Paper Is About

Machine learning is now used extensively in medicine with some success, but it remains unclear what justifies confidence in ML systems — what makes their outputs and development processes epistemically and methodologically legitimate rather than merely convenient. Commentators have responded by drawing a parallel with medicine, suggesting that ML should adopt the standards that govern clinical translation (moving a finding from the lab into safe, warranted clinical use). The paper takes that parallel seriously and asks what it actually licenses: it clarifies which warrants of clinical translation are being gestured at when the analogy is invoked, and works out how those warrants transfer, analogically, to ML.

Key Contributions

  1. Reframes the medicine–ML parallel as a generative analogy. Rather than treating the comparison as a loose metaphor or a rhetorical appeal, the authors use tools from Hesse's work on analogy in science to characterize it as an analogy that can generate claims about the target domain (ML), not just illustrate it.
  2. Makes explicit the warrants that are usually only implied. The paper identifies more precisely the epistemic and methodological warrants of clinical translation that are typically just mentioned in passing when the analogy is invoked.
  3. Specifies how those warrants carry over to ML. It shows in which sense the warrants of clinical translation apply analogically to the context of building ML systems, rather than assuming the transfer is straightforward.
  4. Develops a new form of ML reliabilism. By interpreting the warrants of clinical translation in reliabilist terms, the authors sketch a reliabilist account of ML that is distinct from — though compatible with — existing reliabilist accounts in the philosophy of AI.

Main Findings

  • The analogy is generative, not merely illustrative: On the authors' reading, the medicine–ML comparison is the kind of analogy that can yield new claims about how ML systems should be built and evaluated, which is why it deserves careful analysis rather than dismissal as a loose figure of speech.
  • Clinical translation carries warrants that are usually left unstated: The paper claims that appeals to clinical translation in the ML literature typically name the analogy without spelling out the epistemic and methodological warrants that make clinical translation work; the authors identify these explicitly.
  • Those warrants transfer analogically to ML: The paper argues that the warrants of clinical translation apply to the ML context in a specific, analogical sense — the transfer is qualified, not a claim that ML is simply like medicine.
  • Reliabilism is the natural reading of those warrants: Interpreting the warrants of clinical translation in reliabilist terms is presented as the key move that makes them usable for thinking about ML.
  • A distinct ML reliabilism results: The account the authors develop is claimed to be a new form of ML reliabilism, differentiated from existing reliabilist accounts in philosophy of AI, while remaining compatible with them.
  • Not an empirical result: The abstract reports no experiments, benchmarks, or quantitative findings; the contributions are conceptual and argumentative, and the details of the argument are not available in the abstract.

Methodology in Plain English

This is a conceptual paper, not an experimental one. The authors begin from an existing observation in the literature — that ML should take its standards from clinical translation — and then do two things. First, they borrow tools from Hesse's work on how analogies function in science in order to describe what kind of comparison this is and what it is entitled to do. Second, they unpack what the phrase "clinical translation" actually commits you to: what has to be true, and what has to be done, for a clinical finding to count as properly warranted when it reaches practice. Having laid that out, they ask which of those commitments have recognizable counterparts in the process of building and validating an ML system, and they phrase the answer in reliabilist vocabulary — that is, in terms of whether a process reliably produces the right kind of output. The result is a proposed framework rather than a tested one.

Why This Matters

Impact on research. The paper pushes the medicine–ML comparison from a slogan into an analyzed analogy, which gives philosophers and ML researchers a shared vocabulary for arguing about what "trustworthy" or "warranted" ML should mean. It also connects two literatures that often talk past each other: philosophy of AI (especially reliabilism) and the epistemology of clinical translation.

Real-world applications (directions this framing points toward, not demonstrated in the abstract):

  • Regulatory and approval pathways for medical ML tools, where the question is what evidence justifies deployment.
  • Clinical validation and post-deployment monitoring of ML systems, where the analogue of clinical translation is deciding when a model is ready for practice and how to keep it warranted over time.
  • Documentation and audit practices for ML in high-stakes settings, if warrants are treated as things to be specified and checked rather than assumed from benchmark performance.
  • Training and professional standards for developers and clinical users, if warranted practice is understood in terms of reliable processes rather than outcomes alone.

Industry relevance. Firms deploying ML in regulated sectors face exactly the question the paper addresses — what justifies confidence in a system — and the reliabilist framing offers a way to talk about process reliability rather than just accuracy claims. The abstract does not report any industry study or evaluation, so the relevance here is conceptual rather than demonstrated.

Future Directions

  • Spelling out the ML reliabilism in detail. The abstract announces a new form of ML reliabilism but does not state its criteria, so its content and its precise differences from existing reliabilist accounts in philosophy of AI remain to be examined.
  • Testing the analogy's limits. Since an analogy can only carry so far, the natural next question is where the medicine–ML comparison breaks down and what happens to the transferred warrants at those points.
  • Operationalizing the warrants. Turning epistemic and methodological warrants from abstract standards into something usable in development, review, or regulation is an open practical task the framing invites.
  • Bringing in empirical and normative work. The abstract contains no empirical findings; connecting the account to actual ML development practices, regulatory frameworks, and case studies of medical ML would be a subsequent step.

Target Audience

Philosophers of AI, epistemologists interested in reliabilism and in machine learning, and researchers in the ethics and governance of AI. It is also aimed at ML practitioners and clinical informatics professionals who work on medical ML and who want a principled account of what justifies trust in a system, as well as regulators and policy analysts concerned with how evidence standards from medicine might be adapted for AI. Readers looking for empirical results, benchmarks, or technical ML methods will not find them here — the contribution is conceptual.

Authors’ abstract

In the past few years, machine learning (ML) has been widely (and to an extent, successfully) implemented in medicine. However, uncertainties surrounding ML have made it difficult to establish the bases of its epistemic and methodological warrants. In the literature, a parallel has been drawn between medicine and ML, suggesting that we should model epistemic and methodological standards for ML on the standards of clinical translation. By developing tools from Hesse work, we characterise the nature of this parallel as a generative analogy between the process of clinical translation and the process of building ML systems. We identify more precisely the epistemic and methodological warrants of clinical translation that are typically only mentioned when appealing to the analogy, and we show in which sense such warrants apply analogically to the context of ML. In particular, we interpret warrants of clinical translation in reliabilist terms, and we show how this can inform a new form of ML reliabilism, which is distinct from (though compatible with) existing reliabilist accounts in philosophy of AI.

Read the original paper