Research
Grounding Clinical AI Competency in Human Cognition Through the Clinical World Model and Skill-Mix Framework
Overview Research area: Clinical artificial intelligence, with a focus on how AI competency is defined, evaluated, and bounded — drawing on principles of clinical cognition and human decision-making.
- arXiv
- 2604.08226
- Published
- 2026-04-09
- Authors
- Seyed Amir Ahmad Safavi-Naini, Elahe Meftah, Josh Mohess, Pooya Mohammadi Kazaj, Georgios Siontis, Zahra Atf, Peter R. Lewis, Mauricio Reyes, Girish Nadkarni, Roland Wiest, Stephan Windecker, Christoph Grani, Ali Soroush, Isaac Shiri
AI summary
Overview
Research area: Clinical artificial intelligence, with a focus on how AI competency is defined, evaluated, and bounded — drawing on principles of clinical cognition and human decision-making.
Technical level: Intermediate. The paper is a conceptual and theoretical framework rather than an empirical study; readers need familiarity with clinical AI evaluation and human-factors terminology, but no specialized mathematics.
Scope (one sentence): The paper proposes a formal model of the clinical world and a multi-dimensional skill framework to specify where, for whom, and under what authority a clinical AI system's competency has actually been demonstrated.
What This Paper Is About
Clinical AI currently has no shared formal account of the world it operates in — a gap that limits how competency can be defined or compared. Existing efforts address evaluation, regulation, or system design separately, with nothing connecting them. The paper's goal is to supply that missing common structure so that claims about AI performance can be stated precisely and bounded honestly.
Key Contributions
- The Clinical World Model. A framework that formalizes care as a tripartite interaction among Patient, Provider, and Ecosystem, giving clinical AI a shared referent that evaluation, regulation, and design can all point to.
- Parallel decision-making architectures. Matching architectures for providers, patients, and AI agents that describe how each transforms information into clinical action, grounded in validated principles of clinical cognition rather than in AI-specific assumptions.
- The Clinical AI Skill-Mix. A set of eight dimensions that operationalize competency: five defining the clinical competency space (condition, phase, care setting, provider role, task) and three specifying how AI engages human reasoning (assigned authority, agent facing, anchoring layer).
- An irreducibility claim for competency. Because the dimensions combine into a very large space of distinct coordinates, validation in one coordinate is argued to provide minimal evidence of performance in another.
Main Findings
- Competency is bounded by the world model. The paper's opening premise is that any intelligent agent's competency is limited by its formal account of the world it operates in, and that clinical AI lacks such an account.
- Care is formally tripartite. The Clinical World Model treats care as an interaction among Patient, Provider, and Ecosystem rather than as a one-directional AI-to-clinician pipeline.
- Human and artificial decision-making are modeled in parallel. Rather than treating AI as a special case, the framework describes providers, patients, and AI agents through comparable information-to-action architectures.
- Competency is a coordinate, not a single score. The eight Skill-Mix dimensions combine to yield a space of billions of distinct competency coordinates, per the abstract.
- Validation does not transfer across coordinates. The central structural implication is that evidence gathered in one coordinate gives minimal support for another, making the space irreducible.
- The field's central question shifts. Instead of asking whether AI works, the framework asks in which competency coordinates reliability has been demonstrated, and for whom.
- No empirical results are reported in the abstract. The abstract describes a framework and its structural implications; it contains no datasets, benchmarks, experiments, or comparative baselines, and none should be inferred.
Methodology in Plain English
This is a conceptual paper, not an experimental one. The authors take established principles of clinical cognition — how clinicians actually reason and act under real care conditions — and use them as the foundation for describing how any agent, human or artificial, turns information into clinical action. They then build a formal vocabulary: first a model of the clinical world (Patient, Provider, Ecosystem), then parallel decision architectures for the three agent types, then a set of eight dimensions that together pin down exactly what a competency claim is about. The abstract reports no study design, data collection, or evaluation procedure, because the contribution is the framework itself.
Why This Matters
Research impact. The framework offers a shared vocabulary for a field that currently evaluates AI systems in ways that are hard to compare. If competency is genuinely coordinate-dependent, then benchmark results reported without specifying condition, phase, care setting, provider role, task, authority, agent facing, and anchoring layer are under-specified — a claim that would reshape how evaluation studies are designed and reported.
Real-world applications (directions the framework's stated purpose points toward; the abstract reports no deployed examples or outcomes):
- Regulatory review. Regulators could require that a submission name the competency coordinates in which a device was validated, rather than accepting a general performance claim.
- Hospital procurement and deployment. Health systems could compare candidate tools against the specific coordinates in which they will actually be used — a particular care setting, provider role, and task.
- Human-AI role design. The three engagement dimensions (assigned authority, agent facing, anchoring layer) give teams explicit language for deciding how much authority an AI holds and how it interacts with a clinician's reasoning.
- Safety and scope-of-use documentation. Tool documentation could state where evidence exists and, by implication, where it does not.
Industry relevance. AI vendors, medical device manufacturers, health systems, and regulators all need a way to bound claims about clinical AI. A common grammar for specification and evaluation would affect product labeling, procurement criteria, and post-deployment monitoring, and could reduce the risk of a system validated in one context being relied upon in another.
Future Directions
- Empirical testing of the irreducibility claim. The abstract asserts that validation in one coordinate provides minimal evidence for another. Whether that holds quantitatively across real systems is an open question the framework itself raises.
- Operationalizing the eight dimensions. Turning condition, phase, care setting, provider role, task, assigned authority, agent facing, and anchoring layer into concrete, agreed-upon categories and measurement instruments.
- Mapping existing clinical AI tools into the competency space. Locating current systems at specific coordinates would reveal which regions of the space are well-evidenced and which are empty.
- Adoption by evaluators and regulators. Whether the framework becomes a common grammar depends on institutions agreeing to specify and bound competency claims in these terms, and on how it interacts with existing evaluation and regulatory structures.
Target Audience
Clinicians and clinical informatics leaders deciding how AI enters care; clinical AI and evaluation researchers designing benchmarks and studies; regulators and policymakers writing scope-of-use and validation requirements; human-factors and human-AI teaming researchers; and health system and industry decision-makers responsible for procuring, labeling, or deploying clinical AI. Readers looking for empirical results, benchmark comparisons, or implementation data will not find them in this abstract, which presents a framework and its structural implications.
Authors’ abstract
The competency of any intelligent agent is bounded by its formal account of the world in which it operates. Clinical AI lacks such an account. Existing frameworks address evaluation, regulation, or system design in isolation, without a shared model of the clinical world to connect them. We introduce the Clinical World Model, a framework that formalizes care as a tripartite interaction among Patient, Provider, and Ecosystem. To formalize how any agent, whether human or artificial, transforms information into clinical action, we develop parallel decision-making architectures for providers, patients, and AI agents, grounded in validated principles of clinical cognition. The Clinical AI Skill-Mix operationalizes competency through eight dimensions. Five define the clinical competency space (condition, phase, care setting, provider role, and task) and three specify how AI engages human reasoning (assigned authority, agent facing, and anchoring layer). The combinatorial product of these dimensions yields a space of billions of distinct competency coordinates. A central structural implication is that validation within one coordinate provides minimal evidence for performance in another, rendering the competency space irreducible. The framework supplies a common grammar through which clinical AI can be specified, evaluated, and bounded across stakeholders. By making this structure explicit, the Clinical World Model reframes the field's central question from whether AI works to in which competency coordinates reliability has been demonstrated, and for whom.