Research
Rare Diseases, Common Dilemmas: LLMs Prioritize Equal Resource Distribution over Patient Benefit in Decision-Making
Overview Research area: AI safety and ethics, with a focus on how large language models resolve value conflicts in clinical decision support for rare disease care. Technical level: Intermediate. The p
- arXiv
- 2608.25236
- Published
- 2026-08-25
- Authors
- Minda Zhao, Xu Han, Rishabh Goel, Maya Dagan, Noa Dagan, Adithya Madduri, Payal Chandak, Shilpa Nadimpalli Kobren, Isaac S. Kohane
AI summary
Overview
- Research area: AI safety and ethics, with a focus on how large language models resolve value conflicts in clinical decision support for rare disease care.
- Technical level: Intermediate. The paper assumes some familiarity with the four-principle bioethics framework (autonomy, beneficence, nonmaleficence, justice) and with LLM benchmarking, but its core claims are stated in plain ethical and clinical terms.
- Scope: A benchmark of 208 clinically grounded rare disease vignettes and 11 large language models, used to characterize which ethical values models prioritize, how those priorities shift with the framing of decision-making authority, and how little model identity matters relative to vignette framing.
What This Paper Is About
Clinical decisions frequently involve more than one defensible option, so the real question becomes which ethical value to prioritize, not which answer is factually correct. Rare disease care is an extreme case of this, since scarce evidence, constrained trial slots, high costs, and high-stakes outcomes make ethical tension routine rather than exceptional. The authors ask what ethical values current LLMs implicitly favor when forced to choose between two clinically defensible but ethically irreconcilable next steps in such cases.
Key Contributions
- A clinically grounded benchmark of ethical value alignment in rare disease care. The authors operationalize 208 distinct vignettes derived from authoritative Orphanet and OMIM data, covering real disease entities, symptom profiles, ages of onset, and treatment constraints, and evaluate 11 state-of-the-art LLMs on them. The framing shifts evaluation away from factual correctness toward empirical characterization of latent value prioritization.
- A demonstration of cross-model homogenization toward Justice. Despite differences in architecture, training corpora, and alignment procedures, all evaluated models ranked Justice above Autonomy, Beneficence, and Nonmaleficence. The paper further shows this preference is driven largely by equal resource distribution rather than equitable medical need.
- Identification of a pervasive authority-framing effect. Model reasoning depends on who is specified as the decision-maker, not only on the ethical substance of the case. Models default to Justice in committee-based contexts and shift toward Autonomy and Beneficence when the decision is framed as belonging to a clinician or a patient.
Main Findings
- Justice dominates across every model. Across all 11 evaluated models, Justice achieved the highest normalized win rate, ranging from 57% to 70%. The pattern held in both proprietary and open-weight systems. DeepSeek-V3 and GPT-OSS-20B reached 70%; GPT-4o, GPT-4.1, and Claude 4 were near 69%; Qwen3-14B showed the weakest Justice preference at 57%. The magnitude varied, but the top ordinal rank did not.
- Justice selections are concentrated in one conception of fairness. When Justice-labeled choices were decomposed into need, equity, equality, and maximizing overall benefit, equality (equal resource distribution) formed the largest component of Justice-oriented selections, while need- and equity-based Justice appeared less frequently. The authors highlight that in rare disease care, equal distribution and need-sensitive allocation imply different recommendations when severity and medical necessity are unequally distributed.
- Secondary value orderings differ by model. Once Justice is removed from comparison, Autonomy is consistently preferred over Beneficence and generally favored over Nonmaleficence, though the latter contrast is less decisive. Beneficence versus Nonmaleficence conflicts produce substantial model-level heterogeneity. Nonmaleficence had the widest overall range, from 31% in Qwen3-14B to 58% in DeepSeek-V3.
- Decision-maker framing is the dominant contextual factor. The decision-maker factor showed a large association with selected ethical value (Cramér's V = 0.504, p < 0.001), far above the paper's threshold of V ≥ 0.40 for a large effect. Vignettes specified a Committee, a Medical Team, or an Individual decision-maker.
- Patient context matters, but less. Patient Type, encoded as a decisional-role variable (Maternal-Fetal, Proxy, Self-Directed), showed a medium association (V = 0.206, p < 0.001). Patient Age showed a smaller but significant association (V = 0.181, p < 0.001). The paper describes these results as revising an earlier interpretation that patient factors were negligible.
- Model identity is a negligible driver. Model identity showed a negligible and non-significant association with value selection (V = 0.067, p = 0.416), reinforcing the convergence seen in the win-rate analysis. The same vignette-level framing cues were more strongly associated with choices than which LLM was queried.
- Autonomy and Beneficence are strongly framed by authority. Relative to Committee framing, Individual decision-makers were associated with a 5.71-fold increase in the odds of selecting Autonomy (95% CI: 3.41–9.56, p < 0.001), the largest effect reported, and Medical Teams with a 3.72-fold increase (95% CI: 2.11–6.59, p < 0.001). Beneficence selection increased 4.23-fold for Individual contexts (95% CI: 1.57–11.40, p = 0.004) and 3.46-fold for Medical Team contexts (95% CI: 1.25–9.59, p = 0.017).
- Nonmaleficence shows invariance. Neither Individual (OR = 1.26, 95% CI: 0.57–2.81, p = 0.57) nor Medical Team (OR = 1.83, 95% CI: 0.77–4.38, p = 0.17) contexts significantly altered Nonmaleficence selection.
- Justice regression results are flagged as unreliable. The paper reports that Justice-related results from the logistic regression are unreliable due to severe sample imbalance. The available text is truncated at this point, so the remaining results are not reported here.
- One value pair could not be tested. Justice versus Nonmaleficence is absent from the pairwise analysis because those vignettes failed the quality control loop, value assignment, and/or human clinical review due to inappropriate value tagging, clinical infeasibility, or not representing a true ethical dilemma.
Methodology in Plain English
The pipeline has five phases.
- Data curation. The authors built a disease-level knowledge table from Orphanet (Orphadata) and OMIM. Starting from Orphanet XML sources covering roughly 4,000+ diseases and requiring complete data across gene association, phenotype descriptors, and age-of-onset category, they retained 2,444 diseases. They then harmonized gene symbols and mapped inheritance patterns using OMIM's genemap2. A second quality-control filter removed 181 diseases (7.4%) with unresolved or "unknown" inheritance, yielding 2,263 curated rare diseases.
- Vignette generation. For each candidate vignette, they sampled a disease entity and assembled four seeds: disease name and associated gene, age-of-onset category, a prevalence-ordered symptom profile, and a pre-specified pair of bioethical values to place in conflict. Drafts were generated with GPT-4.1 under a constrained prompt requiring a 4–5 sentence clinically plausible scenario and a forced choice between two actions that were clinically defensible yet ethically irreconcilable. Vignettes were barred from naming the ethical values in the text shown to models.
- Quality control and human review. Drafts passed a diversity gate for redundancy, then two simulated expert agents, a clinician agent and a bioethicist agent, either accepted or sent them back for regeneration. Retained candidates were scored by an LLM-based judge on a 0–10 scale across clinical realism, ethical balance, clarity of the forced choice, and irreconcilability of the trade-off; only candidates scoring at or above 7.0 proceeded. Value alignment was then checked with an existing multi-step framework that assigns promoting or opposing scores across four ethical values, and six reviewers with medical or biomedical backgrounds performed a feasibility audit, with a representative subset assigned to two independent reviewers each. The result was 208 validated vignettes.
- Model evaluation. Eleven LLMs were chosen for provider diversity, access modality, and capability tier: GPT-4o, GPT-4.1, GPT-5, and GPT-OSS-20B from OpenAI; Claude-4-Sonnet from Anthropic; Gemini-2.5-Flash and Gemma-3-27B from Google; LLaMA-3.3-70B from Meta; Qwen3-14B from Alibaba; Mistral-Small-3.2 from Mistral; and DeepSeek-V3 from DeepSeek. Five are closed-source and six open-source. Each model received only the scenario narrative ending in Choice A and Choice B, with a minimal prompt instructing it to respond with a single letter and no explanation. Responses that did not specify a valid choice were re-queried once and then recorded as missing if still noncompliant. Temperature was set to 0.7.
- Analysis. Each vignette was also annotated with three contextual factors extracted mechanically from the prose: decision maker (15 raw labels collapsed to Committee, Medical Team, Individual; 59.6%, 17.8%, 22.6%), patient type (10 labels collapsed to Maternal-Fetal, Proxy, Self-Directed; 25.2%, 63.1%, 11.7%), and patient age (9 labels collapsed to Infant, Pediatric, Adult; 37.1%, 22.8%, 40.1%). Missing values were not imputed. The primary outcome was the ethical value mapped to each chosen action. The authors computed normalized win rates to control for unequal value frequencies, used Cramér's V to screen associations between categorical factors and value selection, and fit binary logistic regressions with Committee as the reference category to estimate decision-maker effects. The logistic regressions were restricted to the closed-source GPT subgroup because it had the lowest inter-model heterogeneity (Cramér's V = 0.0305, p = 0.979). The authors state explicitly that these regressions characterize directional authority-framing patterns rather than normative correctness.
Why This Matters
Impact on research. The paper argues that the dominant evaluation paradigm for medical LLMs is fundamentally epistemic, testing whether answers are correct, harmful, or biased on examinations, medical QA benchmarks, and safety rubrics, while ignoring which defensible ethical trade-offs those answers embody. By treating models as analytical instruments for studying ethical decision-making rather than as clinical agents to be judged against normative standards, it provides a reproducible framework for measuring value prioritization, and shows that the four principles can be decomposed further, as with the four allocation logics underlying Justice.
Real-world applications.
- Clinical decision support in rare disease. If a deployed system silently prefers equal resource distribution over medical need, it could underweight severity and urgency in recommendations about who receives scarce interventions.
- Prior authorization and coverage review. Payers applying evidentiary, reimbursement, and budget-impact standards to rare disease cases would be exposed to a model whose justice reasoning defaults to equality rather than need.
- Shared decision-making tools for patients and caregivers. Rare disease families often hold factual authority their fragmented care teams lack; a system that treats patient input as generic preference expression rather than knowledge-bearing could misrepresent that role.
- Trial slot and resource allocation in specialist centers. Committee-framed decisions were the context in which models defaulted most strongly to Justice, which is precisely the setting where opportunity costs across patients and programs are weighed.
Industry relevance. The paper suggests that institutional pressures surrounding rare disease resource utilization may be silently reflected in LLM-based decision support, and that model outputs could reinforce existing institutional asymmetries rather than supply principled ethical guidance. Because model identity contributed negligibly to value selection while authority framing contributed strongly, the salient design lever for developers is prompt and deployment context, not which model is chosen. The fact that all models, proprietary and open-weight alike, converged on the same top-level preference indicates that this is not a single-vendor problem.
Future Directions
- Resolve the Justice analysis. The logistic regression for Justice is described as unreliable due to severe sample imbalance, and the pairwise Justice versus Nonmaleficence contrast was never tested because those vignettes failed quality control. Both leave the Justice findings incomplete and point to a need for better-balanced vignette sampling or dedicated generation for that value pair.
- Test whether need-sensitive reasoning can be elicited. Since equality dominated Justice selections at the apparent expense of need and equity, an open question is whether prompting, decomposition, or explicit rubric provision can shift models toward need- or equity-based allocation rather than equal distribution.
- Examine whether the authority-framing effect mirrors real institutional behavior. The monotonic gradient from Committee to Medical Team to Individual suggests models have learned to associate concentrated authority with individual-focused values. Whether that learned association tracks actual clinical practice, or reflects patterns in training text, is not established by this work.
- Validate against clinician and patient judgments. The benchmark measures revealed model choices, not whether those choices agree with what clinicians, ethics committees, or rare disease families would decide. The paper does not report such a comparison.
Target Audience
This paper is most useful to AI ethics and AI safety researchers working on value alignment and evaluation methodology; to clinical informatics and health AI developers building or auditing decision support systems; to bioethicists interested in how computational systems operationalize principles such as justice and autonomy; to rare disease clinicians, advocacy organizations, and patient communities affected by resource allocation decisions; and to regulators and health system leaders assessing the ethical behavior of medical LLMs beyond factual accuracy benchmarks.
Authors’ abstract
Clinical decision-making often involves prioritizing ethical values, such as beneficence, non-maleficence, respecting a patient's autonomy, and justice. Recent work has begun to assess how large language models (LLMs) make such subjective, value-laden clinical judgments. However, evaluations of LLM decision-making in rare disease care contexts, where ethical tensions are ubiquitous and where scarce prior information likely impacts LLM behavior, are still lacking. Here, we present a benchmark of 208 clinically grounded rare disease vignettes, each of which presents genuine, high-stakes conflicts. When prompting 11 state-of-the-art LLMs to choose between clinically defensible yet ethically conflicting next steps embedded within these vignettes, we found that all evaluated models consistently prioritized justice over other core bioethical principles. Specifically, models overwhelmingly favor equal resource allocation over need-based considerations, indicating LLMs' limited responsiveness to differences in clinical severity or situational context. We also identify a strong authority-framing effect: models favor justice in committee-based contexts and shift toward beneficence and autonomy only when final decisions are framed as being made by clinicians or patients respectively. Our work suggests that institutional pressures surrounding rare disease resource utilization may be silently reflected in LLM-based decision support systems, with finer ethical considerations disregarded.