Research
Position: Machine Learning for Heart Transplant Allocation Policy Optimization Should Account for Incentives
Overview Research area: Machine learning policy optimization for healthcare resource allocation, specifically US adult heart transplant allocation; positioned at the intersection of machine learning,

- arXiv
- 2602.04990
- Published
- 2026-02-04
- Authors
- Ioannis Anagnostides, Itai Zilberstein, Zachary W. Sollie, Arman Kilic, Tuomas Sandholm
AI summary
Overview
Research area: Machine learning policy optimization for healthcare resource allocation, specifically US adult heart transplant allocation; positioned at the intersection of machine learning, mechanism design, strategic classification, causal inference, and social choice.
Technical level: Intermediate. The paper is a position paper written in largely accessible prose, but it assumes familiarity with concepts such as strategic classification, mechanism design, incentive compatibility, and risk minimization.
Scope (1 sentence): The paper argues that machine learning approaches to organ allocation policy optimization must be "incentive aware," and it documents specific incentive misalignments across the heart transplant decision-making pipeline with analyses of United Network for Organ Sharing (UNOS) data.
What This Paper Is About
Organ allocation in the US is transitioning from handcrafted, rule-based priority systems to machine learning and data-driven optimization, but the authors argue this framing misses a fundamental barrier: allocation is not a static optimization problem but a game involving organ procurement organizations (OPOs), transplant centers, clinicians, patients, and regulators, each responding strategically to policy changes. The paper's goal is to document concrete incentive misalignments in US adult heart transplant allocation, show with data that they have adverse consequences today, and lay out a research agenda for the machine learning community to design incentive-aware policies.
Key Contributions
-
A position statement that allocation policy optimization should be incentive aware. The authors argue that training predictive models on historical data without accounting for strategic shifts risks producing models that fail in practice, and that mechanism design, strategic classification, causal inference, and social choice should be integrated into allocation policy design.
-
A taxonomy of incentive misalignments across the decision-making pipeline, organized into five stages: patient feature manipulation (Section 2), out-of-sequence allocation (Section 3), strategic offer rejection under performance monitoring (Section 4), strategic listing and delisting (Section 5), and value aggregation and strategic preference reporting (Section 6).
-
Empirical analyses of UNOS heart transplant data (2010–2024) supporting the claims, including mortality and transplant timing metrics for highest-urgency candidates, center-level offer and acceptance distributions, evaluation-cycle timing effects, waitlist removal reasons, and multi-listing outcomes.
-
A set of concrete recommendations for each pipeline stage, including reliance on less manipulable features, randomized audits with penalties, hard constraints (rather than discretion) for triggering open offers, refined performance evaluation, credit score systems to limit offer rejections, objective listing and delisting criteria, and mechanism-design-guided preference elicitation focused on "ends" rather than "means."
Main Findings
-
The "device game" inflates urgency. The current US heart allocation rule divides patients into 6 tiers based on estimated medical urgency, deployed in 2018 to create a more granular separation relative to the previous 3-tier system. Because status depends on device utilization, clinicians can alter patient features to inflate medical severity: patients bridged with intra-aortic balloon counterpulsation (IABP) were designated status 2, while patients supported with durable left ventricular assist devices (LVADs) were given reduced priority to status 4. The proportion of patients bridged with an IABP increased from 7.0% to 24.9% following the policy update (a more than threefold increase), which the authors characterize as evidence the policy change triggered a shift in recipient selection for IABP bridging.
-
The highest-urgency system operates on thin margins. Table 1 reports, for highest-urgency (status 1A pre-2018, status 1 post-2018) US adult heart transplant candidates from 2010–2024: 6.5% of deaths occur within 3 days of listing and 13.7% within 7 days; median time to transplant is 26 days versus median time to death of 36 days (a margin of only 10 days); and the interquartile range of time to death is 13 to 118 days. The authors argue this variability shows a single status 1 classification conflates patients at immediate risk with more stable ones.
-
Out-of-sequence allocation has grown dramatically. Although out-of-sequence allocation was historically a rare exception (on the order of 1%), for kidneys the rate rose from 2% in 2020 to 18% in 2023, and across all organs approximately 19% of allocations in 2023 were classified as out of sequence. One cited investigation reports that out-of-sequence allocations predominantly favor certain demographic groups, typically more affluent individuals. After federal scrutiny of OPO out-of-sequence allocations starting in 2025, the rate reportedly plummeted from 20% in 2024 to 9% by early 2026.
-
Performance monitoring may distort center behavior. Transplant centers are evaluated semi-annually by the Scientific Registry of Transplant Recipients (SRTR) on a 5-tier system using three metrics: survival on the waiting list (pretransplant mortality rate), time to transplant (transplant rate), and 1-year graft survival. Data reporting windows close in April and October. Small centers receive only a fraction of the offers large centers receive and generally suffer higher waitlist mortality while accepting offers at a rate nearly 50% higher than large centers. While acceptance rates remain universally low (less than 1%), the authors observe a statistically significant spike in transplant volume and acceptance probability in May, immediately following the April deadline, which they note is consistent with horizon effects — though they state a more rigorous causal analysis is needed to attribute the spike definitively.
-
Waitlist removal is sizable and partly opaque. Among UNOS heart transplant candidates from 2010 to 2024, approximately 20% are removed from the waitlist for reasons other than transplant or death. More than a third of these candidates were removed because the center deemed them too sick, while nearly another third were removed under the generic category of "Other," which the authors highlight as a major limitation of current reporting.
-
Multi-listing confers a measurable advantage. Multi-listing is employed by only 2.16% of patients but multi-listed candidates achieve a higher transplant rate (80.44% versus 73.06%). The mean distance between centers for multi-listed candidates is 379 nautical miles (roughly the distance between Boston and Washington, D.C.), with a maximum exceeding 2,200 nautical miles. The authors report that as the geographic spread between centers increases, transplant rates rise, and that these patients often hold higher urgency status yet suffer lower waitlist mortality despite longer wait times.
-
Preference elicitation is vulnerable to strategic reporting. The continuous distribution framework computes a candidate's priority as a weighted sum over attributes, and the analytic hierarchy process (AHP) was used to set attribute weights through community input. The authors argue stakeholders' responses are likely skewed: small rural centers would benefit from broader sharing while urban centers would not, and candidates and families have incentives to inflate attributes that maximize their own match probability. A cited tentative survey produced a weight of 13.9% for "prior living donor" status, which the authors note seems like it should be zero for a given static patient pool. They also note AHP's limitations in high-dimensional feature spaces and its reliance on an "all else being equal" assumption that is often clinically unsound.
-
Counterarguments are acknowledged. Section 7 presents two alternative views: that clinician discretion acts as a necessary safety valve against imperfect risk metrics, and that mechanism design faces practical limits including the Gibbard-Satterthwaite theorem, impossibility results in kidney exchange about incentivizing centers to reveal all donors, unknown player incentives, and the risk that incentive compatibility becomes moot if agents cannot comprehend a mechanism.
Methodology in Plain English
The authors take a position-paper approach rather than proposing a single new algorithm. They first build a conceptual argument: allocation policy is a game, not just an optimization problem, and strategic behavior exists under rule-based policies and will persist under data-driven ones. They then walk through the allocation pipeline stage by stage, identifying who the strategic actors are, what they stand to gain, and how the current design creates vulnerabilities. To substantiate these claims, they analyze historical UNOS heart transplant data spanning 2010–2024, examining outcomes for highest-urgency candidates, the distribution of offers, transplants, and acceptance rates across centers, timing of transplant volume and acceptance relative to reporting deadlines, reported waitlist removal reasons, and transplant rates and geographic distances for multi-listed versus single-listed candidates. They pair these analyses with concepts from the machine learning and game theory literature — strategic classification, selective verification, repeated risk minimization, causal inference, mechanism design, and preference aggregation — and close with concrete recommendations for each pipeline stage.
Why This Matters
Impact on research: The paper argues that the next generation of allocation policies should be incentive aware, and that machine learning has a key role not only in optimizing policies but in mitigating the adverse effects of misaligned incentives. It identifies specific research gaps, including adapting strategic classification to survival analysis, using causal inference to identify features with a causal link to medical urgency, and what the authors describe as essentially entirely unexplored territory: optimizing national policy while accounting for the regional sub-policies that will be used. It also ties the work to public trust, noting that recent erosion of public confidence in the deceased-donor organ allocation system led to a significant drop in registered donors, and that in a system reliant on altruistic donors, a loss of trust is a loss of supply.
Real-world applications:
- Designing heart (and potentially other organ) allocation policies that are robust to feature manipulation, such as using randomized audits with penalties when device usage cannot be clinically justified, or building models that rely less on manipulable features like device utilization and accrued wait time.
- Replacing discretionary out-of-sequence allocation with hard, transparent triggers, potentially using machine learning to determine optimal thresholds and target centers or patients based on real-time indicators of donor organ state.
- Refining center performance evaluation, for example through more accurate risk adjustment models or a shift from rigid semi-annual evaluation to continuous monitoring.
- Standardizing risk models for waitlist admission and removal, and using counterfactual modeling to quantify when multi-listing benefits the system versus when it undermines equity.
Industry relevance: The paper is relevant to organ procurement organizations, transplant centers and their clinical leadership, regulators such as those overseeing SRTR evaluations and OPTN policy development, and health technology developers building decision-support or policy-optimization systems. The authors note that all algorithmic approaches mentioned have already been deployed in practice in some form, and that misaligned incentives have operational consequences today — including offer rejections that trigger a cycle of organ quality degradation and discard.
Future Directions
- Adapt strategic classification to survival analysis. The authors state that adapting repeated risk minimization and strategic classification frameworks to survival analysis presents distinct research challenges, and call for causal inference to identify features causally linked to medical urgency.
- Build machine learning components for out-of-sequence decisions. This includes determining optimal thresholds to trigger an out-of-sequence allocation, deciding which center or patient should receive the organ based on real-time organ viability indicators (potentially using computer vision during ex-vivo perfusion), and coupling this with dynamic evaluation of recipients in the local region — a problem the authors say is essentially entirely unexplored.
- Design better monitoring and acceptance incentive systems. Options raised include more accurate risk adjustment models, continuous monitoring instead of semi-annual evaluation, credit score systems to incentivize greater acceptance of offers, and — as a more speculative alternative facing considerable obstacles — removing the option for centers to reject offers.
- Apply mechanism design and frugal preference elicitation to value aggregation. The paper suggests frugal preference elicitation from multiple parties, techniques from query learning, and reinforcement learning from human feedback, while emphasizing that preference elicitation should focus on the "ends" (normative objectives) rather than the "means" (attributes and weights), which should be optimized computationally.
- Evaluate before deployment. The authors stress the need for extensive evaluation prior to deployment, citing kidney exchange as an example where this process proved essential for building clinical trust.
Target Audience
Machine learning and mechanism design researchers working on healthcare, allocation, or social-choice applications; transplant clinicians, surgeons, and center administrators; organ procurement organization leadership and policy staff; regulators and policymakers involved in allocation rulemaking and performance evaluation; and health economists or operations researchers studying strategic behavior in resource allocation. The paper is also useful for graduate students seeking an example of how incentive analysis and empirical data analysis can be combined to critique an existing deployed policy.
Authors’ abstract
The allocation of scarce donor organs constitutes one of the most consequential algorithmic challenges in healthcare. While the field is rapidly transitioning from rigid, rule-based systems to machine learning and data-driven optimization, we argue that current approaches often overlook a fundamental barrier: incentives. In this position paper, we highlight that organ allocation is not merely an optimization problem, but rather a complex game involving organ procurement organizations, transplant centers, clinicians, patients, and regulators. Focusing on US adult heart transplant allocation, we identify critical incentive misalignments across the decision-making pipeline, and present data showing that they are having adverse consequences today. Our main position is that the next generation of allocation policies should be incentive aware. We outline a research agenda for the machine learning community, calling for the integration of mechanism design, strategic classification, causal inference, and social choice to ensure robustness, efficiency, fairness, and trust in the face of strategic behavior from the various constituent groups.