Research
From PhysioNet to Foundation Models -- A history and potential futures
From PhysioNet to Foundation Models — A history and potential futures Overview Research area: Open-access biomedical data sharing, physiological signal databases, data-science competitions, and the in

- arXiv
- 2602.15371
- Published
- 2026-02-17
- Authors
- Gari D. Clifford
AI summary
From PhysioNet to Foundation Models — A history and potential futuresOverview
- Research area: Open-access biomedical data sharing, physiological signal databases, data-science competitions, and the intersection of open science with modern machine learning and AI in medicine (with a focus on cardiology).
- Technical level: Beginner-Friendly. The paper is a historical review and perspective article rather than a technical methods paper; it contains no new algorithms, mathematical formulations, or benchmark tables of model results.
- Scope: A single-author account that traces roughly three and a half decades of medical data sharing from mailed tapes and CD-ROMs to networked databases and foundation models, using PhysioNet as the central case study, and lays out the author's view of the field's most promising future directions.
What This Paper Is About
Medical data sharing has moved from mailing magnetic tapes and compact discs containing a handful of carefully labeled recordings to high-speed download of large hospital databases, and now to the push to pour nearly any data source into large models. The author argues this shift brings great potential alongside a new set of problems, including labeling quality, data bias, competition design, prize incentives, funding models, repeatability, and the carbon cost of large AI models. The goal is to document PhysioNet's history and contributions, compare it with other open biomedical data and competition initiatives, and propose where the resource and the broader field should go next.
Key Contributions
-
A firsthand mini-history of PhysioNet. The author, who has supported the resource for a quarter of a century and led its Challenges since 2015, documents the lineage from the MIT-BIH Arrhythmia Database (assembled 1975–1979, released 1980) through the 1999 founding of the NIH-funded Research Resource for Complex Physiologic Signals, including the original three cores, key personnel, and milestones such as distribution on 9-track tape, CD-ROM, and eventually the internet.
-
A structured comparison of biomedical and general data-science competitions. The paper contrasts the PhysioNet Challenges with the KDD Cup (began 1997), the Sage Bionetworks DREAM Challenges (DREAM founded 2006, Sage founded 2009, more than 60 challenges together), and Kaggle (founded 2010, 28.5 million registered users by the end of 2025, acquired by Google in 2017), with particular attention to rules on entries, code release, retrainability, and hidden test data.
-
An analysis of prizes and incentives. The paper traces competition history from ancient Greek contests through the 1714 Longitude Prize, examines evidence on prize size and participation (including a study finding a total prize of between $1,700 and $2,000 optimal), and uses the Ansari XPRIZE and Qualcomm Tricorder XPRIZE as contrasting cases.
-
A forward-looking agenda. The author sets out what he considers the most promising directions for PhysioNet and for massive physiological databases generally — covering foundation models and their AI carbon footprint, Tiny-ML and edge computing, open-access code, public competitions, funding models, and scientific repeatability — and proposes addressing key issues through the PhysioNet Challenges.
Main Findings
-
Open medical data predates the open-access movement: The MIT-BIH Arrhythmia Database, released in 1980, is described as the first open-access health database, containing 48 30-minute Holter monitor recordings chosen for intermittent but clinically significant rhythms, QRS morphology, or noise and artifacts. Its 600 MB of 109 thousand beats were originally distributed to approximately 100 sites on digital 9-track tape.
-
Distribution scaled stepwise: In 1989 the database became available on CD-ROM alongside seven (later nine) additional ECG databases, and by the end of the 1990s approximately 400 copies of these CD-ROMs had been mailed worldwide. Internet access later reduced acquisition time to minutes for those on an academic network.
-
PhysioNet's data volume has grown enormously: As of June 2025, PhysioNet hosts more than 15 TB of data in 215 databases, plus much more data in over 150 databases requiring credentialed access. The paper presents a log-linear plot indicating super-exponential growth in public ECG databases.
-
Labeling has become the bottleneck: As databases grew, meticulously labeling every beat became impractical. The latest database of over 10 million ECGs makes expert labeling cost-prohibitive, forcing reliance on a single expert or algorithm-generated labels, which is likely to encode the biases and limitations of each method and constrain performance.
-
The PhysioNet Challenges have distinctive rules: These 8-to-9-month events typically draw approximately 100 teams averaging four members each, who collectively devote tens of thousands of hours. Teams must submit at least one entry (at most five) during a 4–8 week beta-testing phase, are limited to 10 entries on public data, must submit both the trained model and the training harness, must provide a four-page scientific article reviewed for accuracy, and must publicly defend their approach at a meeting. Training and test time are limited to ensure equitable competition and to value innovation over raw compute.
-
A fully sequestered test set mimics deployment: Submissions are evaluated on an independent test set never released to participants, and neither validation nor test data are published after a Challenge concludes, to prevent publications from overfitting to them and claiming excessive performance. Teams may get a second chance at the test data post-Challenge if they send a draft journal publication, its intended venue, and a description of how their code was modified.
-
The Challenges have produced new datasets: Examples cited include handheld ECGs for identifying atrial fibrillation, multimodal waveform databases for predicting sepsis or coma outcome, and joint ECG image-signal databases for digitizing paper ECGs. In 2025, a Challenge introduced a metric reflecting the downstream capacity of the regional health system in Brazil to perform a definitive blood test.
-
Prize-size effects are non-obvious: One cited study finds an inverted U-shaped relationship, where moderate awards maximize participant numbers because large awards raise perceived competition and lower a solver's expectancy of winning; it identified a total prize of between $1,700 and $2,000 as optimal, though that analysis covered no medical-domain competitions and was limited to total prizes of $4,000 or less. Three claimed advantages of large participant pools are solution diversity, higher quantity and quality of creative solutions, and faster market speed.
-
Large prizes have mixed records: The 1996 Ansari XPRIZE offered and in 2004 awarded $10 million for a reusable three-person spacecraft reaching 100 km twice within two weeks; teams invested over $100 million, and the prize spurred $1.7 billion in total private investment, which the author characterizes as a 10:1 or 170:1 return depending on how output is counted. It also led to the formation of Virgin Galactic and inspired SpaceX. By contrast, in the 2012 Qualcomm Tricorder XPRIZE no team met all the $10 million grand prize requirements; a reduced total of $3.6 million was split between top teams, and the high-profile Scanadu Scout device that emerged failed to receive FDA approval and was discontinued. The author attributes part of the disparity to the far greater regulation of health compared with space.
-
PhysioNet prizes have been deliberately understated: The author states that the Challenges have consistently provided cash prizes in multiple categories, from several hundred to several thousands of dollars, but that the amounts were never publicly divulged before this publication — which led generative AI to incorrectly quote the Challenges as offering no financial prizes.
-
Government prize use has grown, but evidence of impact is limited: A 2020 US Congress note found federal prize money rose from $247,000 in fiscal year 2011 to over $37 million in fiscal year 2018, with the median prize rising from $34,500 to $80,000 over the same period, following the America COMPETES Reauthorization Act of 2010. The same report stated there is limited information on the effectiveness of federal prize competitions in spurring innovation.
-
Kaggle and PhysioNet differ substantially in data governance and entry limits: PhysioNet employs rigorous review before accepting data, requires contributors to have generated the data themselves and to provide institutional review board approval, and accompanies datasets with peer-reviewed publications and usage code. Kaggle allows users substantial discretion over format, quality, licensing, and provenance with relatively limited oversight, and permits at least 2 attempts per day (typically 5–10, up to 100 in some competitions), which the author notes risks overtraining on validation data. Kaggle's terms of service protect users' code and intellectual property and do not grant a license to use it beyond running competitions and analytics.
-
A new partnership is under way: The author reports a current partnership with Kaggle, motivated by reaching a wider and different audience, gaining experience with another platform's challenge design, and enhancing prizes, with results to be reported in a future publication.
Methodology in Plain English
This is a narrative review and perspective piece rather than an experimental study. The author draws on almost three decades of firsthand experience co-developing elements of PhysioNet with its founders, including seven years at MIT and Harvard, leading the George B. Moody PhysioNet Challenges since 2015, and work in global health over the same period. The article began as an invited talk at Computing in Cardiology 2024 and was expanded into a longer journal article that functions both as a mini history of the resource and as a landscape assessment of where big data, foundation models, and open science should go next.
The approach combines personal recollection, historical milestones and dates, tabulated comparisons of databases (number of individuals, average recording length, number of leads, geographic origin of subjects) and of data-science competitions, citation of published work on prizes and incentives, and case-study reasoning from large-prize competitions. No new data collection, model training, or quantitative benchmark experiment is described. Funding and sponsorship are disclosed, and the author declares no conflict of interest. The provided content is truncated while discussing the difficulty of measuring the output of Challenges, so the remaining sections are not available in the text summarized here.
Why This Matters
The paper matters because it documents how the infrastructure, norms, and incentive structures of medical data sharing were built, and argues that the current pace of AI adoption in medicine makes careful guide rails necessary. The author's framing — inverting the old Facebook motto to "when things move fast, things break" — captures the central concern: data-driven clinical decision support is promising, but the speed of change and the rush to pump data into large models create risks around bias, repeatability, and cost that the field has not yet solved.
Real-world applications cited or implied in the paper:
- Cardiac device regulation: The MIT-BIH Arrhythmia Database became a standard used for FDA approval of arrhythmia analysis systems, making open data part of the regulatory pathway for medical devices.
- Atrial fibrillation detection: Challenge work produced handheld ECG databases for identifying atrial fibrillation, relevant to consumer and point-of-care monitoring.
- Critical care prediction: Multimodal waveform databases support predicting sepsis or coma outcome.
- Digitizing paper records: Joint ECG image-signal databases enable converting paper ECGs into usable digital signals.
- Health-system-aware metrics: A 2025 Challenge metric reflected the capacity of Brazil's regional health system to perform a definitive blood test, tying algorithm evaluation to real clinical logistics.
Industry relevance: The paper connects directly to medical device developers, health technology companies, and cloud and hardware vendors, several of which sponsor the Challenges (Alivecor, Amazon Web Services, Google, the Gordon and Betty Moore Foundation, the IEEE Signal Processing Society, and Mathworks). It also speaks to data platform operators, since the author compares PhysioNet's review and entry-limit policies with Kaggle's, and to funders, given the discussion of return on investment in prize-based funding. The paper was supported in part by the National Institute of Biomedical Imaging and Bioengineering and the Director's Office of Data Science Strategy of the NIH under Award Number R01EB030362, and the Challenges received the "Distinguished Achievement Award for Data Reuse" as part of the inaugural DataWorks! prize in 2022 from the Federation of American Societies for Experimental Biology.
Future Directions
- Foundation models versus resource constraints: The paper asks how the field should approach foundation models in the context of the rapidly growing AI carbon footprint, weighing large-scale models against their environmental and resource costs.
- Tiny-ML and edge computing: The author identifies Tiny-ML and edge computing as promising directions, consistent with the Challenges' philosophy of valuing innovation over raw compute and limiting training and test time.
- Prizes, incentives, and funding models: Open questions include exactly how large a prize should be to maximize return on investment in a given biomedical field, since the dose-response appears field- and competition-specific and has yet to be determined; and how competitions can serve as a scalable alternative to conventional funding models.
- Repeatability and measurement of impact: The paper raises scientific repeatability as a concern and notes that measuring the output of Challenges is extremely difficult — the provided content is truncated at exactly this point, and the author's proposed solutions to the field's key issues, including how they might be addressed through the PhysioNet Challenges, are not visible in the text summarized here. A previously unpublicized partnership with Kaggle is also slated for a future publication.
Target Audience
This paper is most useful to biomedical data scientists, clinical researchers, and machine learning practitioners working with physiological signals; competition designers and platform operators who want to understand alternative rules around hidden test sets, entry limits, and retrainable code; research funders and policymakers interested in prize-based incentives and their measured impact; and historians or students of open science who want a firsthand account of how one of the oldest open biomedical data resources was built and sustained. Medical device developers and health technology teams relying on PhysioNet databases for benchmarking will also find the descriptions of dataset provenance, labeling limitations, and governance requirements directly relevant.
Authors’ abstract
Over the last 35 years, the sharing of medical data and models for research has evolved from sneakernet to the internet - from mailing magnetic tapes and compact discs of a handful of well-curated recordings, to the high-speed download of relatively comprehensive hospital databases. More recently, the fervor around the potential for modern machine learning and 'AI' to catapult us into the next industrial revolution has led to a seemingly insatiable desire to pump almost any source of data into large models. Although this has great potential, it also presents a whole set of new challenges. In this article I examine these trends over the last 30 years, drawing on examples from cardiology, one of the oldest data-intensive fields that is undergoing a renaissance via machine learning. From the early days of computerized cardiology, the Research Resource for Complex Physiologic Signals (PhysioNet) has been at the cutting edge of this field. This article, therefore, includes much of the Resource's history and the contributions drawn from 25 years of firsthand experience of co-developing elements of the Resource with its founders. I identify the most promising future directions for the PhysioNet Resource, and more generally, the growing issues and opportunities around dissemination and use of massive physiological databases, associated open access code, and public competitions, along with potential solutions to the key issues facing our field. Topics range from how we should approach foundation models in the context of the rapidly growing AI carbon footprint, to the potential of Tiny-ML and edge computing. I also cover issues around prizes and incentives, funding models, and scientific repeatability, as well as how we might address these issues by leveraging the PhysioNet Challenges, consistent with the philosophy of open-access from the early days of the PhysioNet Resource.