Research
Medical Imaging AI Competitions Lack Fairness
Overview Research area: Medical imaging AI benchmarking, with a focus on dataset representativeness, data accessibility, licensing, and documentation (FAIR principles). Technical level: Intermediate —

- arXiv
- 2512.17581
- Published
- 2025-12-19
- Authors
- Annika Reinke, Evangelia Christodoulou, Sthuthi Sadananda, A. Emre Kavur, Khrystyna Faryna, Daan Schouten, Bennett A. Landman, Carole Sudre, Olivier Colliot, Nick Heller, Sophie Loizillon, Martin Maška, Maëlys Solal, Arya Yazdan-Panah, Vilma Bozgo, Ömer Sümer, Siem de Jong, Sophie Fischer, Michal Kozubek, Tim Rädsch, Nadim Hammoud, Fruzsina Molnár-Gábor, Steven Hicks, Michael A. Riegler, Anindo Saha, Vajira Thambawita, Pal Halvorsen, Amelia Jiménez-Sánchez, Qingyang Yang, Veronika Cheplygina, Sabrina Bottazzi, Alexander Seitel, Spyridon Bakas, Alexandros Karargyris, Kiran Vaidhya Venkadesh, Bram van Ginneken, Lena Maier-Hein
AI summary
Overview
Research area: Medical imaging AI benchmarking, with a focus on dataset representativeness, data accessibility, licensing, and documentation (FAIR principles).
Technical level: Intermediate — readers should be familiar with the general idea of benchmarking challenges and medical imaging modalities, but the analysis itself is about data practices rather than algorithms.
Scope: A large-scale systematic audit of 249 biomedical image analysis challenges (458 tasks, 19 imaging modalities, 2018–2023) assessing whether challenge datasets fairly represent clinical reality and whether they can be legally and practically reused.
What This Paper Is About
Biomedical image analysis competitions have become the main way the field decides what "state of the art" means, and their datasets are reused widely as reference benchmarks. The authors ask whether these competitions are actually fair: do their datasets reflect real-world clinical diversity, and can the data be found, accessed, and legally reused? They conduct a systematic, multi-observer review of 249 challenges to quantify representation, access conditions, licensing, and reporting quality.
Key Contributions
- Large-scale systematic analysis of biomedical imaging challenges: a review of recent challenges spanning 458 medical imaging AI tasks across 19 modalities.
- Empirical evidence of limited representation: quantification of overrepresentation of certain geographic regions, imaging modalities, problem types, and anatomical regions relative to external reference distributions.
- Critical assessment of accessibility, licensing, and documentation: documentation of restrictive access conditions, unclear or inconsistent licensing, and reporting gaps across challenge tasks.
- Framing of fairness along two dimensions: representativeness (whether datasets reflect real-world conditions) and reuse (accessibility, licensing clarity, and documentation, aligned with the FAIR principles with a focus on Accessibility and Reusability).
Main Findings
- Challenge landscape: 249 challenges comprising 458 tasks were identified between 2018 and 2023. Challenges had on average 1.7 tasks (median 1, IQR 1, maximum 10). Most were newly introduced (71%), and only 6% had reached their fifth or higher edition. The majority were hosted at MICCAI (66%), followed by ISBI (14%).
- Venue-linked dataset scale: median training/test dataset sizes were 276/100 cases at MICCAI, 400/113 at ISBI, 2320/1311 at NeurIPS, and 1251/570 at RSNA.
- Overall dataset sizes: across all entries, training/validation/test minimum was 0/1/1, median 288/80/102, and maximum 530,706/18,368/47,227. The median training-to-test ratio was 2.2:1.
- Participation: the median number of teams in the final phase was 12 (maximum 3,308), and the median ratio of participating teams between the test and final phases was 2.3:1. Kaggle had the highest participation (mean 1,289; median 1,347), while MICCAI, NeurIPS, ISBI, and MIDL had medians below 20 teams.
- Submission type shift: 49% of challenges required algorithm or code submission and 46% required result submission. Result submission dominated until 2021 (more than 70% of tasks yearly), but by 2023, 75% of tasks required algorithm or code submission.
- Annotation practices: the median number of annotators was 3 (maximum 92) and the median number of annotators per case was 2 (maximum 7). Experts were involved in 47% of tasks, intermediate-experience annotators in 26%, and trainees including students in 19%. Expert supervision or final review was reported for 31% of tasks, and only 1% used a professional annotation company. Annotator expertise was unclear or not reported for 27% of tasks.
- Geographic concentration: 70% of tasks used data originating from the United States, followed by China (16%), Germany (13%), the Netherlands (10%), and the United Kingdom (10%). Oceania (2.4%), South America (1.3%), and Africa (1.0%) were rarely represented. Relative to global population distribution, North America was overrepresented by a factor of 15.9 and Europe by 7.9, while Africa (0.1), South America (0.2), and Asia (0.5) were underrepresented. Median dataset size varied from 17/10 training/test cases in Japan to 22,601/6,223 in Italy.
- Problem-type concentration: segmentation accounted for 39% of tasks, classification 21%, and detection 13%. Median dataset sizes ranged from 30/35 training/test cases for registration to 1,251/209 for synthesis; segmentation tasks had medians of 175 and 72 cases, classification 450/200, and detection 420/153.
- Modality concentration: MRI (35%) and CT (21%) were most frequent, followed by laparoscopy or endoscopy (13%) and histopathology (11%). Compared with US imaging utilization rates, MRI was 4.4-fold overrepresented and ultrasound severely underrepresented (0.16-fold). MRI imaging rates for older adults were 139 per 1,000 person-years versus 428 for CT and 495 for ultrasound in the US in 2016. Median dataset size ranged from 140/57 training/test cases for MRI to 1,905/177 for laparoscopy- or endoscopy-related tasks.
- Task–modality coupling: detection tasks were most frequently based on laparoscopy or endoscopy data (31%), whereas classification tasks were distributed across modalities with no single modality exceeding 21%.
- Anatomical focus: tasks most frequently concerned the brain (30%), followed by the eyes (10%), colon (10%), and lung (9%).
- Data sourcing and demographics: data were typically collected from a median of one center (mean 6, maximum 500) and a median of two devices (mean 2.5, maximum 18). Where reported, the median sex distribution was 52% male and 47% female. Only 26% of tasks reported age information; among those, median reported mean age was 54 years (mean 49.9, IQR 13), with values ranging from 0 to 100 years.
- Access barriers despite nominal openness: 81% of datasets were publicly available, but access was restricted by mandatory registration (38%), organizer approval (20%), or context limitations (9%) such as country restrictions or time-limited access.
- Licensing quality: on the (Re)usable Data Project 0–5 scale (5 = license clearly permits unrestricted reuse and redistribution), the median score across tasks was 3; 46% of tasks scored 2.5 or lower. Only 59% of tasks provided sufficiently detailed license information and only 41% used an unambiguous license. CC licenses were applied in 51% of cases, often modified or combined with a custom license (19%). The most common licenses were custom licenses (29%), CC BY (18%), CC BY-NC-SA (18%), CC BY-NC-ND (15%), and CC BY-NC (11%). Restrictive licenses dominated (44%) versus permissive licenses (25%). Only 40% of peer-reviewed challenges applied the same license approved by the conference chairs.
- Widespread licensing and access inconsistencies: among the 398 tasks with standard licenses or clearly interpretable licensing information, 80% exhibited some form of inconsistency, categorized as uncertain or borderline cases (43%), inconsistent or misleading (pseudo-open) cases (20%), and potentially non-compliant cases (38%). Specific issues included platform registration under a restrictive license (33%), an additional data usage agreement alongside the stated license (23%), approval-based access under a restrictive license (20%), NoDerivatives clauses (18%), context-limited access under a restrictive license (10%), software-style licenses applied to data (1%), platform registration despite an open license (13%), a declared license with unavailable data (8%), approval-based access despite an open license (5%), context-limited access despite an open license (3%), conflicting license statements across sources (20%), missing or unfindable licenses (16%), conflicting or contradictory license terms (3%), violation of general license terms (2%), and incompatible license combinations within a single source (2%). Table 1 reports no obvious inconsistencies for 20% of tasks.
- Documentation gaps: the challenge website was no longer functional for 3% of tasks; award type was unspecified in 43% and exact prize amounts missing in 56% of monetary cases. Acquisition device information was absent in 43% of tasks. Only 19% of tasks sufficiently described the data splitting strategy, and metadata was not described at all in 25%. De-identification procedures were unreported in 40% of tasks, and ethics committee approval was unreported in 35% of applicable tasks. The study population was undescribed in 41% of tasks and clinical information such as disease status, diagnosis, or treatment history was absent in 40%. Case selection criteria were unreported in 49% and eligibility criteria missing in 50%. Geographic composition of training and test sets was unspecified in 15% and 13% of tasks, respectively. Sex distribution was reported in only 30% of tasks and sufficiently detailed age distribution in only 5%, with 26% reporting partial information such as mean age. Annotation protocols were fully documented in only 40% of tasks, annotation tools were undescribed in 42%, and inter- or intra-rater variability was unreported in 74%.
- Awards and publications: monetary prizes were the most common award (39% of tasks), ranging from 117 EUR to 120,000 EUR (mean 6,046 EUR, median 1,814 EUR, IQR 3,380 EUR). Submission remained possible after the official challenge ended for 31% of challenges. 66% of challenges had an associated publication (57% challenge paper, 25% data publication, 16% both). Median citations per year were 20 (maximum 444) for challenge papers and 22 (maximum 141) for data papers. The median publication delay between opening a challenge for submissions and its challenge paper was 1.6 years (maximum 6.6 years).
Methodology in Plain English
The authors reviewed recent biomedical image analysis challenges and broke each one down into individual tasks (458 tasks across 249 challenges, spanning 2018 to 2023). Information was gathered primarily from challenge websites (41%), challenge papers (24%), and registered challenge design documents (23%). Annotations were performed double-blind by 30 observers. For each task, the team recorded dataset characteristics (origin, modality, problem type, anatomy, size, centers, devices, demographics), access conditions, license terms, and documentation completeness. Geographic and modality representation were then compared against external reference distributions: global population shares and US imaging utilization rates. Licensing and access practices were scored using the (Re)usable Data Project (RDP) framework on a 0–5 scale and categorized into issue types based on the 398 tasks with standard or clearly interpretable licensing information. The authors note that these categories reflect observed practices and inconsistencies rather than definitive legal determinations, since legal frameworks differ and no case-specific legal review was performed.
Why This Matters
Impact on research: Challenge datasets are reused as reference benchmarks, so their limitations propagate into later method development and evaluation. The paper argues that leaderboard success does not equate to clinical readiness, and that benchmarks may preferentially address problems that are easy to measure rather than those of greatest clinical relevance.
Real-world applications:
- Clinical deployment decisions: readers are warned that performance on unrepresentative benchmarks, such as MRI- and segmentation-heavy datasets, is weak evidence of readiness for deployment in settings where ultrasound or X-ray dominate.
- Data governance and licensing: the inventory of problematic practices (conflicting licenses across sources, NoDerivatives clauses, software licenses applied to datasets) gives organizers and institutions concrete failure modes to avoid.
- Dataset reuse and reproducibility: a declared license without accessible data, or approval-based access even under an open license, directly blocks the reuse these benchmarks promise.
- Annotation and reporting standards: the reporting gaps (annotation protocols documented in only 40% of tasks, rater variability unreported in 74%) point to specific improvements that would make results interpretable.
Industry relevance: companies and clinical teams that build on challenge data need confidence that licenses permit commercial or derivative reuse; the finding that only 41% of tasks used an unambiguous license, and that 20% of tasks showed conflicting license statements, exposes concrete legal and practical risk.
Future Directions
- Determine what drives the observed representation patterns — the study design cannot identify whether they stem from annotation effort, data availability, acquisition logistics, or regulatory, institutional, and political frameworks.
- Improve licensing and access practice: reduce unnecessary registration and approval barriers where open licenses are already claimed, and ensure licenses are clear, consistent across sources, and appropriate for data rather than software.
- Close documentation gaps in cohort composition, study design, annotation protocols, and demographic reporting so that datasets are reusable and results interpretable.
- Re-evaluate how benchmark performance is interpreted, since benchmark-specific optimization may not track properties required for clinical deployment, which the paper notes depends on external validation, regulatory approval, and workflow integration.
Target Audience
Challenge organizers and participants, medical imaging AI researchers, dataset curators and data governance staff, journal and conference reviewers who evaluate benchmark claims, and industry teams that rely on public medical imaging datasets for training or validation. Readers interested in the broader question of whether benchmark leaderboards predict clinical usefulness will also benefit.
Authors’ abstract
Benchmarking competitions are central to the development of artificial intelligence (AI) in medical imaging, defining performance standards and shaping methodological progress. However, it remains unclear whether these benchmarks provide data that are sufficiently representative, accessible, and reusable to support clinically meaningful AI. In this work, we assess fairness along two complementary dimensions: (1) whether challenge datasets capture the diversity of real-world clinical data, and (2) whether they are accessible and legally reusable in line with the FAIR principles. To address this question, we conducted a large-scale systematic study of 249 biomedical image analysis challenges comprising 458 tasks across 19 imaging modalities. Our findings reveal limited representation across challenge datasets with respect to geographic location, imaging modalities, and problem types, raising concerns about how well current benchmarks reflect real-world clinical diversity. Despite their widespread influence, challenge datasets were frequently constrained by restrictive or ambiguous access conditions, inconsistent or non-compliant licensing practices, and incomplete documentation, limiting reproducibility and long-term reuse. Together, these shortcomings expose foundational fairness limitations in our benchmarking ecosystem and highlight a disconnect between leaderboard success and clinical relevance.