Skip to content
AI.info

Research

EvalBlocks: A Modular Pipeline for Rapidly Evaluating Foundation Models in Medical Imaging

Overview Research area: Medical imaging, foundation models, and machine-learning tooling/benchmarking (computer vision). Technical level: Intermediate. The paper is a software-framework paper rather t

EvalBlocks: A Modular Pipeline for Rapidly Evaluating Foundation Models in Medical Imaging
arXiv
2601.03811
Published
2026-01-07
Authors
Jan Tagscherer, Sarah de Boer, Lena Philipp, Fennie van der Graaf, Dré Peeters, Joeran Bosma, Lars Leijten, Bogdan Obreja, Ewoud Smit, Alessa Hering

AI summary

Overview

Research area: Medical imaging, foundation models, and machine-learning tooling/benchmarking (computer vision).

Technical level: Intermediate. The paper is a software-framework paper rather than a new-model paper; it assumes familiarity with foundation models, embeddings, k-NN and linear probing, but the pipeline concepts are explained conceptually.

Scope (one sentence): EvalBlocks is an open-source, Snakemake-based pipeline that lets researchers plug datasets, foundation models, aggregation methods, and evaluation strategies together and run reproducible, cached, parallel downstream evaluations of medical imaging foundation models.

What This Paper Is About

Developing medical imaging foundation models requires repeatedly estimating downstream performance, but researchers typically do this with bespoke, ad-hoc scripts that are slow, error-prone, and hard to reproduce. The paper introduces EvalBlocks, a modular "plug-and-play" evaluation framework that turns these one-off workflows into configurable, cacheable, centrally tracked pipeline blocks. The authors demonstrate it by evaluating five foundation models across three patch-level medical imaging classification tasks.

Key Contributions

  1. A modular, extensible, and efficient evaluation framework for foundation models in medical imaging, released as open source software at https://github.com/DIAGNijmegen/eval-blocks.
  2. A pipeline architecture built on Snakemake in which pipeline steps are self-contained blocks grouped into feature models, optional aggregation steps, and evaluation procedures, with inputs and outputs declared declaratively in a configuration file.
  3. A demonstration that evaluates five foundation models on three medical imaging classification tasks, with centrally tracked, single-command reproducible experiments and caching plus parallel execution on shared cluster infrastructure such as Slurm.
  4. Preconfigured public blocks for the five evaluated models, including all necessary preprocessing steps, enabling immediate plug-and-play use.

Main Findings

  • Five models, three tasks: All combinations of foundation models, aggregation methods, and evaluation strategies were run through the pipeline, producing a comprehensive set of metrics and visualizations for each configuration.
  • CT dataset comparison (PANORAMA and AMARA): CT-FM and Curia perform best on PANORAMA, while UMedPT is slightly more accurate on AMARA (Fig. 2, with error bars showing standard deviation across folds).
  • MRI modality findings (PI-CAI): Overall, ADC is the most informative modality for malignancy discrimination, and modality mean aggregation emerges as a well-performing strategy for this task (Fig. 3, which shows four foundation models on PI-CAI).
  • Embedding visualizations (Curia, PANORAMA, first fold): PCA and t-SNE yield no clusters, whereas LDA shows two distinct peaks for the two classes, indicating the model produces linearly separable embeddings in label-dependent directions but not in directions of maximum variance or local neighborhood structure.
  • Efficiency: Caching avoided recomputing embeddings across experiments and, combined with parallel execution, substantially reduced wall-time. No specific speed-up figure or absolute accuracy/AUC values are reported in the paper text.
  • Positioning versus benchmarks: Existing benchmarks (e.g., the clinically relevant tasks of Wang et al., the fairness assessment of Jin et al., and the UNICORN challenge) favor comprehensiveness as static leaderboards, whereas EvalBlocks targets rapid, iterative evaluation during development.

Methodology in Plain English

The pipeline is assembled from independent Snakemake rules that declare what they consume, what they produce, and what resources they need; a rule runs automatically once its inputs exist. Rules fall into three categories: feature models that turn input patches into embeddings, optional aggregation steps, and evaluation procedures. Intermediate outputs are cached so they can be reused, and experiments are described declaratively in a configuration file so a user can run selected combinations or all of them, locally or distributed on a cluster such as Slurm.

For the demonstration, patches of size 224 × 224 × 16 with malignancy labels were extracted from three datasets: AMARA (in-house; 161 malignant and 502 benign pulmonary nodules from 320 patients, with labels determined by pathological confirmation), PANORAMA (675 CT patches of healthy pancreatic tissue and 675 with ductal adenocarcinoma), and PI-CAI (219 MR patches depicting prostate carcinoma and 219 with healthy prostate tissue, from the public test set). Each dataset provides training and test splits across five folds, and input data were preprocessed according to each model author's specifications: three-dimensional-capable models receive the whole patch, two-dimensional architectures receive the central slice, and because DINOv2 and DINOv3 were trained on natural images, their input slices are interpreted as grayscale images with values between 0 and 255.

Five foundation models were evaluated: CT-FM (the only model trained on three-dimensional CT scans as its only modality), Curia (unsupervised training on a large dataset of medical images), UMedPT (the only model trained in a supervised manner), and DINOv2 and DINOv3 (trained on natural images rather than medical data). For the MRI dataset, embeddings were aggregated across modalities by computing the element-wise mean of feature vectors; the framework also supports custom aggregation such as weighted averaging, attention-based fusion, or case-level pooling. Three interchangeable evaluation strategies operate on the optionally aggregated embeddings: a k-Nearest Neighbors classifier with k ∈ {10, 20, 100, 200} reporting accuracy and AUC on the test split (results for k = 20 shown), a single linear layer trained with cross entropy loss at a learning rate of 1e-5 with accuracy and AUC, and visual analyses using linear discriminant analysis, principal component analysis, and t-SNE.

Why This Matters

Impact on research: The framework separates evaluation logistics from model innovation. Because all experiments and results are tracked centrally and reproducible with a single command, comparisons between models and checkpoints become fast and automated, and the ability to run locally made assessment on in-house datasets possible. It bridges the gap between large-scale static benchmarking and practical, iterative experimentation.

Real-world applications:

  • Evaluating and selecting foundation models for clinical imaging tasks such as lung nodule, pancreatic, and prostate cancer classification before deployment in downstream pipelines.
  • Diagnostics and clinical decision support, where the malignancy classification tasks (pulmonary nodules, ductal adenocarcinoma, prostate carcinoma) reflect real screening and diagnostic use cases.
  • Research in data-scarce settings, since the pipeline targets few-shot adaptation of pretrained embeddings via k-NN and linear probing.
  • Shared hospital and academic compute environments, where caching and distributed execution on clusters such as Slurm make evaluation scalable.

Industry relevance: Teams building or adopting medical imaging foundation models need reproducible evidence about which pretrained model, which imaging modality, and which aggregation strategy works best. EvalBlocks offers that infrastructure as open source software, reducing both computational and manual effort, while comparable lightweight evaluation tooling in other domains (Hugging Face's LightEval, NVIDIA's NeMo Evaluator SDK) has shown the value of this kind of tooling for large language models.

Future Directions

  • Extending EvalBlocks to additional task types such as segmentation and detection beyond classification.
  • Integrating with existing popular platforms like Hugging Face to improve community collaboration.
  • Managing the combinatorial explosion of evaluations as the number of datasets and models grows, building further on Snakemake's caching and parallelization and on selective subset execution.
  • Continued expansion of the block library (new datasets, models, aggregation methods, evaluation strategies) so that researchers can focus on architecture, training strategy, and downstream adaptation.

Target Audience

Researchers and engineers developing or applying foundation models in medical imaging, particularly those running many downstream experiments across datasets, models, and preprocessing choices; machine-learning engineers and MLOps practitioners in academic medical centers or industry who need reproducible, cluster-scalable evaluation pipelines; and readers of the German Conference on Medical Image Computing (Bildverarbeitung für die Medizin 2026) proceedings, where this work is published.

Authors’ abstract

Developing foundation models in medical imaging requires continuous monitoring of downstream performance. Researchers are burdened with tracking numerous experiments, design choices, and their effects on performance, often relying on ad-hoc, manual workflows that are inherently slow and error-prone. We introduce EvalBlocks, a modular, plug-and-play framework for efficient evaluation of foundation models during development. Built on Snakemake, EvalBlocks supports seamless integration of new datasets, foundation models, aggregation methods, and evaluation strategies. All experiments and results are tracked centrally and are reproducible with a single command, while efficient caching and parallel execution enable scalable use on shared compute infrastructure. Demonstrated on five state-of-the-art foundation models and three medical imaging classification tasks, EvalBlocks streamlines model evaluation, enabling researchers to iterate faster and focus on model innovation rather than evaluation logistics. The framework is released as open source software at https://github.com/DIAGNijmegen/eval-blocks.

Read the original paper