Skip to content
AI.info

Research

A Multicenter Benchmark of Multiple Instance Learning Models for Lymphoma Subtyping from HE-stained Whole Slide Images

Overview Research area: Computational pathology / computer vision — weakly supervised deep learning for lymphoma subtype classification from hematoxylin and eosin (HE)-stained whole slide images (WSIs

A Multicenter Benchmark of Multiple Instance Learning Models for Lymphoma Subtyping from HE-stained Whole Slide Images
arXiv
2512.14640
Published
2025-12-16
Authors
Rao Muhammad Umer, Daniel Sens, Jonathan Noll, Sohom Dey, Christian Matek, Lukas Wolfseher, Rainer Spang, Ralf Huss, Johannes Raffler, Sarah Reinke, Ario Sadafi, Wolfram Klapper, Katja Steiger, Kristina Schwamborn, Carsten Marr

AI summary

Overview

  • Research area: Computational pathology / computer vision — weakly supervised deep learning for lymphoma subtype classification from hematoxylin and eosin (HE)-stained whole slide images (WSIs).
  • Technical level: Intermediate. Readers benefit from familiarity with multiple instance learning (MIL), pathology foundation models, and WSI preprocessing, though the paper's framing of the clinical problem is accessible.
  • Scope: The paper builds the first multicenter lymphoma subtyping benchmark (four common subtypes plus healthy controls) and systematically benchmarks five publicly available pathology foundation models paired with attention-based and transformer-based MIL aggregators at three magnifications on in-distribution and out-of-distribution data.

What This Paper Is About

Diagnosing lymphoma subtypes usually requires immunohistochemistry, flow cytometry, and molecular genetic tests on top of HE-stained slides — tests that need costly equipment and trained personnel and that delay treatment. This paper asks whether deep learning applied directly to routinely available HE-stained whole slide images can distinguish common lymphoma subtypes, and it tests how well such models hold up when moved from one hospital to another. Because no comprehensive multicenter benchmark for lymphoma subtyping on HE slides existed, the authors assembled one themselves and evaluated it systematically.

Key Contributions

  1. The first multicenter lymphoma benchmark covering the four most frequent subtypes — Chronic Lymphocytic Leukemia (CLL), Follicular Lymphoma (FL), Mantle Cell Lymphoma (MCL), and Diffuse Large B-Cell Lymphoma (DLBCL) — plus healthy control tissue (NEG), totaling 999 HE WSIs from 609 patients across four German centers (Munich, Kiel, Augsburg, Erlangen).
  2. A deep MIL benchmarking pipeline that pairs six pretrained vision encoders (ResNet50 as baseline, H-optimus-1, H0-mini, Virchow2, UNI2, Titan) with four aggregator variants (AB-MIL, TransMIL, TransMIL + BEL, MoA).
  3. A systematic magnification study comparing 10x, 20x, and 40x resolution, including cross-magnification aggregation, to determine how much resolution lymphoma subtyping actually needs.
  4. An in-distribution versus out-of-distribution evaluation: models trained on the Munich cohort are tested on held-out Munich data and on external cohorts from Kiel, Augsburg, and Erlangen, quantifying generalization loss across clinical sites. The code is released at https://github.com/RaoUmer/LymphomaMIL.

Main Findings

  • In-distribution accuracy exceeds 80%: On in-distribution test sets, models achieved multiclass balanced accuracies exceeding 80% across all magnifications, with the five foundation models performing similarly to one another and AB-MIL and TransMIL aggregators showing comparable results.
  • 40x resolution is sufficient: The magnification study found that 40x resolution is sufficient for accurate classification, with no performance gains from higher resolutions or from cross-magnification aggregation.
  • Large out-of-distribution drop: On external test sets, performance dropped substantially to around 60% balanced accuracy. The paper's limitations section frames this as a 20% decrease and attributes it to scanner-specific color variation, staining differences, and shifts in patient cohorts across clinical sites.
  • UNI2 with AB-MIL generalizes most robustly: UNI2 paired with AB-MIL achieved the best overall performance on the OOD Kiel cohort (AUC 0.88 ± 0.05, F1 0.61 ± 0.06, BACC 0.62 ± 0.04) and on the OOD Augsburg cohort (AUC 0.82 ± 0.04, F1 0.58 ± 0.05, BACC 0.60 ± 0.06). Across all OOD evaluations, UNI2 + AB-MIL showed the most robust and stable generalization behavior. The authors suggest Titan's slide-level token reduction improves efficiency but limits capture of subtle, spatially localized lymphoma patterns under domain shift.
  • Erlangen results are unstable by construction: On the Erlangen cohort, which contains only CLL cases, UNI2 combined with TransMIL yielded the highest accuracy (AUC 0.90 ± 0.05, F1 0.75 ± 0.35, BACC 0.83 ± 0.24). The authors attribute the high standard deviation to the very small sample size (29 WSIs), the single-class composition, and data acquired with a different scanner, and they report these results to highlight the difficulty of OOD evaluation on small single-class cohorts rather than as evidence of reliable generalization.
  • Aggregator differences are modest: TransMIL combined with UNI2 or Titan was better than AB-MIL at 40x and 20x, with a slight performance decline at 10x, but AB-MIL remained the most computationally efficient option for training and inference compared to transformer-based MIL methods.
  • BEL adds little: TransMIL + BEL showed only minimal gains over TransMIL. The authors attribute this to a mismatch between the BEL objective and lymphoma morphology: BEL regularizes bag-level embeddings, whereas lymphoma subtyping relies on heterogeneous, spatially dispersed patterns rather than a single dominant bag embedding. Further parameter tuning did not lead to improvements.
  • Mixtures of aggregator (MoA) did not win: The multi-resolution fusion strategy MoA did not outperform the standard AB-MIL or TransMIL baselines. On the combined 40x/20x/10x setting with UNI2 it reached AUC 0.94 ± 0.03, F1 0.76 ± 0.08, and BACC 0.77 ± 0.07.
  • In-distribution encoder comparison (Munich, 40x, AB-MIL): Titan reached AUC 0.96 ± 0.01, F1 0.83 ± 0.06, BACC 0.82 ± 0.05; H-optimus-1 0.96 ± 0.01, 0.79 ± 0.03, 0.79 ± 0.03; H0-mini 0.95 ± 0.01, 0.79 ± 0.02, 0.78 ± 0.01; Virchow2 0.96 ± 0.01, 0.77 ± 0.04, 0.77 ± 0.04; UNI2 0.95 ± 0.01, 0.79 ± 0.02, 0.79 ± 0.03; ResNet50 baseline 0.89 ± 0.03, 0.66 ± 0.06, 0.67 ± 0.06.
  • In-distribution encoder comparison (Kiel, 40x, AB-MIL): H-optimus-1 reached AUC 0.97 ± 0.01, F1 0.81 ± 0.09, BACC 0.81 ± 0.09; Titan 0.96 ± 0.02, 0.83 ± 0.06, 0.81 ± 0.05; UNI2 0.96 ± 0.03, 0.82 ± 0.07, 0.82 ± 0.06; Virchow2 0.96 ± 0.02, 0.72 ± 0.11, 0.74 ± 0.09; H0-mini 0.92 ± 0.05, 0.66 ± 0.15, 0.69 ± 0.13; ResNet50 baseline 0.85 ± 0.05, 0.64 ± 0.08, 0.65 ± 0.04.
  • All foundation models cluster together in-distribution: The authors conclude that current pathology foundation models have reached a similar level of representation quality for lymphoma subtyping, while the baseline ResNet50 trails them.
  • Efficiency varies widely: Measured parameters, FLOPs, and throughput (see Table 4) were: H-optimus-1, 1,135M parameters, 591.81 GFLOPs, 82.5 ± 4.3 images/s; H0-mini, 87M, 44.60 GFLOPs, 885.1 ± 7.1 images/s; Virchow2, 632M, 329.11 GFLOPs, 134.3 ± 7.9 images/s; UNI2, 681M, 360.71 GFLOPs, 134.6 ± 7.6 images/s; Titan, 85M, 36.76 GFLOPs, 43.95 ± 2.3 slides/s.
  • Attention maps are noisy and heterogeneous: Attention distributions remain highly heterogeneous across whole slides, indicating informative regions are not trivially separable. High attention was given to non-lymphoma areas in CLL and to poorly preserved areas in DLBCL; UNI2 avoided artifact patches in the highest attention scores compared to Titan.
  • Not directly comparable to other benchmarks: The authors state that their lymphoma subtyping findings should be interpreted as complementary to the "all-task" benchmark rather than directly comparable, because that benchmark does not include lymphoma subtyping and TCGA offers only a single lymphoma subtype (DLBCL) with limited histological diversity.

Methodology in Plain English

The researchers gathered archived HE-stained lymphoma slides from four German pathology institutes. The Munich cohort contains 513 WSIs from 261 patients scanned at 40x with an Aperio AT2 scanner; Kiel contains 177 WSIs from 151 patients scanned at 40x with a Hamamatsu scanner; Augsburg contains 290 de-identified WSIs from 168 patients scanned at 40x with a Philips scanner; and Erlangen contains 29 WSIs from 29 patients, CLL only. Munich, Kiel, and Augsburg include FL, CLL, MCL, and DLBCL; Munich and Kiel additionally include a healthy-tissue control class (NEG), while Augsburg does not. Some Kiel slides come from the same patient because the same biopsy was scanned twice rather than sectioned differently.

For preprocessing, they used Trident, a DeepLabV3-based segmentation model, to separate tissue from background in a stain-agnostic way and to extract non-overlapping 256x256-pixel patches at a chosen magnification. Pen marks, blurring, compression, and water bubbles were present in the slides.

Each patch is then compressed into a single feature vector by a frozen pretrained encoder — no fine-tuning of the encoder. The stacked patch features for a slide form a "bag," and a trainable aggregation layer combines them into one slide-level representation used for the five-class prediction. The four aggregation variants tested were AB-MIL (attention-based), TransMIL (transformer-based), TransMIL + BEL, and MoA (a multi-resolution fusion approach). Supervision is weak: only slide-level labels are available, no patch-level annotations.

Data were split patient-wise for five-fold cross-validation with an 80% / 10% / 10% train / validation / test ratio, so all images from one patient stay in the same fold. Class imbalance was addressed by balanced sampling at the patient level during aggregator training rather than slide-level label balancing. Models were trained for 50 epochs with the AdamW optimizer, a warm-up phase of 10 epochs raising the learning rate linearly from 0 to 0.0001, then cosine decay down to 1e-6, weight decay increased from 0.04 to 0.4, and a momentum factor of 0.9. Performance was measured with Area Under the ROC Curve (AUC), Macro F1-score, and Balanced Accuracy (BACC), with confusion matrices used to interpret error distribution.

For the Augsburg OOD evaluation, the label space is mismatched (no NEG class), so the authors report four-class metrics using the original predictions and consider the top-2 class for cases where NEG was the top-1 prediction. Attention scores were converted to percentiles per predicted class and mapped back to slide coordinates, with overlapping patches averaged, to produce heatmaps via Trident.

Why This Matters

Impact on research. The paper establishes a reproducible, multicenter benchmark for a disease area the authors argue is underrepresented in existing pathology benchmarks, and it releases an automated pipeline so that others can evaluate new models externally. Its central scientific message is cautionary and useful: strong in-distribution numbers (above 80% balanced accuracy) do not transfer automatically across sites (around 60% balanced accuracy), and neither more magnification nor fusion of multiple magnifications fixes that gap.

Real-world applications.

  • Triage and case prioritization: An AI-augmented pathology workflow could guide initial screening to prioritize cases requiring expert review, potentially reducing turnaround times for lymphoma diagnosis.
  • Reducing auxiliary testing burden: Because HE slides are inexpensive and widely available compared with immunohistochemistry, flow cytometry, and molecular genetic tests, accurate HE-based subtyping could help minimize the number of auxiliary tests needed for accurate diagnosis.
  • Efficient use of specialized diagnostics: Flagging likely subtypes early could allow more efficient allocation of costly specialized diagnostic tests and trained personnel.
  • Cross-site model auditing: The released benchmarking pipeline lets hospitals and developers test whether a model trained at one center degrades at another before deployment.

Industry relevance. Pharmaceutical and diagnostics companies, digital pathology vendors, and clinical AI developers all depend on knowing whether pathology foundation models generalize across scanners and staining protocols. The efficiency table is directly actionable for deployment planning on constrained hardware: the models differ by more than two orders of magnitude in parameter count (from 85M to 1,135M) and roughly an order of magnitude in throughput (from 82.5 to 885.1 images/s), while in-distribution accuracy differences among them are small.

Future Directions

  1. Larger multicenter studies with standardized protocols: The authors state that only German centers were included and that larger European and worldwide initiatives are needed.
  2. Coverage of rare lymphoma subtypes: The current benchmark covers four common subtypes and healthy controls; the authors call for expanding to rare subtypes that are also affected by class imbalance in the current cohorts.
  3. Domain adaptation and stain normalization: To close the roughly 20% performance drop across sites, the paper proposes incorporating stain normalization or domain adaptation techniques to improve cross-site generalization, plus prospective clinical validation studies.
  4. Better patch selection and multi-scale modeling: Because attention distributions remain highly heterogeneous across slides and high-attention regions sometimes fall on non-lymphoma or poorly preserved areas, the authors argue that more advanced and principled patch-selection or multi-scale modeling strategies are needed, and they leave this to future work. They also suggest integrating complementary modalities such as immunophenotyping data or clinical parameters, and note that attention behavior on non-lymphoma areas in CLL and poorly preserved areas in DLBCL might be avoidable with more classes during training.

Target Audience

This paper is most useful to computational pathology and medical imaging researchers working on weakly supervised whole slide image classification, to clinical AI developers and digital pathology vendors deciding which foundation models and MIL aggregators to deploy, and to pathologists and translational oncologists interested in what HE-stained-slide-only lymphoma subtyping can realistically deliver today. Machine learning engineers focused on domain shift and out-of-distribution generalization will also find the multicenter evaluation design instructive. Researchers looking for a ready-made baseline pipeline for a new cohort are a direct audience, since the code is publicly available. The paper is less suited to readers seeking a deployable clinical tool: the authors are explicit that the OOD results should not be read as reliable generalization, and that larger multicenter studies with standardized protocols are needed before clinical utility can be claimed.

Authors’ abstract

Timely and accurate lymphoma diagnosis is essential for guiding cancer treatment. Standard diagnostic practice combines hematoxylin and eosin (HE)-stained whole slide images with immunohistochemistry, flow cytometry, and molecular genetic tests to determine lymphoma subtypes, a process requiring costly equipment, and skilled personnel, causing treatment delays. Deep learning methods could assist pathologists by extracting diagnostic information from routinely available HE-stained slides directly, yet comprehensive benchmarks for lymphoma subtyping on multicenter data are lacking. In this work, we present the first multicenter lymphoma benchmark, covering four common lymphoma subtypes and healthy control tissue. We systematically evaluate five publicly available pathology foundation models (H-optimus-1, H0-mini, Virchow2, UNI2, Titan) combined with attention-based (AB-MIL) and transformer-based (TransMIL) multiple instance learning aggregators across three magnifications (10x, 20x, 40x). On in-distribution test sets, models achieve multiclass balanced accuracies exceeding 80% across all magnifications, with foundation models performing similarly, and aggregation methods showing comparable results. The magnification study reveals that 40x resolution is sufficient, with no performance gains from higher resolutions or cross-magnification aggregation. However, on out-of-distribution test sets, performance drops substantially to around 60%, highlighting significant generalization challenges. To advance the field, larger multicenter studies covering additional rare lymphoma subtypes are needed. We provide an automated benchmarking pipeline to facilitate such future research. Our paper codes is publicly available at https://github.com/RaoUmer/LymphomaMIL.

Read the original paper