Research
Atlas 2 -- Foundation models for clinical deployment
Overview Research area: Computational pathology / medical computer vision — self-supervised vision foundation models for histopathology whole slide images. Technical level: Intermediate. Scope: This r

- arXiv
- 2601.05148
- Published
- 2026-01-08
- Authors
- Maximilian Alber, Timo Milbich, Alexandra Carpen-Amarie, Stephan Tietz, Jonas Dippel, Lukas Muttenthaler, Beatriz Perez Cancer, Alessandro Benetti, Panos Korfiatis, Elias Eulig, Jérôme Lüscher, Jiasen Wu, Sayed Abid Hashimi, Gabriel Dernbach, Simon Schallenberg, Neelay Shah, Moritz Krügener, Aniruddh Jammoria, Jake Matras, Patrick Duffy, Matt Redlon, Philipp Jurmeister, David Horst, Lukas Ruff, Klaus-Robert Müller, Frederick Klauschen, Andrew Norgan
AI summary
Overview
- Research area: Computational pathology / medical computer vision — self-supervised vision foundation models for histopathology whole slide images.
- Technical level: Intermediate.
- Scope: This report introduces Atlas 2, Atlas 2-B, and Atlas 2-S, three pathology vision foundation models trained on 5.5 million histopathology whole slide images and evaluated across eighty public benchmarks for prediction performance, robustness, and resource efficiency.
What This Paper Is About
Pathology foundation models have improved computational pathology, but tradeoffs between prediction performance, robustness to confounding variation, and computational cost have limited their use in routine clinical diagnostics. This paper addresses those tradeoffs by scaling up training data, model capacity, and training recipe, then distilling the result into two lightweight variants. The goal is a model family that is simultaneously state of the art, robust, and cheap enough to deploy.
Key Contributions
- Atlas 2: A 2 billion parameter Vision Transformer (ViT, patch-token size 8) trained on a multi-centric corpus of 5.5 million de-identified histopathology whole slide images from Charité - Universitätsmedizin Berlin, LMU Munich, and Mayo Clinic — described as the largest pathology foundation model dataset to date. Tiles were extracted at 0.25, 0.5, 1.0, and 2.0 microns per pixel (40x, 20x, 10x, and 5x magnifications).
- Two distilled lightweight models: Atlas 2-B (ViT-B, 86 million parameters) and Atlas 2-S (ViT-S, 22 million parameters), distilled from Atlas 2 and reported as 24 and 91 times smaller in parameter count and 3.4 and 9 times more resource efficient for standard-size image inference.
- A comprehensive multi-framework evaluation: eighty public benchmark datasets drawn from five public evaluation frameworks — eva, HEST, Plismbench, PathoROB, and Patho-Bench — plus additional tumor microenvironment (TME) and MSI prediction benchmark datasets (MSI CRC, MSI STAD, TCGA Uniform) evaluated on an in-house framework.
- A combined performance-robustness and performance-efficiency analysis: Figure 1 positions the Atlas 2 family on the state-of-the-art Pareto fronts for both robustness and resource efficiency.
Main Findings
- Best overall large-model performance: Atlas 2 achieves the best performance in 22/27 tasks and second-best in 3 additional tasks, with an average of 44.8% on HEST (a 1.6 percentage point improvement over second-best Pluto-4G) and 82.9% on eva (a 1.2 percentage point improvement over Pluto-4G, on the eva benchmark subset evaluated in that prior work).
- Best robustness among large models: Atlas 2 reaches an average robustness score of 85.7%, a 9.7 percentage point improvement over the closest contender, Virchow2.
- Improvements over H-Optimus-1: On the shared benchmarks HEST, CRC-100k, and MHIST, Atlas 2 improves by 1, 1.5, and 4.5 percentage points respectively.
- Patho-Bench molecular tasks: Atlas 2 is best in 15/24 molecular related tasks and at least second-best in 19/24, with an average improvement of 3.6 percentage points over the second-best model, Virchow2.
- Patho-Bench morphology tasks: Atlas 2 is top in 14/19 tasks and at least second-best in 16/19, with an average difference of 3.3 percentage points over the second-best model.
- Treatment response: Atlas 2 shows the highest performance, a 1.8 percentage point improvement over the second-best model.
- Survival prediction is the weak spot: H-Optimus-0 is the best-performing model on survival prediction, and the paper notes that all models perform close to chance level (C-Index of 50).
- Small models are competitive: Atlas 2-B and Atlas 2-S are best-performing in 27/27 and 25/27 tasks respectively within their compute category. Atlas 2-B runs standard-size image inference more than twice as fast as Pluto-4G (174 vs. 79 img/s) while performing on par, and outperforms the larger Atlas and UNI2-H (66.3 vs. 65.4 and 65.3). Atlas 2-S outperforms the larger H-Optimus-0 and Virchow2 (64.8 vs. 64.0 and 63.9) while being 6.6 and 4.3 times more resource efficient respectively.
- Small models are more robust than the large one: Atlas 2-B and Atlas 2-S reach robustness averages of 86.2 and 85.9, above Atlas 2's 85.7.
- H0-mini comparison: On the CLS+MEAN representation, average robustness on PathoROB is 11.6 percentage points better for Atlas 2-B (93.1 vs. 81.5). On the CLS token evaluation for Plismbench, Atlas 2-B outperforms H0-mini by 4.4 percentage points (58.5 vs. 54.1).
Methodology in Plain English
The team collected 5.5 million de-identified whole slide images from three institutions' digital archives and cut image tiles out of them at four magnification levels. They trained a large vision transformer (2 billion parameters) using a self-supervised recipe adapted from their earlier RudolfV and Atlas approaches, with parts based on DINOv2 and selective improvements from DINOv3, running on NVIDIA B200 and NVIDIA A100 GPUs. They then distilled that large model into two smaller architectures, ViT-B and ViT-S, to produce deployment-friendly versions.
For evaluation, they froze the encoders and only trained small task-specific prediction heads on the resulting embeddings, which isolates the quality of the learned representations. Unless stated otherwise, results use a concatenation of the CLS token and the mean of patch tokens ("CLS+MEAN"); CLS-only results are in Appendix A.2. They benchmarked 15 foundation models using publicly available weights (including UNI2-H, H-Optimus-0, Midnight-12k, Virchow2, Prov-GigaPath, UNI, Phikon-v2, Hibou-B, and several ViT-B/ViT-S models) and compared against 5 more models (Pluto-4G, Pluto-4S-8, Pluto-4S-16, H-Optimus-1, H0-mini) using published results because weights were unavailable.
Evaluation harnesses: HEST uses Ridge Regression with PCA (256 factors) on embeddings, scored by Pearson correlation. eva uses linear heads for patch classification, patch-token segmentation, and ABMIL for slide-level tasks, scored by balanced accuracy and Dice score without background. PathoROB measures robustness against non-biological confounders. Plismbench measures representation consistency across scanners and stains using n = 8,139 tiles. Patho-Bench uses the ABMIL protocol on a subset of 53 tasks drawn from its 95 tasks across seven subcategories, repeating each task three times to account for head initialization and optimization variance.
Why This Matters
Research impact. The paper is a direct test of whether the usual performance / robustness / efficiency trilemma in pathology foundation models can be reduced rather than merely traded off. Its dataset scale (5.5 million whole slide images from three institutions), model scale (2 billion parameters), and evaluation breadth (eighty benchmarks across five public frameworks) set a reference point for future comparisons, and the finding that distilled small models can be more robust than their teacher is a notable empirical result.
Real-world applications (as reported or directly supported by the paper's benchmarks):
- Tumor microenvironment profiling from routine H&E, via the authors' Atlas H&E-TME application and the derived OpenTME dataset of cell-level profiles across thousands of TCGA slides.
- Gene expression prediction directly from H&E-stained tissue (the HEST tasks), including tasks spanning kidney, colon, breast, lung, lymph node, pancreas, prostate, rectum, and skin cancer.
- Mutation and molecular subtyping prediction from histology, including MSI prediction and HER2 status.
- Tumor grading and cancer detection, such as Prostate cancer ISUP/Gleason grading (PANDA, Gleason Arvaniti) and lymph node metastasis detection (CAMELYON16, PCAM).
Industry relevance. The models are framed explicitly around clinical deployment constraints: GPU inference cost, throughput in images per second, and robustness to scanner, staining, and lab processing variation that would otherwise degrade performance in routine diagnostics. A model that is roughly a magnitude more efficient while matching larger competitors matters for labs processing large slide volumes under cost and latency limits.
Future Directions
- Survival prediction remains unsolved. All evaluated models score near chance level (C-Index of 50), and H-Optimus-0 rather than Atlas 2 leads this task, so improvements here are clearly still needed.
- Further robustness benchmarking. H0-mini could not be compared on every robustness benchmark due to a lack of reference values and model access, leaving gaps in the comparison landscape.
- Broader head-to-head comparisons. Only a limited comparison to H-Optimus-1 was possible, and comparisons to Pluto-4G, Pluto-4S-8, Pluto-4S-16, H-Optimus-1, and H0-mini relied on published numbers rather than re-evaluation.
- From foundation models to deployed applications. The authors state further improvements and benchmarking are warranted, and point toward applications built on the Atlas family (Atlas H&E-TME, OpenTME) as the route to scalable, quantitative tissue analysis in practice.
Target Audience
Researchers and engineers in computational pathology and medical computer vision; machine learning practitioners working on self-supervised and distilled vision models; pathologists and clinical informatics teams assessing whether a foundation model is deployable under real resource and robustness constraints; and industry teams building diagnostic or biomarker-prediction products on histopathology data. Readers interested primarily in benchmark design (multi-framework evaluation, robustness indexing, cross-scanner consistency measurement) will also find the evaluation protocol sections useful.
Authors’ abstract
Pathology foundation models substantially advanced the possibilities in computational pathology --- yet tradeoffs in terms of performance, robustness, and computational requirements remained, which limited their clinical deployment. In this report, we present Atlas 2, Atlas 2-B, and Atlas 2-S, three pathology vision foundation models which bridge these shortcomings by showing state-of-the-art prediction performance, robustness, and resource efficiency in a comprehensive evaluation across eighty public benchmarks. Our models were trained on the largest pathology foundation model dataset to date comprising 5.5 million histopathology whole slide images, collected from three medical institutions Charité - Universitätsmedizin Berlin, LMU Munich, and Mayo Clinic.