Skip to content
AI.info

Research

CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models

Overview Research area: Computational pathology and representation learning — specifically, benchmarking pathology foundation models (PFMs) at single-cell resolution on whole-slide images. Technical l

arXiv
2608.21060
Published
2026-08-21
Authors
Bokai Zhao, Yiyang Zhang, Hanqing Chao, Yawei Ma, Long Bai, Tai Ma, Minfeng Xu, Ming Song, Tianzi Jiang

AI summary

Overview

Research area: Computational pathology and representation learning — specifically, benchmarking pathology foundation models (PFMs) at single-cell resolution on whole-slide images.

Technical level: Advanced. The paper assumes familiarity with whole-slide image (WSI) processing, spatial transcriptomics (Xenium), linear probing, and cross-domain transfer protocols.

Scope in one sentence: CellPath-Bench is a benchmark that keeps 30 pathology-specific and general-purpose foundation models frozen and measures how much cell-type information is linearly decodable from their nucleus-anchored features, and how well that information transfers across tissue sections, datasets, and organs.

What This Paper Is About

Existing PFM benchmarks evaluate downstream utility through patch-level classification, region-level dense prediction, or whole-slide prediction, but those scores depend jointly on the evaluation's spatial unit, feature aggregation, prediction head, and adaptation strategy. That makes it hard to tell whether a model actually preserves cell-level information or just exploits tissue composition and regional texture.

The paper builds a controlled benchmark where the cellular reference, spatial readout definitions, and classifier capacity are held constant across heterogeneous model architectures, so differences can be attributed to the frozen representation itself. It then measures both absolute cell-type decodability and cross-domain transferability across sections, datasets, and anatomical organs.

Key Contributions

  1. A coordinate-aligned framework for measuring cell representations in frozen foundation models. Frozen WSI feature fields are sampled at registered nuclear coordinates, and different spatial readouts are compared using a unified multiclass linear probe. Within this framework, Nuc performance measures absolute cell-type decodability, Cell Representation Advantage (CRA) measures sensitivity to spatial readout within tissue sections, and Cell Representation Transferability (CRT) summarizes relative model performance across sections, datasets, and organs.

  2. A quality-controlled, multi-organ cellular reference. Starting from 52 candidate Xenium datasets, the authors construct an evaluation panel of 25 spatially registered H&E–Xenium tissue sections spanning 11 organs and 7,079,283 cells, with molecularly informed annotations derived from 14 labeled scRNA-seq references (one normal-tissue atlas and 13 cancer-specific references), harmonized fine- and coarse-grained taxonomies, and independent histological concordance assessment.

  3. A large-scale, multidimensional evaluation of foundation models. 30 pathology-specific and general-purpose foundation models are benchmarked through 304,920 controlled linear-probe runs across spatial readouts, magnifications, taxonomic granularities, and evaluation protocols, assessing both within-section decodability and transfer across sections, datasets, and organs.

  4. A panel-relative diagnostic view rather than a single ranking. By separating absolute decodability, sensitivity to spatial readout, and cross-domain generalization, the benchmark produces multidimensional capability profiles that a single downstream score cannot capture.

Main Findings

  • Nucleus-anchored sampling beats patch averaging everywhere. Under the intra-section (IS) protocol at 20x, the Nuc readout consistently outperformed the Mean readout across all 30 foundation models and both cell-type taxonomies, and positive CRA was observed across all 30 models. The authors interpret this as nucleus-anchored sampling retaining cell-type information that is less accessible after patch-level spatial averaging.

  • CRA and CRT are not fully aligned. The joint CRA–CRT landscape in Figure 1 shows that within-section Cell Representation Advantage and cross-domain Cell Representation Transferability reveal complementary properties; models are not ordered identically on the two axes.

  • CRA magnitudes are modest and vary widely. At 20x, CRA ranged from 10.14 ± 6.05 percentage points for H-Optimus-1 (29 significant wins) down to 4.61 ± 3.75 for OmiCLIP (0 significant wins). Virchow2 (10.12 ± 5.82), H-Optimus-0 (9.96 ± 5.89), UNI2 (9.94 ± 6.03), and SEAL(UNI2) (9.89 ± 6.03) followed, each with 25 significant wins.

  • Highest absolute decodability also came from H-Optimus-1. Under Table 4's organ-balanced IS results at 20x, H-Optimus-1 reached 36.73 ± 8.51 Macro-F1 on the fine taxonomy with Nuc (versus 26.95 ± 6.52 for Mean) and 56.88 ± 8.96 on the coarse taxonomy with Nuc (versus 47.81 ± 5.78 for Mean).

  • Coarse taxonomies score higher but rankings stay stable. The coarse-grained taxonomy consistently yielded higher Macro-F1 than the fine-grained taxonomy, yet model rankings remained highly consistent, with Spearman's rho of 0.994, 0.956, 0.949, and 0.933 for IS, IOCD, MOCV, and LOOO respectively.

  • 20x largely preserves cell-type information; 10x does not. Under IS, performance retention relative to 40x was 99.6% at 20x and 92.9% at 10x.

  • Fusing context did not consistently help. The Cls readout was less informative than Nuc, and combining Nuc with Mean or Cls provided no consistent improvement across the seven representation modes tested.

  • Top cross-domain performers under all three transfer protocols. SEAL(UNI2) led IOCD (fine Macro-F1 28.82 ± 8.28, Macro-AUROC 68.19 ± 10.21; coarse F1 49.20 ± 8.84; 27 significant wins), MOCV (fine F1 29.14 ± 5.30, AUROC 74.91 ± 4.38; coarse F1 60.19 ± 7.49; 27 significant wins), and LOOO (fine F1 23.65 ± 7.41, AUROC 68.98 ± 8.36; coarse F1 53.51 ± 10.30; 25 significant wins). UNI2 tracked it almost exactly.

  • Leave-one-organ-out is the hardest setting. Fine-taxonomy LOOO F1 was consistently lower than IOCD and MOCV F1 for models across the panel; OmiCLIP was lowest at 15.90 ± 4.74 (fine) and 44.70 ± 8.38 (coarse), with 0 significant wins in all three cross-domain protocols.

  • Scale of pretraining and parameters alone did not determine cell-level capability. The evaluated models span the PV, PVL, PVO, and GV paradigms with parameter counts from 21M (Kaiko ViT-S/8 and ViT-S/16, Lunit ViT-S) to 1100M (H-Optimus-1, H-Optimus-0, Prov-GigaPath), and the paper notes that models without comparable pretraining-scale metadata are omitted from the CRA–CRT landscape figure.

Methodology in Plain English

Building a molecular ground truth. The authors started with 52 Xenium spatial transcriptomics datasets from 10x Genomics that also have high-resolution H&E images spatially registered to the transcriptomic measurements. They assembled 14 scRNA-seq references to serve as annotation sources, prioritizing same-patient or tissue-matched references and otherwise matching by organ and disease state. Reference-specific labels were harmonized into 68 molecular cell types using a marker knowledge base curated from more than 300 publications covering primary, secondary, and negative markers, then transferred onto the Xenium sections with SpCAST. Every section was manually reviewed for expected marker patterns without applying a universal numerical threshold.

Quality control. Sections with fewer than two valid classes were excluded. The provider-supplied spatial alignment was inspected by overlaying Xenium cell coordinates on 20 sampled H&E regions per section, with no re-registration performed. Separately, CellViT++ predicted nuclear morphology from H&E alone; those predictions were matched to Xenium-derived labels by mutual nearest neighbors and mapped to a shared three-class space, and section-level balanced accuracy was used as a morphology–annotation concordance score purely for ranking. The 25 highest-scoring sections formed the primary panel S25. Robustness to panel size was checked using nested subsets S10 ⊂ S15 ⊂ S20 ⊂ S25. CellViT++ neither generated nor modified the labels and did not affect cell-level eligibility.

Two classification tasks. The 68 molecular cell types were deterministically mapped to a nine-class fine-grained taxonomy (T_fine) and a three-class coarse-grained taxonomy (T_coarse).

Coordinate-aligned feature extraction. At a given magnification, each WSI is split into non-overlapping patches matching the encoder's native input size, and each cell is assigned to the patch containing its registered nuclear center. The frozen encoder produces a spatial token grid (and a class token where the architecture supports one). Three base readouts are defined: Nuc, which bilinearly samples the feature grid at the registered nuclear center; Mean, which mean-pools the token grid for that patch; and Cls, the class token. Nuc uses the coordinate as a spatial anchor and does not imply a nucleus-restricted receptive field. Up to four fusions are also tested (Nuc+Mean, Nuc+Cls, Nuc|Mean, Nuc|Cls), created by direct addition or concatenation without normalization, rescaling, or projection — giving up to seven representation modes. For encoders without a native class token, the Cls readout and its derived fusions are omitted.

Four transfer protocols at ascending hierarchy. IS (intra-section) trains on cells inside a region of interest and tests on cells outside it. IOCD (intra-organ cross-dataset) holds out one complete tissue section from an eligible organ while the remaining sections of that organ supply training and validation. MOCV (multi-organ cross-validation) assigns complete sections to a fixed, organ-aware fold manifest and pools test predictions from all sections in the held-out fold. LOOO (leave-one-organ-out) excludes all sections of a target organ entirely, measuring zero-shot cross-organ transfer. Cells sharing a held-out spatial or anatomical unit are strictly excluded from probe fitting.

Probing and statistics. For each combination of model, condition, and evaluation unit, a fresh multiclass linear probe is fit on training representations, selected on validation, and evaluated once on the test partition; the encoder stays frozen throughout. A "run" is one fixed combination of model, magnification, representation mode, taxonomy, protocol, evaluation unit, and random seed, with seeds applied to data partitioning rather than just probe initialization. Macro-F1 is the primary metric and Macro-AUROC is secondary. Protocol summaries balance anatomical or cross-validation units rather than pooling all cells, so the reported standard deviations reflect variation across organs or folds, not across pooled cells. Model comparisons use two-sided paired Wilcoxon signed-rank tests over common valid evaluation units, with a significant win requiring p < 0.05 and a median paired difference favoring the model.

Scale of the experiment. Thirty models spanning pathology vision (PV), pathology vision–language (PVL), pathology vision–omics (PVO), and general-purpose vision (GV) were evaluated at 40x, 20x, and 10x under both taxonomies and the seven representation modes. This produced 157,500 runs for IS (25 sections, 5 seeds), 86,940 for IOCD (23 folds across 9 organs, 3 seeds), 18,900 for MOCV (5 folds, 3 seeds), and 41,580 for LOOO (11 folds across 11 organs, 3 seeds) — 304,920 in total. For models reporting only patch or image–text pair counts, a standardized conversion of 1,000 patches/pairs per WSI was applied to report WSI-equivalent pretraining counts.

Why This Matters

Impact on research. The benchmark reframes PFM evaluation around the frozen representation rather than the task pipeline. Because the cellular reference, readout definitions, probe configuration, and metrics are held constant, differences between models can be attributed to the representation itself. This gives model developers a diagnostic signal that region- or slide-level scores can obscure — for example, a model may score well on slide-level tasks by exploiting tissue composition even when local cellular variation is not strongly preserved. The paper also shows that decodability, sensitivity to spatial readout, and cross-domain generalization are related but not interchangeable properties, so a single aggregate ranking is an incomplete description of a model.

Real-world applications:

  • Cell-type and cell-state annotation in digital pathology pipelines, where the quality of a frozen backbone determines how much downstream fine-grained annotation is even possible.
  • Cross-site and cross-cohort model selection, since IOCD, MOCV, and LOOO probe how well cell-type information survives a change of section, dataset, or organ.
  • Spatial biology and tissue microenvironment studies, where nucleus-anchored, spatially indexed representations are the substrate for relating cells to their tissue context.
  • Backbone procurement decisions, where teams must choose among 30 models with different architectures, parameter counts, and pretraining paradigms under a fixed compute budget.

Industry relevance. The panel spans pathology-specific models from academic and industry groups as well as general-purpose vision models, and the benchmark is publicly released with a website for browsing results. It provides a standardized, reproducible way to audit a candidate encoder before committing to a large-scale annotation or clinical-research pipeline, and it points to specific design choices — such as preserving coordinates through feature sampling rather than relying on patch averaging — that matter at the representation level.

Future Directions

  • Explaining why some models retain more cell-level information. The paper documents substantial model-dependent differences in cell-type decodability and cross-domain generalization but, based on the content available, does not establish which pretraining data, objectives, or architectural choices cause them.

  • Motivating better aggregation. Since combining Nuc with Mean or Cls gave no consistent improvement across the seven representation modes, a natural next question is whether learned or otherwise designed readouts could combine nucleus-anchored and contextual information more effectively.

  • Closing the fine-grained gap. Fine-taxonomy Macro-F1 is far below coarse-taxonomy Macro-F1 under every protocol, and LOOO fine F1 is the lowest of all settings; improving discrimination among closely related cell types and generalizing it to unseen organs remain open problems.

  • Extending the reference. The panel covers 25 sections, 11 organs, and 7,079,283 cells drawn from Xenium data, and the authors already test robustness with the nested subsets S10 through S25 and defer taxonomy mappings, cell-eligibility rules, section-level scores, and configuration details to the Supplementary Material; broader organ coverage and other spatial platforms are natural extensions.

  • Connecting representation diagnostics to downstream outcomes. The paper argues that cell-level representation capability and higher-level downstream utility are related but not interchangeable; testing whether CRA and CRT predict performance on concrete clinical tasks is left open.

Target Audience

Researchers and engineers working on computational pathology, whole-slide image modeling, and foundation-model evaluation; developers of pathology foundation models who want diagnostics beyond downstream task scores; computational biologists using spatial transcriptomics as a cellular reference; and applied teams in digital pathology or pharma who need to select and audit a frozen encoder before building downstream analysis. Readers without a background in WSI processing, linear probing, or spatial transcriptomics will find the setup sections demanding, but the CRA–CRT framing and the plain-language interpretation of readouts are accessible.

Authors’ abstract

Pathology foundation models (PFMs) are increasingly used as general-purpose backbones, yet existing benchmarks cannot systematically diagnose their whole-slide cellular representation capabilities, including the decodability of cell-type information and the transferability of such information across tissue sections, datasets, and anatomical organs. We introduce CellPath-Bench, a cellular-resolution benchmark that evaluates frozen PFMs themselves. Following quality control of 52 candidate Xenium datasets, we construct a panel of 25 spatially aligned H\&amp;E--Xenium tissue sections spanning 11 organs and 7,079,283 cells, harmonized into fine- and coarse-grained taxonomies. CellPath-Bench samples frozen WSI feature maps at registered nuclear coordinates and evaluates them using standardized multiclass linear probes. Cell Representation Advantage (CRA) measures the within-section advantage of nucleus-anchored representations over patch-level mean pooling, while Cell Representation Transferability (CRT) characterizes the generalization of cell-type decodability across tissue sections, datasets, and organs. We benchmark 30 pathology-specific and general-purpose foundation models through 304,920 runs across spatial readouts, magnifications, taxonomic granularities, and evaluation protocols. The results reveal substantial model-dependent differences in cell-type decodability and its cross-domain generalization, yielding distinct multidimensional capability profiles. CellPath-Bench provides a standardized framework for auditing cellular information in frozen PFM representations.

Read the original paper