Skip to content
AI.info

Research

Data Efficiency and Transfer Robustness in Biomedical Image Segmentation: A Study of Redundancy and Forgetting with Cellpose

Overview Research area: Data-centric biomedical image segmentation; specifically training-data efficiency (coreset selection) and knowledge retention (catastrophic forgetting) in the generalist Cellpo

Data Efficiency and Transfer Robustness in Biomedical Image Segmentation: A Study of Redundancy and Forgetting with Cellpose
arXiv
2511.04803
Published
2025-11-06
Authors
Shuo Zhao, Jianxu Chen

AI summary

Overview

Research area: Data-centric biomedical image segmentation; specifically training-data efficiency (coreset selection) and knowledge retention (catastrophic forgetting) in the generalist Cellpose framework.

Technical level: Intermediate — the paper assumes familiarity with instance segmentation, transfer learning, and embedding-based dataset analysis, but its methods (patch quantization, finetuning, replay) are conceptually simple.

Scope: A diagnostic empirical study — not a new architecture — measuring how much of Cellpose's training data is actually needed and how badly its performance degrades when finetuned across biomedical imaging domains.

What This Paper Is About

Generalist biomedical segmentation models such as Cellpose are widely reused across labs, cell types, and imaging modalities, but two questions are rarely quantified: how much of the training data is redundant, and how much knowledge is lost when the model is adapted to a new domain. The authors use Cellpose as a case study to measure both, testing whether compact training subsets preserve accuracy and whether simple replay or domain ordering can prevent forgetting. The goal is a data-centric diagnosis of existing training and adaptation workflows, not a critique of Cellpose itself.

Key Contributions

  1. A dataset quantization (DQ) strategy using pretrained MAE features, bin partitioning, and a submodular gain criterion to build compact, diverse training subsets, showing that roughly 10% of the Cellpose (Cyto) training data reaches near-saturated segmentation performance.
  2. A demonstration that finetuning Cellpose from Cyto to MoNuSeg (Histo) causes catastrophic forgetting, with source-domain IoU collapsing from 0.771 ± 0.182 to 0.047 ± 0.083.
  3. Validation of both redundancy and forgetting phenomena under more challenging cross-domain conditions using the NeurIPS 2022 Cell Segmentation Challenge dataset (MultiInst).
  4. Analysis of zero-shot transfer, sequential multi-stage transfer paths, and a lightweight DQ-based replay baseline (5–10% of source data) that restores source performance while preserving target accuracy.

Main Findings

  • Redundancy is substantial. On the Cyto dataset, performance saturates early: precision moves only from 0.891 ± 0.113 at 1% of data to 0.903 ± 0.101 at 10%, dice rises from 0.823 ± 0.171 at 10% to 0.858 ± 0.144 at 40%, and at 100% dice is 0.855 ± 0.157 and IoU 0.771 ± 0.182 — essentially the same as at 40% (IoU 0.772 ± 0.175). The authors conclude that more than 60% of the training data is functionally redundant.
  • DQ selects more diverse patches, but metrics are comparable to random sampling. Visual inspection and t-SNE projections of MAE embeddings show DQ-selected patches covering a broader region of latent space than random samples. Numerically on Cyto, however, random sampling was slightly higher at matched rates (for example IoU 0.722 ± 0.015 vs DQ 0.669 ± 0.208 at 1%, and 0.771 ± 0.006 vs 0.743 ± 0.203 at 30%). On MultiInst, DQ led at 30% (IoU 0.418 ± 0.316 vs 0.373 ± 0.022) and 50% (0.407 ± 0.342 vs 0.381 ± 0.018). The authors position DQ's advantages as feature diversity, determinism, and reproducibility rather than metric superiority.
  • Zero-shot transfer is moderate. A model trained only on 30% of Cyto generalized to MultiInst at IoU 0.428 ± 0.337, dice 0.516 ± 0.365, recall 0.453 ± 0.354, accuracy 0.832 ± 0.121, and PQ 0.341 ± 0.319, versus the 100%-data model's IoU 0.419 ± 0.291, dice 0.530 ± 0.308, precision 0.810 ± 0.213, recall 0.440 ± 0.303, accuracy 0.820 ± 0.128, and PQ 0.308 ± 0.287. The DQ model showed better recall, accuracy, and PQ but higher variance.
  • Forgetting is severe and asymmetric. After Cyto → Histo finetuning, Cyto IoU fell to 0.047 ± 0.083, dice to 0.080 ± 0.129, and PQ to 0.014 ± 0.038, while Histo dice reached 0.817 ± 0.037. The reverse direction (Histo → Cyto) forgot far less: Cyto dice stayed at 0.856 ± 0.149 while Histo dice was 0.561 ± 0.151. Joint training on Cyto + Histo was balanced (Cyto dice 0.850 ± 0.153; Histo dice 0.814 ± 0.032).
  • Dataset compression alone does not prevent forgetting. DQ-trained models finetuned on full Histo all degraded, with Cyto IoU between 0.042 and 0.086 regardless of whether the subset was 1%, 10%, 30%, or 50%. Larger subsets gave no clear resilience.
  • Selective replay works; full replay backfires. Adding 5–10% of Cyto to full Histo training raised Cyto dice from 0.080 ± 0.129 (0% replay) to 0.813 ± 0.195 (5%) and 0.831 ± 0.164 (10%), with Histo accuracy stable near 0.80. Using 100% Cyto plus 100% Histo dropped Histo IoU to 0.310 ± 0.123 and dice to 0.459 ± 0.155.
  • Training order matters. In the three-stage paths, Path A (Cyto → Histo → MultiInst) gave the best MultiInst result (IoU 0.479) but forgot prior tasks (Cyto IoU 0.465, Histo IoU 0.232). Path B (Cyto → MultiInst → Histo) excelled on Histo (IoU 0.682) but lost Cyto (IoU 0.142). Path C (MultiInst → Cyto → Histo) the authors describe as the most balanced (Cyto IoU 0.160, Histo IoU 0.680, MultiInst IoU 0.295). Starting with the heterogeneous MultiInst domain produced more transferable representations, and simpler domains like Cyto were forgotten more easily when exposed early.

Methodology in Plain English

The authors ran a controlled empirical study on three public datasets, renamed for clarity: Cyto (the original Cellpose set, 540 annotated microscopy images), Histo (MoNuSeg, 37 H&E histopathology images), and MultiInst (the NeurIPS 2022 Cell Segmentation Dataset, multi-site and multi-modality). All images were patched with a 224 × 224 sliding window at stride 112, yielding 10,063 Cyto patches, 2,997 Histo patches, and 128,458 MultiInst patches.

To test redundancy, they extracted patch features with a pretrained Masked Autoencoder, partitioned the feature space into non-overlapping bins, and within each bin selected patches maximizing a submodular gain (Equation 1), then uniformly sampled a fixed proportion ρ from each bin (Equation 2) to form the coreset. They trained Cellpose at quantization rates from 1% to 100% and compared against random sampling across five runs, using IoU, dice, precision, recall, accuracy, and panoptic quality. Training used the official Cellpose interface with grayscale input, learning rate 0.1, weight decay 1 × 10⁻⁴, 500 epochs, and checkpoints every 50 epochs, on a single NVIDIA A100 GPU (40GB).

To test forgetting, they finetuned Cyto-trained models on Histo and vice versa with no retention mechanism, then repeated the experiment adding 1–50% of Cyto back as replay. They also ran three of the six possible three-domain orderings to see how sequence affects retention and final accuracy. Code is available at https://github.com/MMV-Lab/biomedseg-efficiency.

Why This Matters

Research impact. The paper reframes segmentation quality as a data-composition problem, not only an architecture problem. It quantifies two effects — redundancy and catastrophic forgetting — that are frequently assumed rather than measured in biomedical imaging, and it provides a reproducible baseline (DQ plus simple replay) that other continual-learning and coreset methods can be compared against.

Real-world applications:

  • Annotation budgeting in wet labs: if ~10% of a 540-image training set approaches saturated performance, teams can redirect the large majority of pixel-level labeling effort elsewhere.
  • Cross-institution model reuse: labs that finetune a shared generalist model on their own tissue type can see source performance collapse to IoU 0.047, which would make a deployed model unreliable after each adaptation round.
  • Cloud or edge training under compute limits: compact coresets plus 5–10% replay give retention benefits at a fraction of full-retraining cost.
  • Multi-modality pipeline design in consortia: the ordering results give concrete guidance (start with heterogeneous data such as MultiInst, place specialist data such as Histo last) for groups adapting one model across fluorescence, CODEX, and histopathology.

Industry relevance. Any organization selling or deploying biomedical imaging models faces the same trade-off: retraining on new customer data while preserving validated performance on previous domains. The replay prescription here (5–10% source data, not 100%) is cheap, requires no architectural change, and directly informs how product teams should schedule and retain finetuning runs.

Future Directions

  • Combining DQ-based replay with continual-learning machinery such as knowledge distillation or memory-based regularization, to move beyond the lightweight baseline.
  • Replacing the fixed DQ criterion with uncertainty- or diversity-aware sampling to improve coreset quality and reduce the higher variance observed under domain shift in DQ-trained models.
  • Modular or partially separated architectures that isolate domain-specific components and reduce interference between generalist and specialist knowledge.
  • Systematic curriculum research: the paper tests only three of the six possible domain orderings, and unstated combinations, longer domain sequences, and adaptive ordering rules remain open.

Target Audience

Researchers and engineers working on biomedical image segmentation, data-centric machine learning, and continual or transfer learning. It is also directly useful to practitioners who deploy Cellpose or similar generalist models across multiple labs or imaging modalities and need to know how much annotation and retention effort is actually required. Readers looking for a novel architecture will not find one here; the value is in the empirical measurements and the practical replay and ordering prescriptions.

Authors’ abstract

Generalist biomedical image segmentation models such as Cellpose are increasingly applied across diverse imaging modalities and cell types. However, two critical challenges remain underexplored: (1) the extent of training data redundancy and (2) the impact of cross domain transfer on model retention. In this study, we conduct a systematic empirical analysis of these challenges using Cellpose as a case study. First, to assess data redundancy, we propose a simple dataset quantization (DQ) strategy for constructing compact yet diverse training subsets. Experiments on the Cyto dataset show that image segmentation performance saturates with only 10% of the data, revealing substantial redundancy and potential for training with minimal annotations. Latent space analysis using MAE embeddings and t-SNE confirms that DQ selected patches capture greater feature diversity than random sampling. Second, to examine catastrophic forgetting, we perform cross domain finetuning experiments and observe significant degradation in source domain performance, particularly when adapting from generalist to specialist domains. We demonstrate that selective DQ based replay reintroducing just 5-10% of the source data effectively restores source performance, while full replay can hinder target adaptation. Additionally, we find that training domain sequencing improves generalization and reduces forgetting in multi stage transfer. Our findings highlight the importance of data centric design in biomedical image segmentation and suggest that efficient training requires not only compact subsets but also retention aware learning strategies and informed domain ordering. The code is available at https://github.com/MMV-Lab/biomedseg-efficiency.

Read the original paper