Research
QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation
QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation Overview Research area: Computer vision / biomedical image analysis — specifically instance segmentation of overlappi

- arXiv
- 2608.29253
- Published
- 2026-08-29
- Authors
- Yaroslav Prytula, Anton Popov, Dmytro Fishman
AI summary
QCell: Recombining and Aligning Cell Queries for Overlapping Instance SegmentationOverview
Research area: Computer vision / biomedical image analysis — specifically instance segmentation of overlapping cells in microscopy images, built on query-based (DETR-style) mask transformer architectures.
Technical level: Advanced. The paper assumes familiarity with MaskDINO, denoising (DN) training, Hungarian matching, contrastive learning (InfoNCE), amodal instance segmentation, and metrics such as AP, AP50, AP75, DICE, F1, and AJI.
Scope: The paper proposes a query-based model, QCell, that adds instance decomposition/recombination and contrastive query alignment on top of MaskDINO, and introduces a new Organoids benchmark for densely overlapping cell segmentation.
What This Paper Is About
When cells overlap in microscopy images, the overlap region contains blended visual evidence from multiple cells, producing weak boundaries and ambiguous ownership of pixels. Standard instance segmentation models output a single mask per instance and lack explicit mechanisms for reasoning about which parts of a cell are visible, hidden, or shared. QCell's goal is to "de-overlap" cells — recovering complete object structure under overlap while keeping the queries of neighboring instances distinguishable from one another.
Key Contributions
- Query-level instance decomposition and recombination with consistency regularization. Each instance query is decomposed into amodal (full extent), visible (non-overlapped), and invisible (occluded) sub-representations via three MLP heads, then recombined into a refined full-instance query. A consistency loss constrains the refined mask to be recoverable from the XOR of the thresholded visible and invisible predictions (with stop-gradient on the pseudo-target).
- DN-guided contrastive query learning. An instance-discriminative InfoNCE loss and a cosine alignment loss operate on projected query embeddings, using denoising (DN) queries as stable positive anchors for the same instance and DN queries of other instances as negatives.
- A new Organoids benchmark for overlapping cell instance segmentation in brightfield microscopy, containing 1,186 training images, 1,199 validation images, and 201 test images at 540×540 resolution, with up to 105 instances per training image and an average of 96 instances per test image (maximum 223). The dataset is available upon request.
- Empirical validation across three benchmarks showing best-overall performance on ISBI2014, Revvity-25, and Organoids, with +2.2 AP and +2.7 AJI over the MaskDINO baseline on ISBI2014.
Main Findings
- ISBI2014 cytoplasm segmentation: QCell achieves 65.9 AP, 91.9 AP50, 69.1 AP75, 92.4 DICE, 92.3 F1, and 78.6 AJI. Relative to the MaskDINO baseline (63.7 AP, 89.0 AP50, 65.4 AP75, 92.3 DICE, 90.0 F1, 75.9 AJI), this is +2.2 AP, +2.3 F1, and +2.7 AJI.
- Revvity-25: QCell reaches 52.9 AP, 86.0 AP50, 59.3 AP75, 89.4 DICE, 86.4 F1, and 73.6 AJI — the best AP and AJI among compared methods. The MaskDINO baseline reaches 52.3 AP and 73.5 AJI there.
- Organoids: QCell obtains 51.0 AP, 69.1 AP50, 56.4 AP75, 92.5 DICE, 71.6 F1, and 63.2 AJI, versus MaskDINO at 49.7 AP, 68.7 AP50, 54.7 AP75, 92.2 DICE, 71.5 F1, 63.1 AJI.
- Ablation — instance recombination: Adding the instance recombination loss raises the baseline from 63.7 to 65.1 AP and from 65.4 to 68.1 AP75. Adding consistency regularization further lifts performance to 66.6 AP, 69.2 AP75, and 77.1 AJI.
- Ablation — contrastive losses: The instance-discriminative loss alone reaches 65.5 AP; the cosine alignment loss alone reaches 66.2 AP; combining both reaches 67.0 AP and 77.9 AJI, indicating the two are complementary.
- DN oracle analysis: Replacing matched main-query predictions with their corresponding DN-query predictions at test time improves results from 63.7 to 68.6 AP (±1.2) and from 75.9 to 80.0 AJI (±0.6), motivating DN queries as contrastive anchors.
- Performance under heavy overlap (ISBI2014 subset with pairwise ground-truth IoU ≥ 0.5): The baseline reaches 11.67 AP; adding recombination reaches 12.29 AP; contrastive learning alone reaches 12.63 AP; combining both reaches 13.68 AP and 10.97 AP75 — a +2.01 AP and +3.33 AP75 gain over the baseline.
- Object-level error analysis: QCell achieves the lowest object-based false-negative rate (FNo) and highest pixel-based true-positive rate (TPp) across all three datasets. The paper reports that DICE stays comparable to MaskDINO across benchmarks, suggesting gains come primarily from improved de-overlapping rather than boundary refinement.
- Cost: QCell uses 45M parameters and 182G FLOPs, versus MaskDINO at 44M parameters and 163G FLOPs.
- Mask2Former query-count sensitivity on Organoids: Mask2Former scored 34.7 AP with N=100 queries and 31.3 AP with N=300 queries; the authors note that increasing query count degraded Mask2Former performance in their experiments.
- PCTrans caveat: The authors trained and report PCTrans using its provided setup and observe limited generalization to overlapping amodal cell segmentation. Its AP stays low because the model does not produce per-instance confidence scores, so "AP has no meaningful ranking."
Methodology in Plain English
QCell starts from MaskDINO, a query-based segmentation framework in which a set of learned queries attends to multi-scale image features and each query produces a class score, a box, and a mask via dot product with pixel features.
Two ideas are layered on top:
1. Break each query into parts and put it back together. Instead of one query representation per instance, three small MLP heads produce an amodal sub-query, a visible sub-query, and an invisible (occluded) sub-query. Each generates its own mask. The three sub-queries are then fused by a learned projection into a refined query, whose mask is supervised against the full amodal ground truth. Because the fusion must recover the complete object, the sub-queries are pushed to capture complementary information. A consistency term adds the constraint that the refined mask should match the XOR of the thresholded visible and invisible masks — with gradients flowing only through the refined prediction (the XOR target is a fixed pseudo-target via stop-gradient).
2. Make queries of different overlapping cells look different. During training, MaskDINO's denoising queries are noisy versions of ground-truth boxes and labels, so every ground-truth instance is guaranteed several DN representations regardless of how well matching went. QCell projects both matched content queries and DN queries into a shared space. For each matched query, DN queries of the same instance are positives and DN queries of all other instances are negatives. An InfoNCE-style loss encourages discriminative features, and a cosine alignment loss pulls positives toward cosine similarity 1 while pushing negatives toward orthogonality. Both losses are applied across decoder layers.
The total objective adds these terms to the standard MaskDINO loss with weights λ=2.0 for the discriminative loss and λ=5.0 for the alignment loss, temperature τ=0.1, and the decomposition/refinement BCE and Dice coefficients set to 5.0 with a consistency weight of 1.0.
Experimental setup: All experiments use Detectron2 implementations with a ResNet-50-FPN ImageNet-pretrained backbone. R-CNN-based models use SGD with momentum 0.9 and an initial learning rate of 10⁻³ with 1k iterations of linear warm-up. Query-based models use AdamW with an initial learning rate of 10⁻⁴, weight decay 0.05, and a backbone learning-rate multiplier of 0.1, trained for 60k iterations with batch size 2, dropping the learning rate by a factor of 0.1 at 50k and 55k iterations. Query-based models use 100 queries on ISBI2014 and Revvity-25 and 300 on Organoids. Checkpoints are selected on validation performance and test results are averaged over three random seeds. All experiments ran on a single NVIDIA H200 Tensor Core GPU with 141 GB of HBM3e memory.
Datasets used:
- ISBI2014: 16 real extended depth-of-focus cervical cytology images and 945 synthetic images at 512×512 with nuclei and cytoplasm annotations; the paper follows the challenge split of 45 synthetic images for training, 90 for validation, and 810 for testing, benchmarking on cytoplasm annotations only.
- Revvity-25: 110 high-resolution 1080×1080 brightfield images, averaging 27 manually labeled, expert-validated cancer cells each, totaling 2,937 annotated instances.
- Organoids: the newly introduced benchmark described above.
Why This Matters
Overlapping and semi-transparent cells are the norm, not the exception, in dense cultures and cytology specimens. Because single-mask models entangle neighboring cells, downstream measurements of cell morphology, spatial organization, and population-level behavior can be corrupted. QCell targets that failure mode directly at the query level rather than relying on local regions of interest or hand-designed shape priors — which the paper argues are problematic for cells with extreme morphological diversity.
Potential real-world applications (drawn from the domains and datasets in the paper):
- Cervical cytology screening, where the ISBI2014 benchmark originates and where cell overlap is the central segmentation challenge.
- Cancer cell analysis in brightfield microscopy, using the Revvity-25 dataset's expert-validated cell borders and overlap annotations.
- Organoid research and drug development, where dense, highly overlapping 3D-derived cultures must be separated into individual instances (up to 223 instances per image in the new benchmark).
- General microscopy pipelines across brightfield, phase-contrast, and fluorescence modalities that feed downstream morphometric measurements.
Industry relevance: The author affiliations include the University of Tartu, Ukrainian Catholic University, Igor Sikorsky Kyiv Polytechnic Institute, and two companies (STACC OÜ and Better Medicine OÜ), indicating direct interest from the biomedical imaging and digital pathology sector. The modest compute overhead — 45M parameters and 182G FLOPs versus 44M/163G for MaskDINO — matters for practical deployment.
Future Directions
- Making pseudo-masks more accurate. The paper notes that because hidden regions must be inferred, reconstructed masks can have moderate IoU, which can limit AP75 and threshold-averaged AP even as more cells are detected overall. Improving the precision of the consistency pseudo-target is an open problem.
- Generalizing to more modalities and annotation types. QCell was benchmarked on cytoplasm annotations on ISBI2014 and on brightfield data for Revvity-25 and Organoids; whether it transfers to fluorescence, phase-contrast, or nuclei-specific settings is not established in the reported content.
- Handling the query-count trade-off. The authors observed that increasing object queries from 100 to 300 degraded Mask2Former on Organoids and chose 300 queries for QCell there — understanding how many queries a dense scene really needs remains unresolved.
- Extending to models with different mask representations. The paper reports that PCTrans's categorical mask formulation (one label per pixel, no per-instance confidence) does not fit the overlapping amodal task and yields unrankable AP. Adapting amodal de-overlapping to such representations is left open.
- Wider availability of the Organoids benchmark. Additional dataset details are deferred to supplementary material, and the dataset itself is available only upon request — broader access would help independent comparison.
Target Audience
Researchers and practitioners working on instance segmentation, amodal/occlusion-aware segmentation, and biomedical microscopy analysis. It is most useful to readers already comfortable with DETR-family architectures (DINO, MaskDINO, Mask2Former) and contrastive representation learning. Those focused on digital pathology, cytology, organoid imaging, or cell-tracking pipelines would benefit from the benchmark and the de-overlapping objective, while the two proposed modules are general enough to interest anyone studying instance separation in dense, visually entangled scenes.
Note: the supplied paper content is truncated partway through Table 6 (object-level error analysis), so the full FNo/TPp figures for QCell and the remaining baselines are not available in the content reviewed here.
Authors’ abstract
Instance segmentation of overlapping cells in microscopy remains challenging due to semi-transparent structures that produce weak boundaries and mixed visual evidence in overlap regions. Existing methods address this through local regions of interest or shape priors but lack global reasoning across overlapping objects. We present QCell, a novel query-based model that de-overlaps cell instances in microscopy scenes. Our approach combines (i) an instance recombination module that decomposes and recombines query representations in latent space, enabling the model to reason about complete object structure under overlap, and (ii) a contrastive query alignment objective that combines distinctive instance feature learning and separation of overlapping cell queries. We additionally introduce a new Organoid dataset benchmark for overlapping cell segmentation. We show that QCell outperforms state-of-the-art methods across multiple benchmarks, achieving +2.2 AP and +2.7 AJI on ISBI2014. Code is available at https://github.com/SlavkoPrytula/QCell