Research
INTERACT-CMIL: Multi-Task Shared Learning and Inter-Task Consistency for Conjunctival Melanocytic Intraepithelial Lesion Grading
Overview Research area: Computational digital pathology and computer vision for ophthalmology, specifically multi-task deep learning for grading Conjunctival Melanocytic Intraepithelial Lesions (CMIL)

- arXiv
- 2512.22666
- Published
- 2025-12-27
- Authors
- Mert Ikinci, Luna Toma, Karin U. Loeffler, Leticia Ussem, Daniela Süsskind, Julia M. Weller, Yousef Yeganeh, Martina C. Herwig-Carl, Shadi Albarqouni
AI summary
Overview
- Research area: Computational digital pathology and computer vision for ophthalmology, specifically multi-task deep learning for grading Conjunctival Melanocytic Intraepithelial Lesions (CMIL) from H&E-stained conjunctival biopsy images.
- Technical level: Advanced. The paper assumes familiarity with multi-task learning, foundation-model feature extractors, KL-divergence-style regularizers, and macro F1 evaluation under class imbalance.
- Scope: The paper presents INTERACT-CMIL, a five-head framework built on frozen CHIEF pathology embeddings that jointly predicts five CMIL grading criteria, trained with combinatorial partial supervision and an inter-dependence loss on a newly curated 486-crop multi-center dataset from three German university hospitals.
What This Paper Is About
CMIL are precursors of conjunctival melanoma, and grading them accurately determines treatment and melanoma risk. Current practice relies on subjective histopathological review that shows significant inter-observer variability, particularly for low-grade lesions. The paper's goal is to build an objective, reproducible, and interpretable model that predicts the five interrelated CMIL grading axes together rather than in isolation, so that predictions stay diagnostically coherent.
Key Contributions
- First deep learning framework designed specifically for CMIL classification, integrating the WHO 4th/5th edition grading paradigms with the C-MIN scoring system (the paper states CMIL is largely unexplored computationally due to rarity, limited data, and taxonomic complexity).
- A multi-head architecture with combinatorial partial supervision (SFCS — Shared Feature Learning with Combinatorial Partial Supervision): five task heads, of which three are selectively activated per training iteration, systematically cycling through all possible three-head combinations rather than randomly sampling.
- An inter-dependence loss (
L_dep) that aligns the model's predicted joint task distribution with the empirical co-occurrence statistics in the data, using a divergence term over all possible label combinations, added to the classification loss asL_total = L_cls + λ · L_dep. - A newly curated, expert-consensus, multi-center benchmark dataset of 486 annotated conjunctival biopsy crops from three German university hospitals (Tübingen, Bonn, Erlangen), described as the first multi-center benchmark for computational CMIL grading.
Main Findings
- Best performance on all five axes: INTERACT-CMIL achieves mean macro F1 of 0.7617 (WHO4), 0.8766 (WHO5), 0.7066 (horizontal spread), 0.7613 (vertical spread), and 0.7032 (atypia) over 5-fold cross-validation.
- Gains over the strongest baseline (BaseCHIEF), reported as relative improvements: +55.10% (WHO4), +24.71% (WHO5), +15.19% (horizontal spread), +24.99% (vertical spread), and +12.46% (atypia). The abstract cites gains up to 55.1% (WHO4) and 25.0% (vertical spread).
- Baseline comparison: BaseCHIEF (frozen CHIEF embeddings, single-task) reached WHO4 0.4911, WHO5 0.7029, HS 0.6134, VS 0.6091, AP 0.6253. BaseCNN (ResNet-18, ImageNet-pretrained, five heads jointly optimized) was weaker still: WHO4 0.4293, WHO5 0.5848, HS 0.3434, VS 0.2902, AP 0.3687.
- Pretrained pathology features matter substantially: using pathology-pretrained features (BaseCHIEF) over a ResNet-18 trained from scratch increased mean macro F1 by over 20% across tasks.
- Ablation — dependency loss: removing the inter-dependence loss produced the largest single drop, with a mean macro F1 decrease of approximately 2–3%, including −1.6 (atypia) and −0.7 (vertical spread) in absolute terms (e.g., atypia 0.7032 → 0.6869; vertical spread 0.7613 → 0.7542).
- Ablation — temperature scaling: disabling both dependency and temperature scaling yielded a further decline of approximately 5–6% across tasks (WHO5: 0.8613 → 0.8172; VS: 0.7542 → 0.7103; WHO4: 0.7397 → 0.6924).
- Ablation — selective multi-head supervision: removing it produced the most significant cumulative drop, averaging over 8% below the full model, with WHO4 falling from 0.7617 to 0.6120. The paper describes selective head updates as preventing gradient domination and improving balance between easier (HS) and harder (VS, AP) tasks.
- Component attribution: the paper attributes approximately 3% of the total improvement to dependency modeling, approximately 4% to temperature scaling, and approximately 4% to selective supervision, jointly explaining an aggregate gain of approximately 10% over the strongest baseline (note: this aggregate is smaller than the per-task relative gains reported above, which are stated against BaseCHIEF).
- Dataset characterization: Bonn contributed the largest share of crops (333), followed by Erlangen (135) and Tübingen (18). Label distributions show a predominance of higher-grade lesions (WHO5=2, WHO4=3), with C-MIN scores spanning 0–10 and peaking around 6–8. Vertical spread and atypia showed strong dependency, with 300 samples showing zero discrepancy.
- Class counts per task: macro F1 was computed over 3 classes for WHO5, 4 for WHO4, 4 for horizontal spread, 3 for vertical spread, and 3 for atypia; train/validation/test splits were patient-disjoint. ROC curves and AUC are presented in Figure 3 as secondary metrics.
- Training setup: implemented in PyTorch; training converged within 100 epochs on a single NVIDIA GPU. The value of λ, the batch size, and the specific temperature-scaling procedure are not reported in the provided text.
Methodology in Plain English
Each histopathological image patch is passed through a frozen CHIEF foundation-model encoder, producing a 768-dimensional feature vector. A set of shared layers compresses that into a 256-dimensional shared representation, intended to capture morphology relevant to every diagnostic criterion. Five separate classification heads then read from that shared representation: WHO4, WHO5, horizontal spread, vertical spread, and atypia.
The distinctive training trick is combinatorial partial supervision. Instead of updating all five heads on every iteration, only three are activated at a time, and the training loop systematically cycles through every possible three-head combination. The paper argues this regularizes the shared representation and prevents any single head from dominating the gradient signal.
The second trick is the inter-dependence loss. Within a batch, the model forms the predicted joint distribution over the three active tasks by taking outer products of each head's softmax probabilities, and forms the empirical joint distribution the same way from the one-hot ground-truth labels. A divergence term then pushes the predicted joint toward the empirical joint, so the model is penalized when its combination of head outputs looks statistically unlike what actually co-occurs in the data. The total training objective is the sum of the active heads' classification losses plus λ times this dependency term; temperature scaling is an additional component enumerated in the ablation study. Only the active heads and shared layers are updated each iteration.
Why This Matters
Impact on research. The paper positions CMIL as an underexplored computational pathology problem and provides what it describes as the first deep learning framework for it, along with the first multi-center expert-annotated benchmark. It also tests a general idea — explicitly modeling dependencies between correlated diagnostic labels — in a setting where those dependencies are clinically well-defined, which is transferable to other multi-criteria grading tasks.
Real-world applications.
- Clinical decision support for ophthalmic pathologists grading conjunctival biopsies, providing a second read on WHO4, WHO5, C-MIN spread, and atypia.
- Standardizing grading across institutions and reducing the inter-observer variability the paper identifies as a major problem, especially for low-grade lesions.
- Structured, multi-criteria outputs that map onto existing WHO and C-MIN reporting conventions rather than a single opaque score.
- Building multi-center digital ocular pathology infrastructure, since the dataset itself demonstrates that cross-site staining and acquisition heterogeneity can be pooled.
Industry relevance. The work matters to digital pathology platform vendors, computational pathology tool developers, and clinical AI groups interested in multi-task architectures for rare diseases with small, heterogeneous datasets. The finding that frozen pathology foundation-model features plus careful multi-head scheduling beat an end-to-end ImageNet CNN is directly relevant to anyone deciding whether to fine-tune or freeze large encoders on small clinical cohorts.
Future Directions
- Extend the dataset beyond 486 crops across three institutions, which the authors explicitly name as future work.
- Move to whole-slide representations rather than sampled patches, also named in the paper.
- Apply domain adaptation and self-supervised pretraining to improve robustness under clinical variability.
- Open questions the paper leaves unaddressed: the value of λ is not reported, so the sensitivity of the dependency loss to its weighting is unclear; the temperature-scaling procedure is enumerated in the ablation but not detailed in the methodology as presented; and the paper's own claim that horizontal spread is "near-saturated" sits alongside a reported macro F1 of 0.7066, leaving room to examine how much headroom remains on that axis. Additionally, an AUC analysis is deferred to a figure without numeric values in the text.
Target Audience
Ophthalmic pathologists and ocular oncologists interested in computational grading of conjunctival lesions; computer vision and machine learning researchers working on multi-task learning and small-cohort medical imaging; computational pathology groups evaluating foundation-model feature extractors against end-to-end CNNs; and clinical AI developers looking for a concrete, expert-annotated multi-criteria benchmark with patient-disjoint splits and reported cross-validation statistics.
Authors’ abstract
Accurate grading of Conjunctival Melanocytic Intraepithelial Lesions (CMIL) is essential for treatment and melanoma prediction but remains difficult due to subtle morphological cues and interrelated diagnostic criteria. We introduce INTERACT-CMIL, a multi-head deep learning framework that jointly predicts five histopathological axes; WHO4, WHO5, horizontal spread, vertical spread, and cytologic atypia, through Shared Feature Learning with Combinatorial Partial Supervision and an Inter-Dependence Loss enforcing cross-task consistency. Trained and evaluated on a newly curated, multi-center dataset of 486 expert-annotated conjunctival biopsy patches from three university hospitals, INTERACT-CMIL achieves consistent improvements over CNN and foundation-model (FM) baselines, with relative macro F1 gains up to 55.1% (WHO4) and 25.0% (vertical spread). The framework provides coherent, interpretable multi-criteria predictions aligned with expert grading, offering a reproducible computational benchmark for CMIL diagnosis and a step toward standardized digital ocular pathology.