Research
Semi-Supervised Multi-Task Learning for Interpretable Quality As- sessment of Fundus Images
Overview Research area: Medical computer vision — automated retinal image quality assessment (RIQA), combining multi-task learning with semi-supervised pseudo-labeling for fundus (retina) photographs.

- arXiv
- 2511.13353
- Published
- 2025-11-17
- Authors
- Lucas Gabriel Telesco, Danila Nejamkin, Estefanía Mata, Francisco Filizzola, Kevin Wignall, Lucía Franco Troilo, María de los Angeles Cenoz, Melissa Thompson, Mercedes Leguía, Ignacio Larrabide, José Ignacio Orlando
AI summary
Overview
Research area: Medical computer vision — automated retinal image quality assessment (RIQA), combining multi-task learning with semi-supervised pseudo-labeling for fundus (retina) photographs. Published in Computers in Biology and Medicine; arXiv:2511.13353v1 [cs.CV], 17 Nov 2025, under a CC BY-NC-ND 4.0 license. The work is a collaboration between the Pladema Institute / Yatiris Group (UNCPBA-CONICET, Tandil, Argentina) and the ophthalmology service of Hospital de Alta Complejidad en Red "El Cruce" Dr. Néstor Carlos Kirchner (Florencio Varela, Argentina).
Technical level: Intermediate. The core ideas (transfer learning, teacher-student pseudo-labeling, multi-task heads, Grad-CAM) are standard CNN concepts, but interpreting the experimental design requires familiarity with classification metrics and dataset splits.
Scope: The paper proposes and evaluates a semi-supervised, multi-task deep learning framework that predicts overall fundus image quality together with the specific acquisition defects (illumination, clarity, contrast) that caused a bad capture, using a ResNet-18 backbone and pseudo-labels instead of new expert annotations.
What This Paper Is About
Automated screening systems for eye disease depend on the quality of the retinal images they receive, but most existing quality-assessment tools output only a single overall grade (e.g., good/bad) and say nothing about why an image is unusable. Telling a technician that a photo failed without telling them whether the problem was poor illumination, defocus, or low contrast is not actionable. The problem is that detailed defect labels are expensive, so they exist only in small datasets, which limits both model performance and generalization.
The paper's goal is to obtain interpretable RIQA models — ones that report overall quality and the specific imaging conditions to fix — without paying the full cost of expert labeling. It does this by combining manual overall-quality labels with pseudo-labels for quality details generated by a Teacher model trained on a small annotated set.
Key Contributions
-
A semi-supervised multi-task RIQA scheme. A Teacher model trained on a small expert-annotated set (MSHF) produces pseudo-labels for image quality details; these are used to fine-tune a network that was pre-trained only for overall quality assessment, adding an auxiliary multi-label branch.
-
Evidence that pseudo-label noise is comparable to inter-observer variability. Under this regime the multi-task model improves the primary task, statistically matches expert performance on the auxiliary task, requires far fewer manual labels, and offers greater interpretability than its single-task counterpart.
-
An analysis of how multi-task learning improves the primary task and produces more informative Grad-CAMs, aligning explanations better with the actual image regions that drove a prediction.
-
A public release of new expert annotations of capture conditions for a subset of EyeQ images (referred to as EyeQ-D), to support future research. The labels are released at a GitHub repository listed in the paper.
Main Findings
-
Multi-task beats single-task on overall quality. Using the same ResNet-18 backbone and identical training conditions, MT-EyeQ reached F1 0.875, precision 0.877, recall 0.874 and accuracy 0.891, versus 0.863 / 0.866 / 0.863 / 0.882 for the single-task ST-EyeQ baseline. Every EyeQ improvement was statistically significant (p<0.05).
-
The same trend holds on a second dataset. On DeepDRiD, MT-DeepDRiD reached F1 0.778, precision 0.721, recall 0.845, accuracy 0.735 versus 0.763 / 0.745 / 0.782 / 0.732 for ST-DeepDRiD. The F1 gain was significant (p<0.05); the other metrics increased without statistical significance, except precision, where the single-task model was slightly higher.
-
The approach is competitive with published state-of-the-art EyeQ results. The paper compares against a long list of prior methods, including Leonardo et al. 2022 (F1 0.878, Pr 0.879, Re 0.878, Acc 0.894), Xu et al. 2023 (0.872 / 0.876 / 0.871 / 0.889), QuickQual (0.867 / 0.877 / 0.861 / 0.886), König et al. 2024 (0.830 / 0.810 / 0.850 / 0.910) and Guo et al. 2024 (0.846 / 0.867 / 0.827 / 0.866). MT-EyeQ's F1 of 0.875 is higher than all but one of these on F1, and it is the only listed method combining semi-supervised training with both an additional task and CAM-based interpretability.
-
It outperforms the strongest comparable multi-label competitor. König et al. 2024, the closest prior approach (predicting overall quality plus multiple acquisition details in a multilabel setup), yields lower F1, precision and recall than both ST-EyeQ and MT-EyeQ.
-
It outperforms a re-trained lightweight baseline on DeepDRiD. QuickQual, re-trained on DeepDRiD images, obtained F1 0.742, Pr 0.698, Re 0.790, Acc 0.698 — lower than MT-DeepDRiD on all four metrics.
-
Quality-detail predictions are statistically comparable to the Teacher. The multi-task model achieved performance statistically comparable to the Teacher for most detail prediction tasks (p>0.05).
-
The model performs similarly to experts on the new EyeQ-D subset. In a newly annotated EyeQ subset of 160 test images labeled by eight experienced ophthalmologists (labels assigned by majority vote), the model's performance was similar to the experts', suggesting pseudo-label noise aligns with expert variability.
-
Per-class analysis favors the multi-task model. On EyeQ, gains were consistent across classes. The multi-task model notably improved results for the "Usable" class and achieved the highest accuracy for the "Reject" class, while QuickQual was highest on the "Good" class. The largest improvement was for the "Usable" category, where F1 increased by more than 2 points (the sentence is truncated in the supplied text). On DeepDRiD, the model improved classification for both the "Good" and "Bad" classes compared with re-trained QuickQual.
-
Binary reformulation of the good-quality task. Yi et al., modeling RIQA as a binary task merging "usable" and "reject" into a bad-quality class, reported F1 0.895; Guo et al. reported 0.926. MT-EyeQ achieved 0.940 under this setting.
Methodology in Plain English
The pipeline has four steps, illustrated in Figure 2 of the paper.
-
Train a Teacher on a small, richly annotated set. Task A is multilabel prediction of three binary quality details: illumination, clarity and contrast, each labeled good (1) or bad (0). The Teacher is a ResNet-18 fine-tuned on MSHF (802 retinal images) with a sigmoid multi-label head, minimizing binary cross-entropy. No data augmentation was used, consistent with common practice in Noisy Student training.
-
Generate pseudo-labels on a larger, differently labeled set. The Teacher is applied to EyeQ or DeepDRiD — datasets where overall quality is manually annotated but detail labels are absent — producing a label-augmented set containing manual overall-quality labels plus predicted detail labels.
-
Pre-train a single-task model for overall quality. A ResNet-18 is trained with a cross-entropy loss to classify overall quality: three classes on EyeQ (rejectable, usable, good) and two on DeepDRiD (poor, good).
-
Add the auxiliary branch and fine-tune jointly. A new multi-label head for the detail task is randomly initialized and all parameters are fine-tuned with a weighted sum of the binary cross-entropy on pseudo-labels and the cross-entropy on manual overall-quality labels, with weights λ_A and λ_B. Data augmentation follows a RandAugment-inspired strategy with rotations, horizontal and vertical flips, and small color perturbations; each training image underwent at most seven transformations.
All models used ImageNet-pretrained ResNet-18 backbones in PyTorch 1.11, trained for up to 115 epochs with SGD (momentum 0.9, initial learning rate 0.01, halved at epochs 30, 60 and 80). Model selection used F1-score for overall quality on the validation set, and λ_A and λ_B were tuned with the same criterion. Deeper or more complex backbones such as Vision Transformers were deliberately excluded because of overfitting risk on the limited training sets. Single-task baselines (ST-EyeQ, ST-DeepDRiD) were the pre-fine-tuning models; to control for training duration, their training was extended by an additional 115 epochs, but the best-performing single-task models were those obtained before this extension. Evaluation used F1, precision, recall and accuracy, with macro F1 for multiclass EyeQ and the bad-quality class as positive for binary DeepDRiD. Statistical significance was assessed with one- or two-tailed Wilcoxon signed-rank tests and bootstrap resampling.
Why This Matters
Impact on research. The paper addresses a practical bottleneck in medical imaging: multi-task models promise more interpretable outputs, but each added task normally adds an annotation burden. Demonstrating that pseudo-labels for auxiliary tasks can substitute for manual labels — and that their noise is on the order of inter-observer variability — provides a template applicable beyond fundus imaging. The released EyeQ-D annotations give the community a new benchmark for detail-level quality labels, and the analysis linking multi-task learning to more faithful Grad-CAMs connects model architecture choices to explainability quality.
Real-world applications.
- Point-of-care screening. A technician capturing retinal images in a diabetic retinopathy screening program can receive immediate feedback that an image failed because of illumination and clarity, rather than being told only that it was rejected.
- Immediate image recapture. Because feedback is available at acquisition, the patient does not need to be recalled for a second visit — the paper frames this explicitly as enabling correction of capture settings in real time.
- Large-scale screening program triage. Programs that have accumulated very large image collections (EyeQ alone contains 28792 images captured with more than 40 cameras) can apply such a model without commissioning detail-level annotation for the entire corpus.
- Training and quality assurance. Detail-level defect outputs plus Grad-CAM heatmaps can support technician training and auditing of capture practices across cameras and sites.
Industry relevance. The method is deliberately lightweight: a ResNet-18 backbone, ImageNet-pretrained weights, PyTorch 1.11, and only a handful of augmentation operations. That makes it deployable on modest hardware alongside existing screening platforms. The value proposition for commercial retinal screening and tele-ophthalmology products is reduced annotation cost, higher throughput through fewer unusable images, and interpretable outputs that clinicians can verify rather than trust blindly.
Future Directions
-
Extension to more quality details and other imaging modalities. The framework was demonstrated with k=3 details (illumination, clarity, contrast); the same pseudo-labeling strategy could in principle cover additional acquisition defects or be transferred to other retinal imaging modalities, though this is not tested.
-
Rigorous validation of the pseudo-label-noise hypothesis. The paper argues that pseudo-label noise aligns with expert variability, based on comparison with eight ophthalmologists on 160 EyeQ-D images. A broader multi-center study of inter-observer variability would test how far that claim generalizes.
-
Closing the gap with the strongest prior methods. On F1, Leonardo et al. 2022 (0.878) remains marginally higher than MT-EyeQ (0.875), and the paper notes that the best single-task checkpoints were reached before the training-time extension, suggesting headroom in training schedule and hyperparameter tuning.
-
Backbone and augmentation trade-offs. The authors excluded Transformers and deeper backbones to avoid overfitting on limited training sets; whether larger-scale data or stronger regularization could let more expressive architectures improve results while preserving interpretability remains open.
-
Independent replication on the released EyeQ-D labels. The value of the public release depends on other groups adopting it as a common benchmark, which the paper anticipates but cannot yet demonstrate.
Target Audience
This paper is most useful to researchers and engineers working on medical image analysis, particularly automated retinal screening pipelines and image quality assessment, who need interpretable model outputs without commissioning new expert annotations. It will also interest clinical informaticists and ophthalmology technologists concerned with capture quality at the point of care, and machine learning practitioners looking for a concrete case study of multi-task learning combined with teacher-student pseudo-labeling. Readers should be comfortable with standard CNN training terminology and classification metrics; the paper's distinctive contribution — that pseudo-label noise from a small Teacher can stand in for expert detail labels — is best appreciated by those already familiar with semi-supervised learning methods such as Noisy Student.
Authors’ abstract
Retinal image quality assessment (RIQA) supports computer-aided diagnosis of eye diseases. However, most tools classify only overall image quality, without indicating acquisition defects to guide recapture. This gap is mainly due to the high cost of detailed annotations. In this paper, we aim to mitigate this limitation by introducing a hybrid semi-supervised learning approach that combines manual labels for overall quality with pseudo-labels of quality details within a multi-task framework. Our objective is to obtain more interpretable RIQA models without requiring extensive manual labeling. Pseudo-labels are generated by a Teacher model trained on a small dataset and then used to fine-tune a pre-trained model in a multi-task setting. Using a ResNet-18 backbone, we show that these weak annotations improve quality assessment over single-task baselines (F1: 0.875 vs. 0.863 on EyeQ, and 0.778 vs. 0.763 on DeepDRiD), matching or surpassing existing methods. The multi-task model achieved performance statistically comparable to the Teacher for most detail prediction tasks (p > 0.05). In a newly annotated EyeQ subset released with this paper, our model performed similarly to experts, suggesting that pseudo-label noise aligns with expert variability. Our main finding is that the proposed semi-supervised approach not only improves overall quality assessment but also provides interpretable feedback on capture conditions (illumination, clarity, contrast). This enhances interpretability at no extra manual labeling cost and offers clinically actionable outputs to guide image recapture.