Research
Reliable Mislabel Detection for Video Capsule Endoscopy Data
Overview Research area: Medical computer vision — dataset quality and label-noise handling for Video Capsule Endoscopy (VCE), with downstream anomaly detection. Technical level: Intermediate. The pape

- arXiv
- 2602.06938
- Published
- 2026-02-06
- Authors
- Julia Werner, Julius Oexle, Oliver Bause, Maxime Le Floch, Franz Brinkmann, Hannah Tolle, Jochen Hampe, Oliver Bringmann
AI summary
Overview
Research area: Medical computer vision — dataset quality and label-noise handling for Video Capsule Endoscopy (VCE), with downstream anomaly detection.
Technical level: Intermediate. The paper assumes familiarity with CNNs, focal loss, Gaussian Mixture Models, and anomaly detection metrics, though the pipeline itself is described step by step.
Scope: The paper proposes and clinically validates a two-stage pipeline that detects, corrects, and filters mislabeled frames in the two largest public VCE datasets, showing improved anomaly detection after cleaning.
What This Paper Is About
Deep models for medical imaging need large, accurately labeled datasets, but medical labels must come from scarce specialist physicians, and class boundaries can be genuinely ambiguous. This paper builds a framework that finds likely mislabeled images in VCE datasets, then either corrects the label or removes the sample before training. It validates the framework both on a dataset where the researchers injected known label noise and on a real dataset where the mislabels were pre-existing, with expert gastroenterologists reviewing the flagged samples.
Key Contributions
- A mislabel detection and cleaning pipeline combining repeated CNN training with three-component Gaussian Mixture Models on per-sample loss, followed by separate label-correction and sample-filtering steps.
- A controlled validation on the Kvasir-Capsule dataset (47,238 labeled and 4,694,266 unlabeled frames) using deliberately injected label noise at 1%, 5%, and 10%, where ground truth is known.
- Application of the pipeline to the Galar dataset (3,513,539 labeled images), filtering 167,709 samples (4.8%) and correcting 31,650 samples (0.9%), with released corrected dataset splits.
- A clinical review in which 100 pipeline-flagged samples were re-annotated by three experienced gastroenterologists (two of whom were involved in creating the original Galar dataset), yielding a Precision@100 score of 78.
Main Findings
- Controlled noise detection works well: On Kvasir-Capsule, 456 of 471 injected 1%-noise samples, 2262 of 2360 at 5%, and 4355 of 4722 at 10% were corrected and/or filtered. Only 15, 98, and 367 noisy samples respectively remained undetected, while 916, 975, and 991 non-noisy samples were filtered out.
- Loss separates clean from noisy samples: The GMM component with the lowest mean loss holds the largest, correctly labeled group; a flat third component with the highest loss values and strong outliers captures presumably noisy labels (shown for the 5% injection, correction step).
- Better separated latent clusters: t-SNE plus PCA (reducing to 50 dimensions) showed that corrections were concentrated on samples lying inside the opposite class cluster or in transitional regions between classes, producing more coherent clusters.
- Filtering the development set gives the largest gain: Training on the cleaned development set with an untouched test set reached 93.83% accuracy and 71.58% F1, versus 90.99% accuracy and 53.70% F1 for uncleaned data. Baselines reported 37.01% and 54.38% F1.
- Correcting alone helps, filtering helps more: Corrected development set (uncleaned test) gave 91.48% accuracy and 64.23% F1; the filtered development set gave the highest accuracy and F1 of the study.
- Filtering the test set boosts precision: The best precision, 89.88%, came from filtering both development and test sets, at 91.72% accuracy, 73.67% F1, and 68.05% sensitivity.
- Model confidence rises with cleaning: Maximum confidence went from 0.85 (uncleaned) to 0.89 (corrected) to 0.96 (filtered). Baseline confidence values are not reported in the table.
- Expert review confirms most flags: Of the 100 suspected samples, 78% were confirmed to have incorrect labels; 49% were originally labeled normal but confirmed as anomalies, and 29% were originally labeled anomalous but confirmed as normal. Only 1% of anomalies and 21% of normal data flagged by the pipeline were re-evaluated as correctly labeled.
Methodology in Plain English
The work runs in two stages. First, a controlled experiment: the researchers take Kvasir-Capsule, where true labels are known, and deliberately flip a small fraction of labels (1%, 5%, or 10%). Which samples were flipped is recorded, so the pipeline can be scored exactly. Noise was injected mainly by picking samples from mid- and high-uncertainty quantiles, ranked by a combined normalized average prediction confidence and entropy across epochs and three independent training runs, with flips distributed proportionally per class to preserve the original class distribution.
Second, the same pipeline runs on Galar to find pre-existing mislabels. The pipeline works like this: train a CNN three times on the raw data, collect the per-sample loss per epoch, and fit a three-component Gaussian Mixture Model with the sklearn implementation. The component with the lowest mean is treated as correctly labeled samples, the middle component as hard-to-learn samples, and the highest-mean component as the wrongly labeled distribution. Each sample gets a noise probability, and a noise-reduction score defined as the difference between its noise probability before and after a potential correction. The samples with the highest noise reduction are corrected first (labels are flipped in the binary case), then three more CNN-and-GMM rounds are run, and finally the samples with the highest noise probability are filtered out. This fuses a correction step with a filtering step, following the idea of an earlier work.
The classifier is MobileNetV3, chosen for low model complexity (about 1 M parameters) and suitability for embedded devices, trained with focal loss to counter class imbalance, using the HANNAH framework, the AdamW optimizer, learning rate and weight decay of 1×10⁻⁴, 15 epochs on Galar and 10 epochs on the smaller Kvasir-Capsule. Official Kvasir-Capsule splits were used so no patient appears in both development and test sets; Galar used splits generated as in prior work.
For the clinical check, samples were ranked by noise-reduction score, the top 500 considered, and 100 selected: 70 showing normal mucosa and 30 showing pathological findings, drawn from more than 50 distinct videos/patients with a maximum of three samples per video and at least 100 frames between samples to avoid near-duplicate frames. Two of the three reviewing gastroenterologists were among the original Galar dataset authors. A label was treated as wrong if at least two physicians agreed it was wrong.
Why This Matters
Data cleaning is usually invisible in reported results, yet this paper shows it can move F1 from 53.70% to 71.58% on Galar without changing the classification architecture. That reframes the bottleneck in medical VCE machine learning as a data-quality problem as much as a modeling problem — a useful reframing for a field where annotations are expensive and class boundaries are subjective.
Real-world applications:
- On-device anomaly detection in swallowable capsules: The long-term goal is real-time detection on the capsule itself, which requires reliable models trained on trustworthy labels.
- Clinical dataset curation: Hospitals and consortiums building VCE archives could use the pipeline to audit existing annotations before releasing data.
- Annotation quality assurance: Flagging likely mislabels directs limited gastroenterologist time to the samples that most need review, rather than re-checking everything.
- Gastrointestinal screening and diagnosis: Better-trained models support detection of angiectasia, polyps, and blood in the small intestine, where VCE is specifically applied.
Industry relevance: the released corrected splits and per-sample anomaly/healthy labels for the entire Galar dataset (167,709 filtered, 31,650 corrected) let other groups train on cleaned data directly. The emphasis on a compact MobileNetV3 of roughly 1 M parameters and the HANNAH framework also points toward embedded, low-power deployment rather than cloud-only inference. The work was partly funded by the German Federal Ministry of Research, Technology and Space (BMFTR) in the project MEDGE (16ME0530).
Future Directions
- Re-annotate all identified noisy samples with gastroenterologists rather than only the 100-sample subset, which the authors explicitly acknowledge as a small sample size to be increased.
- Validate the pipeline on other medical imaging datasets beyond VCE, since the framework is presented generally for medical datasets.
- Determine systematic rules for when a noisy sample should be corrected versus filtered, beyond the current ranking-based approach.
- Establish whether training on corrected test sets is methodologically preferable, since filtering the test set raised precision to 89.88% but lowered accuracy to 91.72% and F1 comparisons to baselines become less direct.
Target Audience
Medical imaging and clinical machine learning researchers dealing with label noise; dataset curators and annotation teams in gastroenterology; engineers building embedded or on-device diagnostic models for capsule endoscopy; and anyone working on noisy-label learning, confident learning, or dataset cleaning who wants a medically grounded case study with expert-verified validation.
Authors’ abstract
The classification performance of deep neural networks relies strongly on access to large, accurately annotated datasets. In medical imaging, however, obtaining such datasets is particularly challenging since annotations must be provided by specialized physicians, which severely limits the pool of annotators. Furthermore, class boundaries can often be ambiguous or difficult to define which further complicates machine learning-based classification. In this paper, we want to address this problem and introduce a framework for mislabel detection in medical datasets. This is validated on the two largest, publicly available datasets for Video Capsule Endoscopy, an important imaging procedure for examining the gastrointestinal tract based on a video stream of lowresolution images. In addition, potentially mislabeled samples identified by our pipeline were reviewed and re-annotated by three experienced gastroenterologists. Our results show that the proposed framework successfully detects incorrectly labeled data and results in an improved anomaly detection performance after cleaning the datasets compared to current baselines.