Research
Cataract-LMM Large-Scale Multi-Source Multi-Task Benchmark for Deep Learning in Surgical Video Analysis
Cataract-LMM: A Multi-Source, Multi-Task Benchmark for Cataract Surgery Video Analysis Overview Research area: Computer vision for computer-assisted surgery (CAS), specifically surgical video analysis
- arXiv
- 2510.16371
- Published
- 2025-10-18
- Authors
- Mohammad Javad Ahmadi, Iman Gandomi, Parisa Abdi, Seyed-Farzad Mohammadi, Amirhossein Taslimi, Mehdi Khodaparast, Hassan Hashemi, Mahdi Tavakoli, Hamid D. Taghirad
AI summary
Cataract-LMM: A Multi-Source, Multi-Task Benchmark for Cataract Surgery Video AnalysisOverview
Research area: Computer vision for computer-assisted surgery (CAS), specifically surgical video analysis of phacoemulsification cataract surgery — workflow recognition, surgical scene segmentation, instrument–tissue interaction tracking, and automated skill assessment.
Technical level: Advanced. The paper assumes familiarity with instance segmentation architectures (Mask R-CNN, YOLO), video action-recognition backbones (SlowFast, X3D, MViT, Video Swin), temporal action-segmentation models (MS-TCN/TeCNO, ASFormer), and domain-adaptation evaluation protocols.
Scope in one sentence: The paper introduces and technically validates Cataract-LMM, a 3,000-video, two-center cataract surgery dataset with four annotation layers (surgical phase labels, instance segmentation masks, instrument–tissue interaction tracking, and rubric-based skill scores), together with benchmark and cross-center domain-adaptation baselines for four deep-learning tasks.
What This Paper Is About
Deep-learning systems for computer-assisted surgery need large, richly annotated surgical video datasets that reflect real clinical variability — different hospitals, different surgeons, different skill levels, different equipment. Existing cataract surgery benchmarks, such as CATARACTS, CaDIS, CatRel, SICS-105, Sankara-MSICS, and Cataract-1K, are each drawn from a single clinical center and generally cover only one or two annotation types. The authors' goal is to close that gap by releasing a multi-center dataset of 3,000 phacoemulsification procedures with four complementary annotation layers, and by showing through benchmark experiments that the data can support workflow recognition, scene segmentation, interaction tracking, and automated skill scoring — including evaluation of how models transfer from one surgical center to another.
Key Contributions
-
A large multi-source surgical video corpus. 3,000 phacoemulsification cataract surgery videos recorded between December 2021 and March 2025 at two centers — Farabi Eye Hospital (2,930 videos) and Noor Eye Hospital (70 videos) — totaling 1,134.2 hours and capturing surgeons across experience levels (residents, fellows, expert attendings) and two different microscope-mounted camera systems.
-
Four complementary annotation layers on task-specific subsets. Temporal phase labels for 13 surgical phases across 150 videos; instance segmentation masks for 12 classes (10 instruments plus Pupil and Cornea) on 6,094 frames from 150 videos; spatiotemporal instrument–tissue interaction tracking for 170 capsulorhexis clips (469,118 frames); and quantitative skill scores for the same 170 clips using a 6-indicator, 5-point rubric adapted from GRASIS and ICO-OSCAR.
-
A paired skill–kinematic resource. The tracking subset and the skill-assessment subset use the same 170 capsulorhexis clips, which links dense instrument motion data (persistent instance IDs, functional keypoints, trajectories) directly to independently adjudicated expert skill ratings and adverse-event flags.
-
Benchmark and domain-adaptation baselines for four tasks. The paper specifies baseline protocols for phase recognition (clip-level and video-level), instance segmentation (supervised and zero-shot), object tracking validation, and skill classification, including an explicit cross-center setup where models train on Farabi data and are tested on all 21 Noor videos as an out-of-distribution set.
Main Findings
-
Dataset scale and diversity exceed prior cataract benchmarks. Table 1 compares Cataract-LMM (3,000 cases, multi-source, 720×480 and 1920×1080 resolution, 30 and 60 fps, 2021–2025) against CATARACTS (50 cases), CaDIS (25), CatRel (22), SICS-105 (105), Sankara-MSICS (53), and Cataract-1K (1,000) — all listed as single-source. Cataract-LMM is also the only dataset in that table listed with tracking and skill-assessment annotations.
-
Annotation reliability was measured before full annotation. Phase labels were produced by three ophthalmology residents (Years 2–4) who achieved a Global Fleiss' Kappa of κ = 0.924 on a 13-video held-out validation set (after a 5-video pilot set). Instance segmentation annotators (eight ophthalmology residents, Years 2–4) achieved a mean Intersection over Union of 0.874 and semantic class accuracy of 0.992 on a blind validation subset of n = 250.
-
Tracking annotations showed high inter-rater reliability. On a stratified validation subset of 10 video clips, a blinded inter-rater reliability study yielded Association Accuracy (AssA) of 98.4, Detection Accuracy (DetA) of 82.7, and Higher Order Tracking Accuracy (HOTA) of 90.2.
-
Skill ratings showed excellent agreement. Three independent blind raters produced an overall Intraclass Correlation Coefficient (ICC) of 0.87 before adjudication. Final scores came from double-blind tri-rating, supervisor adjudication when raters disagreed by more than one point on the 5-point scale, and arithmetic averaging within tolerance (≤ 1 point).
-
A data-driven binary skill split was defined. K-Means clustering (K = 2) on the continuous overall scores of the 170 clips produced a Lower-Skilled group (n = 63, mean score = 3.12 ± 0.38) and a Higher-Skilled group (n = 107, mean score = 4.24 ± 0.37).
-
Cross-center evaluation was built into the phase-recognition protocol. Training (80 videos) and validation (26 videos) came exclusively from Farabi Hospital; the 44-video test set combined 23 unseen Farabi videos (in-distribution) and all 21 Noor videos (out-of-distribution).
-
Instance segmentation was benchmarked at three granularities. Task 1 merges all 10 instruments into a single "Instrument" class (3 classes), Task 2 merges only the most similar instruments (9 classes), and Task 3 treats all 12 classes as distinct. Models are evaluated with mask mAP over IoU thresholds 0.50 to 0.95 following the COCO protocol.
-
Position-based interaction analysis was specified for the tracking data. Instrument tip coordinates are expressed relative to the pupil centroid (Δx, Δy), fitted with a two-dimensional Gaussian Kernel Density Estimator separately for Lower-Skilled and Higher-Skilled cohorts, and tip-to-pupil Euclidean distances (d = √(Δx² + Δy²)) are summarized with a normalized histogram and kernel density curve annotated with mean and median.
-
Quantitative benchmark results are not reported in the provided content. The paper states that score distribution and rubric construct validity are characterized in a "Technical Validation" section, and that baselines were run for all four tasks, but the truncated text supplied here ends before that section. The actual accuracy, F1, mAP, HOTA comparison, and skill-classification performance numbers for the baselines are therefore not available in this content.
Methodology in Plain English
Building the corpus. The team prospectively collected 3,000 cataract surgery recordings at two Tehran hospitals. Variability was deliberately introduced in two ways: surgeons of different experience levels were included (residents, fellows, and expert attendings), and two different microscope-camera systems were used — a Haag-Streit HS Hi-R NEO 900 recording at 720×480 and 30 fps at Farabi, and a ZEISS ARTEVO 800 recording at 1920×1080 and 60 fps at Noor. Videos were screened for technical quality (excluding incomplete procedures, poor focus, or excessive glare), then de-identified.
Choosing which videos to annotate. Rather than only sampling randomly — which the authors argue risks over-representing routine cases — the team used a content-aware, purposeful sampling strategy supervised by a senior ophthalmic surgeon. It targeted procedural heterogeneity (stochastic workflow variation), visual complexity (specular reflections and occlusions), and behavioral diversity (a balanced spread of proficiency from novice to expert).
Annotating phases. A 13-phase taxonomy covering the procedure from Incision to Tonifying-Antibiotics (plus an Idle phase) was defined. Annotators were trained with gold-standard "video anchor" clips, then annotated on a custom platform called SurgiNote using a coarse-to-fine protocol — locating transitions at second-level resolution first, then refining frame by frame. Quality was controlled by auditing roughly 10% of videos fully and reviewing transition points in the remaining 90%, with weekly dispute-resolution sessions for disagreements.
Annotating segmentation. Frames were sampled from the 150 phase-annotated videos with a minimum 0.5-second gap between frames from the same video, covering all 13 phases and all instruments, while deliberately retaining visually difficult frames (high inter-instrument similarity, boundary ambiguity, specular reflections). Polygon masks were drawn in Roboflow by eight ophthalmology residents, with an ontology defining how tight polygons should be for hard classes such as the transparent Cornea. 15% of annotations were audited by two senior reviewers.
Annotating tracking. For 170 capsulorhexis clips selected to span a wide range of psychomotor proficiency, an AI-assisted pipeline generated first-pass masks: a YOLO11-L model fine-tuned on the segmentation subset, plus custom tracking logic built on BoT-SORT with a tracklet memory buffer and Kalman filtering to bridge occlusions. Instrument tips were computed algorithmically with a constraint-based geometric regression using least-squares line fitting on the mask body, rather than manual pixel clicking. Human annotators then verified spatial boundary adherence and corrected any identity switches.
Annotating skill. A panel of three consultant ophthalmic surgeons and two medical education experts adapted six indicators from GRASIS and ICO-OSCAR — Instrument Handling, Motion, Tissue Handling, Microscope Use, Commencement of Flap, and Circular Completion — each scored on a 5-point scale with novice, intermediate, and competent anchors. Adverse events were recorded as a binary flag plus qualitative descriptions and risk factors. The overall proficiency score is the unweighted mean of the six final indicator scores.
Running the benchmarks. For phase recognition, the authors compared two-stage pipelines (ImageNet-pretrained ResNet50 and EfficientNet-B5, plus frozen DINO and CLIP ViT-B/16 features, modeled with LSTM or GRU on 10-frame clips) against end-to-end Kinetics-400-pretrained models (SlowFast, X3D, R(2+1)D, MC3, R3D, MViT, Video Swin) and stronger video-level architectures (TeCNO's MS-TCN and ASFormer). Videos were downsampled to 4 fps, and the visually similar, underrepresented Viscoelastic and Anterior Chamber Flushing phases were merged. For instance segmentation, COCO-pretrained Mask R-CNN (ResNet-50), YOLOv8-L, and YOLOv11-L were trained on a 70/20/10 video-level split, and SAM and SAM2 were evaluated zero-shot with ground-truth bounding-box prompts (an oracle localization signal, described as an upper bound on zero-shot quality). For skill classification, 3D-CNNs (X3D-M, SlowFast R50, R(2+1)D-18, R3D-18), CNN-LSTM and CNN-GRU hybrids, and TimeSformer were trained on 100-frame snippets at 10 fps and aggregated by averaging predicted posterior probabilities. Kinematic analysis used cumulative instrument path length, computed as the sum of frame-to-frame Euclidean distances of the tip, as a proxy for motion economy.
Why This Matters
Impact on research. The paper targets a specific structural weakness in surgical computer vision: benchmarks that are single-center, single-task, or both, which make it hard to tell whether a model has learned surgery or learned one hospital's camera, lighting, and surgical style. By combining two centers, two camera systems, multiple experience levels, four annotation layers, and paired skill-and-tracking labels on the same clips, Cataract-LMM enables directly measuring cross-center generalization and linking model outputs to objective competency scores rather than only segmentation or classification accuracy.
Real-world applications (as supported by the paper's framing):
- Postoperative assessment and standardized feedback for surgical trainees, using automated skill scores and motion-derived metrics such as path length.
- Intraoperative guidance and real-time workflow recognition, supported by the clip-level (causal) phase-recognition benchmark designed to simulate real-time deployment.
- Post-hoc video indexing and retrospective analysis of surgical archives, supported by the video-level (non-causal) benchmark.
- Surgical scene understanding for instrument detection and tracking, relevant to instrument handling and safety-event documentation such as radial capsular tears or zonular dehiscence.
Industry relevance. Surgical robotics and CAS companies need annotated video of the kind of messy, variable procedures their products will encounter, not curated ideal-case footage. The two-source design (with S1/S2 source identifiers embedded in filenames specifically to enable domain-adaptation benchmarking) and the modular Hugging Face release structure make it usable for teams evaluating whether a model will hold up at a new hospital site.
Future Directions
- Close the reported-data gap on baseline performance. The paper promises a Technical Validation section characterizing baseline results for phase recognition, instance segmentation, tracking, and skill assessment; the supplied content stops before that section, so the actual comparative numbers for all four benchmarks remain to be examined.
- Extend beyond binary skill classification. The authors explicitly note that Cataract-LMM retains full continuous scores and six individual rubric indicators, enabling regression on continuous scores, multi-class categorization via expert-defined thresholds or data-driven clustering, and assessment using motion-derived kinematic features.
- Scale annotation coverage. The annotations currently cover subsets of the 3,000-video corpus: 150 videos for phase labels and segmentation, and 170 clips for tracking and skill. Expanding annotation to the full corpus, or to additional procedures beyond the capsulorhexis phase used for tracking and skill labels, is an open direction implied by the dataset's structure.
- Generalize the multi-task formulation across more centers and countries. The paper's out-of-distribution test is Farabi to Noor (70 videos, one alternative center). Whether the same multi-task models transfer across additional hospitals, equipment vendors, and populations is untested.
Target Audience
This paper is most useful to computer vision and medical-imaging researchers building multi-task or domain-adaptive models for surgical video; to clinical researchers and surgical educators interested in objective, video-based competency assessment; and to engineers at surgical robotics, ophthalmology device, and CAS companies who need a realistic benchmark for workflow recognition, instrument segmentation, and tracking. Clinicians in ophthalmology training programs may also benefit from the rubric design, but the benchmark and architecture details assume a machine-learning background.
Authors’ abstract
Computer-assisted surgery research requires large, deeply annotated video datasets that capture clinical and technical variability. Existing cataract surgery resources lack the diversity and annotation depth required to train generalizable deep-learning models. To address this gap, we present a dataset of 3,000 phacoemulsification cataract surgery videos acquired at two surgical centers from surgeons with varying expertise. The dataset provides four annotation layers: temporal surgical phases, instance segmentation of instruments and anatomical structures, instrument-tissue interaction tracking, and quantitative skill scores based on competency rubrics adapted from ICO-OSCAR and GRASIS. We demonstrate the technical utility of the dataset through benchmarking deep learning models across four tasks: workflow recognition, scene segmentation, instrument-tissue interaction tracking, and automated skill assessment. Furthermore, we establish a domain-adaptation baseline for phase recognition and instance segmentation by training on one surgical center and evaluating on a held-out center. Ultimately, these multi-source acquisitions, multi-layer annotations, and paired skill-kinematic labels facilitate the development of generalizable multi-task models for surgical workflow analysis, scene understanding, and competency-based training research.