Skip to content
AI.info

Research

SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding

Overview Research area: Surgical computer vision and medical multimodal large language models (LLMs), specifically benchmark dataset construction for surgical scene understanding. Technical level: Int

SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding
arXiv
2511.21339
Published
2025-11-26
Authors
Tae-Min Choi, Tae Kyeong Jeong, Garam Kim, Jaemin Lee, Yeongyoon Koh, In Cheul Choi, Jae-Ho Chung, Jong Woong Park, Juyoun Park

AI summary

Overview

Research area: Surgical computer vision and medical multimodal large language models (LLMs), specifically benchmark dataset construction for surgical scene understanding.

Technical level: Intermediate. Readers should be comfortable with multimodal LLM training pipelines (pre-training plus instruction tuning, LoRA), Visual Question Answering, and segmentation metrics such as IoU/mIoU.

Scope in one sentence: The paper introduces SurgMLLMBench, a unified benchmark that merges six surgical video datasets — including the newly collected MAVIS micro-surgery dataset — under a single annotation taxonomy spanning stage, phase, step, instrument action, and pixel-level instrument segmentation, together with baseline results for two multimodal LLMs.

What This Paper Is About

Existing surgical datasets are mostly built for Visual Question Answering with inconsistent label taxonomies, and most lack pixel-level segmentation, so models trained on them cannot be compared cleanly and cannot produce fine-grained visual evidence. The authors build a single benchmark that aligns laparoscopic, robot-assisted, and micro-surgical data into one schema of four tasks (phase recognition, step classification, instrument-centered action detection, and instrument segmentation) plus template-generated VQA prompts, and they test whether one model trained on this corpus works across domains and transfers to a dataset it never saw.

Key Contributions

  1. SurgMLLMBench, a unified multimodal benchmark that integrates six surgical datasets — Cholec80, EndoVis2018, AutoLaparo, GraSP, MISAW, and the new MAVIS — under one annotation schema covering stage, phase, step, instrument action, and pixel-level instrument segmentation (112.80 video hours, 561,418 total frames as reported in the paper's dataset table).
  2. The MAVIS dataset (Micro-surgical Artificial Vascular anastomosIS): 19 videos of 1 mm artificial vessel anastomosis performed by three expert micro-surgeons (seven videos from Surgeon 1, seven from Surgeon 2, five from Surgeon 3), recorded at 1920 × 1080 and sampled at 1 FPS, totaling 10,652 frames. MAVIS introduces a hierarchical stage–phase–step workflow annotation, with Stage reported as the first such attribute across surgical domains.
  3. A reproducible integration and VQA generation pipeline that converts heterogeneous sources into COCO-style frame-level metadata, unifies labels, and produces structured question–answer pairs from fixed prompt templates across five query types rather than using a generative language model.
  4. Quantitative baselines with OMG-LLaVA and LLaVA under two training configurations (per-dataset fine-tuning versus a single instruction-tuned model on SurgMLLMBench) plus a generalization test on the held-out MAVIS dataset.

Main Findings

  • One cross-domain model is competitive: A single model instruction-tuned on SurgMLLMBench (MAVIS excluded) performs comparably to, or surpasses, per-dataset fine-tuned models on EndoVis2018 and AutoLaparo. For example, on AutoLaparo, instruction-tuned LLaVA reached 24.67 and 82.56 versus 23.00 and 57.21 for per-dataset fine-tuned LLaVA, and on EndoVis2018 instruction-tuned LLaVA reached 43.65 and 78.38 versus 34.31 and 58.88.
  • Not all datasets improve: On Cholec80, instruction-tuned OMG-LLaVA dropped to 55.78 and 82.53 from 76.51 and 83.08 under per-dataset fine-tuning; on MISAW and GraSP, LLaVA trained on SurgMLLMBench also showed lower action accuracy (for example 2.95 on GraSP) — the authors attribute this to imbalanced action labels, where semantically related actions across datasets are named differently (Idle vs. Still, Retraction vs. Pull).
  • Generalization to an unseen dataset and a novel task works best for LLaVA: On MAVIS, which was excluded from instruction tuning, LLaVA improved from 67.67 stage / 65.09 phase / 37.59 step / 54.67 count with MAVIS-only fine-tuning to 75.70 / 69.43 / 44.80 / 68.47 when initialized from the SurgMLLMBench instruction-tuned checkpoint (3 epochs of MAVIS fine-tuning).
  • OMG-LLaVA degrades when initialized from the benchmark checkpoint: On MAVIS it fell from 62.84 stage / 63.63 phase / 37.28 step / 55.46 count / 53.96 segmentation (MAVIS-only training) to 43.90 / 55.72 / 29.48 / 47.51 / 36.89 after SurgMLLMBench initialization, which the authors link to a visual gap between datasets limiting transfer of

Authors’ abstract

Recent advances in multimodal large language models (LLMs) have highlighted their potential for medical and surgical applications. However, existing surgical datasets predominantly adopt a Visual Question Answering (VQA) format with heterogeneous taxonomies and lack support for pixel-level segmentation, limiting consistent evaluation and applicability. We present SurgMLLMBench, a unified multimodal benchmark explicitly designed for developing and evaluating interactive multimodal LLMs for surgical scene understanding, including the newly collected Micro-surgical Artificial Vascular anastomosIS (MAVIS) dataset. It integrates pixel-level instrument segmentation masks and structured VQA annotations across laparoscopic, robot-assisted, and micro-surgical domains under a unified taxonomy, enabling comprehensive evaluation beyond traditional VQA tasks and richer visual-conversational interactions. Extensive baseline experiments show that a single model trained on SurgMLLMBench achieves consistent performance across domains and generalizes effectively to unseen datasets. SurgMLLMBench will be publicly released as a robust resource to advance multimodal surgical AI research, supporting reproducible evaluation and development of interactive surgical reasoning models.

Read the original paper