Skip to content
AI.info

Research

HeiCo-FOCUS: A Clinically Grounded Dataset for Long-Context Video Understanding

Overview Research area: Computer vision / multimodal machine learning — specifically long-context video understanding, vision-language model (VLM) evaluation, and surgical (medical) AI. Technical leve

HeiCo-FOCUS: A Clinically Grounded Dataset for Long-Context Video Understanding
arXiv
2610.10156
Published
2026-10-07
Authors
Leon Mayer, Lucas Luttner, Patrick Godau, Kai Fritzsche, Annika Reinke, Leonie Boland, Jule Brandt, Janne Heinecke, Chloe K. Nobuhara, Niklas Holzwarth, Evangelia Christodoulou, Marcel Knopp, Dominik Michael, Pascale Piermarco, Saliq Neyaz, Korhan Derin Özarslan, Jakob Hennighausen, Carlos Aumente-Maestro, Tim Rädsch, Dheeraj Baji, Peter Maximilian Full, Finn Aichholz, Justus Veit Erpenbeck, Linus Finn Schott, Bastian Winkelhausen, Claas de Boer, Bianca Güttner, Anneli Hummel, Gregor Just, Max Kirchner, Chenyang Li, Rozenn Raffaut, Ariel Rodriguez, Danush Kumar Venkatesh, Kevin Wang, Jinjing Xu, Mona Sheikh Zeinoddin, Salman Khan, Thomas M. Pausch, Stefanie Speidel, Danail Stoyanov, Daniel A. Hashimoto, Fiona R. Kolbinger, Thomas G. Weiser, Lena Maier-Hein

AI summary

Overview

  • Research area: Computer vision / multimodal machine learning — specifically long-context video understanding, vision-language model (VLM) evaluation, and surgical (medical) AI.
  • Technical level: Advanced. The paper assumes familiarity with VLM benchmarking, video temporal grounding, and clinical evaluation terminology.
  • Scope in one sentence: The paper introduces HeiCo-FOCUS, a clinically grounded surgical video benchmark of 30,000 visual question answering pairs built to test whether VLMs can maintain cumulative temporal consistency over procedures lasting up to hours.

What This Paper Is About

Existing video benchmarks mostly test short-term reasoning within clips of a few minutes, so they do not measure whether a model can track the same objects across an entire procedure as those objects are inserted, manipulated, occluded, and removed. HeiCo-FOCUS targets this gap by turning a real patient-safety problem — ensuring that foreign objects (sponges, sutures, and similar items) are retrieved before surgery ends — into a structured long-context evaluation task over 30 Heidelberg Colorectal surgeries. The goal is to give the surgical AI community a trusted, expert-verified benchmark that separates local perception from long-horizon, cumulative reasoning.

Key Contributions

  1. A new long-context surgical VQA benchmark. HeiCo-FOCUS contributes 30,000 visual question answering pairs built on all 30 Heidelberg Colorectal surgeries of the HeiCo dataset (10 sigmoid resections, 10 rectal resections, 10 proctocolectomies), covering roughly 96 hours of surgical video.
  2. A capability taxonomy and a three-track evaluation design. Questions span five core capabilities (object recognition and identity matching, temporal grounding, aggregation, event and procedural understanding, complex reasoning) and are split into a FRAME track (single images), a SEGMENT track (video segments up to 5 min), and a PROCEDURE track (video from the beginning up to a time point t, ranging 5–296 min).
  3. Persistent object identities and an expert-driven annotation pipeline. Instead of generating questions from existing labels with a language model, the authors annotated individual object instances across time, involving 39 domain experts (7 expert surgeons, 18 medical students, 14 surgical AI researchers) across six pipeline stages.
  4. A rigorous, bias-aware evaluation of ten frontier VLMs plus a text-only baseline. The benchmark explicitly filters questions that can be answered without images and reports stratified results by capability, track, cost, and a human baseline.

Main Findings

  • The benchmark is far from solved. Across the video tracks, only around half of the evaluated models clearly outperform a text-only baseline. The text-only baseline (GPT-5.6 Sol without visual input) scored 20.7% on FRAME, 30.4% on SEGMENT, and 27.8% on PROCEDURE.
  • Event and procedural understanding is the strongest capability; temporal grounding is the weakest. Event and procedural understanding achieved the highest scores (mean Accuracy 56.5% across all models), while temporal grounding obtained a mean Accuracy of 19.7% across all evaluated models in the abstract, with 26.9% on the SEGMENT track and 12.6% on the PROCEDURE track.
  • FRAME track results. Gemini 3.8 Flash led with 44.8% [39.1, 50.1] Accuracy, followed by Gemini 3 Flash (43.2% [39.0, 48.0]) and GPT-5.6 Sol (38.9% [34.3, 44.0]) — gains of 24.1, 22.5, and 18.2 percentage points over the baseline. Muse Spark 1.3 and Gemini 3.5 Flash Lite followed with gains of 15.0 and 14.3 percentage points. The remaining five models (Claude Sonnet 5, Grok-4.3, MiMo-V 2.5, DeepSeek V4 Flash Vision, and Nova 2 Lite) stayed within 6 percentage points of the baseline.
  • SEGMENT track results. Mean Accuracy was 41.7% across the ten models, 11.3 percentage points above the baseline on average. Gemini 3.8 Flash reached 59.4% [53.9, 64.9], GPT-5.6 Sol 58.3% [53.2, 63.1], and Gemini 3 Flash 49.1% [45.3, 52.5], improving on the baseline by 29.0, 27.9, and 18.7 percentage points.
  • PROCEDURE track results. Mean Accuracy was 37.9% across the ten models versus a 27.8% baseline. Gemini 3.8 Flash reached 56.7% [52.4, 62.0], GPT-5.6 Sol 51.0% [46.7, 55.1], and Muse Spark 1.3 43.8% [38.0, 49.3] — gains of 28.9, 23.2, and 16.0 percentage points. DeepSeek V4 Flash Vision, MiMo-V 2.5, and Nova 2 Lite scored below the baseline on this track.
  • Fine-grained capability weaknesses. Object recognition and identity matching proved more challenging than aggregation, with object attributes yielding the lowest scores (22.7% averaged over all models). Temporal grounding dropped on the PROCEDURE track to 12.6%, less than half its SEGMENT value.
  • Human raters also struggle on long-horizon questions. Three raters (two surgical residents and one medical student) answered 300 questions on the test videos and reached 62.0%, 69.8%, and 57.4% Accuracy on the FRAME, SEGMENT, and PROCEDURE tracks. On the same 300 questions, the best model on each track scored 12.0, 12.1, and 6.9 percentage points lower, respectively. Human raters reached only 35% on event aggregation and 40% on duration estimation (20 questions each).
  • Accuracy did not follow price. The most accurate model, Gemini 3.8 Flash, cost about a fifth of GPT-5.6 Sol per question on the video tracks.
  • Annotation scale. Stage 4 produced 324,273 manually curated frames containing 79,682 verified foreign-object bounding boxes, of which 55,629 carried instance identities across time; clips were annotated for visibility only because instance linking was infeasible even for expert surgeons. The taxonomy contained 85 question templates and was verified on about 500 diverse sample questions in the final pipeline stage.

Methodology in Plain English

The authors started from HeiCo, an existing clinically curated dataset of 30 Heidelberg colorectal surgeries with transparent provenance. Every procedure includes 10 sigmoid resections, 10 rectal resections, and 10 proctocolectomies, and the original HeiCo split is retained, with sigmoid resections recommended as the held-out test set (4,000 FRAME, 4,000 SEGMENT, and 2,000 PROCEDURE questions).

Building the questions took six stages. First, short clips were screened for foreign objects. Second, frames sampled at 1 frame per second were annotated with bounding boxes and class labels. Third, automated methods corrected and completed those boxes. Fourth, the most resource-intensive stage, 18 medical students assigned consistent instance identities to each object across time, so that each object's trajectory through the procedure is known. Fifth, anchor moments — insertion and retrieval events, long absences, cluttered scenes — were derived from these trajectories and used as the basis for question generation, combining expert-written questions with automatically instantiated templates, balanced by drawing iteratively from the least-represented capability, template, and answer value. Sixth, the whole pipeline was verified end-to-end by experts.

Each procedure contributed 1,000 VQA pairs (400 FRAME, 400 SEGMENT, 200 PROCEDURE), with questions assigned a primary capability category and optional secondary tags. To prevent questions being answerable from text alone, three models (Mistral Small 2603, Gemini 2.5 Flash Lite, and DeepSeek V4 Flash) answered every candidate without the image or video, and a question was discarded if all three answered correctly.

Evaluation covered ten frontier VLMs via cloud APIs: GPT-5.6 Sol, Claude Sonnet 5, Gemini 3.8 Flash, Gemini 3 Flash, Gemini 3.5 Flash Lite, Meta Muse Spark 1.3, xAI Grok-4.3, DeepSeek V4 Flash Vision, Amazon Nova 2 Lite, and Xiaomi MiMo-V 2.5. Inputs were adapted per model, provided either as native video or chronologically ordered frame sequences; SEGMENT windows were sampled at 1 frame per second and PROCEDURE windows were represented by 300 uniformly spaced frames, with all visual inputs standardized to 210 × 360 pixels. Accuracy was the primary metric, with open-ended, multiple-choice, and matching answers graded for semantic equivalence by an LLM judge (gpt-oss-120b, validated against two human raters). Uncertainty was reported with bootstrapped 95% confidence intervals, and the hierarchical structure of the data (multiple questions per video or segment) was accounted for through hierarchical aggregation and bootstrapping.

Why This Matters

  • Impact on research. The paper argues that existing surgical VQA datasets are limited to frame-level or short-clip questions, are often generated from rule-based or AI-derived sources, and rarely undergo systematic expert verification. HeiCo-FOCUS provides a benchmark that combines long-context video understanding, clinically grounded question design, and rigorous expert validation, and it explicitly measures cumulative temporal consistency that short-clip benchmarks miss.
  • Clinical safety. Retained foreign objects can cause serious complications, and the benchmark is anchored to the objective, verifiable task of confirming that objects inserted during a procedure are retrievable at the end.
  • Research prioritization. The authors note that laparoscopic cholecystectomy accounts for more than 50% of surgical AI studies despite complications in roughly 2–3%, and deliberately chose colorectal procedures — which have complications in up to one third of patients, up to 15 times higher complication rates, 5-fold longer procedures, and on average three times more foreign object instances per video.
  • Benchmark quality standards. Compared with more than 400 analyzed medical imaging AI benchmarks (median 3 annotators), HeiCo-FOCUS involved more than an order of magnitude more domain experts.
  • Industry relevance. The results show that current frontier VLMs are not yet reliable for long-horizon surgical video analysis, with accuracy varying by more than 20 percentage points between models and cost not tracking accuracy — relevant for anyone building or procuring clinical video AI. Code is released at https://github.com/IMSY-DKFZ/orena-focus, with data to be published on Hugging Face (access via the corresponding authors until then), under a CC BY-NC-SA 4.0 license to comply with the original HeiCo license.

Future Directions

  • Expand procedural diversity. The authors note the dataset derives from a limited number of colorectal procedures and may introduce biases related to surgical style, instrumentation, patient population, or recording conditions; they call for expanding procedures, institutions, and surgical settings, and state that benchmark performance should not be interpreted as evidence of safe generalization.
  • Grow the reasoning question pool. Future work should emphasize increasing the number and diversity of expert-generated reasoning questions, the category that required the most additional clinical expertise beyond instance annotations.
  • Study prompt sensitivity and run-to-run variability. The authors state they did not systematically investigate prompt sensitivity or variability across repeated model runs, and highlight the need to prioritize efficient evaluation protocols given the monetary and environmental costs of evaluating long-context VLMs on extended surgical videos.
  • Resolve remaining annotation and taxonomy debates. The paper acknowledges that some annotation and taxonomy decisions may remain debatable, and provides a community feedback mechanism through a restricted post-release review period for reporting ambiguous or erroneous cases.

Target Audience

This paper is most useful to researchers and engineers working on vision-language models for long-form video, medical and surgical AI practitioners evaluating model readiness for clinical workflows, benchmark designers interested in expert-verified annotation pipelines, and clinical stakeholders — surgeons and patient-safety researchers — concerned with retained foreign objects and the reliability of automated surgical video analysis.

Authors’ abstract

Recent advances in Vision-Language Models (VLMs) have led to rapid progress in video understanding across a wide range of benchmark tasks. However, existing evaluations largely focus on short-term reasoning, failing to assess a critical capability: maintaining cumulative temporal consistency over extended time horizons. To close this evaluation gap, we introduce HeiCo-FOCUS, a clinically grounded dataset for evaluating long-context video understanding through the task of Foreign Object Contextual Understanding in Surgery. Built on a dataset of Heidelberg Colorectal surgeries, this task requires models to continuously track multiple objects as they are inserted, manipulated, occluded, and removed over procedures lasting up to hours. HeiCo-FOCUS comprises 30,000 visual question answering (VQA) pairs covering five core capabilities: object recognition, temporal grounding, aggregation, event and procedural understanding, and complex reasoning. The dataset was constructed through a rigorous multi-stage annotation pipeline involving large-scale crowd annotation and 39 surgical domain experts to ensure high quality and clinical relevance. To systematically probe model behavior, we introduce a multi-track evaluation framework that progressively increases temporal and contextual demands from single frames to full procedures. Experiments with ten frontier VLMs show that HeiCo-FOCUS tasks are far from solved: only around half of the models clearly outperform a text-only baseline. Across the video tracks, models perform best on event and procedural understanding (mean Accuracy: 56.5% across all models), while temporal grounding remains particularly challenging for all evaluated models (mean Accuracy: 19.7%). We therefore expect HeiCo-FOCUS to serve as a catalyst for the development of models capable of reliable, temporally consistent reasoning over hours-long videos.

Read the original paper