Skip to content
AI.info

Research

A Multi-View Pipeline and Benchmark Dataset for 3D Hand Pose Estimation in Surgery

A Multi-View Pipeline and Benchmark Dataset for 3D Hand Pose Estimation in Surgery arXiv: 2601.15918v1 [cs.CV], 22 Jan 2026 — Valery Fischer, Alan Magdaleno, Anna-Katharina Calek, Nicola Cavalcanti, N

A Multi-View Pipeline and Benchmark Dataset for 3D Hand Pose Estimation in Surgery
arXiv
2601.15918
Published
2026-01-22
Authors
Valery Fischer, Alan Magdaleno, Anna-Katharina Calek, Nicola Cavalcanti, Nathan Hoffman, Christoph Germann, Joschua Wüthrich, Max Krähenmann, Mazda Farshad, Philipp Fürnstahl, Lilian Calvet

AI summary

A Multi-View Pipeline and Benchmark Dataset for 3D Hand Pose Estimation in Surgery

arXiv: 2601.15918v1 [cs.CV], 22 Jan 2026 — Valery Fischer, Alan Magdaleno, Anna-Katharina Calek, Nicola Cavalcanti, Nathan Hoffman, Christoph Germann, Joschua Wüthrich, Max Krähenmann, Mazda Farshad, Philipp Fürnstahl, Lilian Calvet (University Hospital Balgrist, University of Zurich; ETH Zürich)

Overview

Research area: Surgical computer vision — multi-view 3D hand pose estimation and benchmark dataset construction.

Technical level: Intermediate. The pipeline is modular and training-free, so the high-level logic is accessible, but full appreciation requires familiarity with 2D/3D pose estimation, triangulation, and multi-view geometry.

Scope: The paper proposes a training-free, multi-view pipeline for 3D hand pose estimation in operating-room video and releases a new annotated surgical benchmark dataset for evaluating such methods.

What This Paper Is About

Estimating the 3D pose of a surgeon's hands from video would enable skill assessment, workflow analysis, and touch-free control of surgical systems, but surgical scenes are hostile to existing pose estimators: lighting is bright, focused, and glare-producing; instruments and staff cause frequent occlusions; and surgical gloves remove the texture cues that hand models rely on. There is also a shortage of annotated surgical data with 3D ground truth for training or evaluating models. This paper addresses both problems by building a pipeline that uses only off-the-shelf pretrained models (no domain-specific fine-tuning) and by releasing a multi-view surgical benchmark dataset with manual 2D annotations and triangulated 3D ground truth.

Key Contributions

  1. A modular, training-free pipeline for accurate multi-view 3D hand pose estimation in surgical environments, requiring no retraining or fine-tuning on surgical data.
  2. A new real multi-view surgical dataset comprising over 68,000 frames and 3,000 manually annotated 2D hand poses with corresponding triangulated 3D ground truth, captured in a replica operating room under six defined levels of scene complexity.
  3. A comprehensive evaluation including quantitative and qualitative comparison against state-of-the-art 2D and whole-body pose methods, plus an ablation study over the optimization loss terms.
  4. A constrained 3D optimization formulation combining reprojection, temporal smoothness, shape consistency, and (optionally) biomechanical priors, solved frame-by-frame with L-BFGS-B.

Main Findings

  • Large 2D accuracy gain: The proposed method achieves the lowest Mean Joint Error (MJE) in Table 1 at 11.8, compared with 29.5 (Sapiens), 27.5 (RTMPose), 23.5 (DWPose), 20.9 (Sapiens→HM), 16.9 (RTM→HM), and 17.0 (DW→HM). It also leads on mPCK_2D (22.8 vs. 7.7 / 12.6 / 14.0 / 17.4 / 20.8 / 20.9) and AP (70.3 vs. 67.3 / 66.9 / 66.9 / 66.9 / 65.4 / 66.7).
  • Large 3D accuracy gain: In Table 2, the full pipeline ("Ours: RTM + T → HM") reports MPJPE 8.5 mm, MRE 8.2 px, mPCK_2D 76.6, and mPCK_3D 74.7, versus MPJPE 35.8 / 51.9, MRE 39.8 / 58.1, mPCK_2D 39.8 / 42.3, and mPCK_3D 40.8 / 41.4 for the RTM and RTM→HM baselines.
  • Headline reductions: The abstract reports a 31% reduction in 2D mean joint error and a 76% reduction in 3D mean per-joint position error relative to baselines.
  • Tracking is the key ingredient: The authors attribute the improvement largely to the tracker's object-agnostic design, which localizes hands without prior knowledge of the target and yields more reliable bounding boxes than whole-body hand estimates alone.
  • Ablation — temporal smoothness helps consistently: The temporal smoothness constraint "consistently yields significant improvements."
  • Ablation — shape consistency helps on fast motion: The shape consistency term provides comparable average performance but proves particularly effective for fast hand motions in L2 sequences, where it mitigates degradation of 2D keypoints; it is therefore retained in the baseline configuration.
  • Ablation — biomechanics hurts: Including the biomechanical constraint decreases performance, likely because the most plausible biomechanical shape does not precisely match the multi-view 2D predictions.
  • Baseline optimization weights: Results in Table 2 use λ_reproj = 1.0, λ_smooth = 20.0, λ_shape = 50.0, and λ_biomech = 0.

Methodology in Plain English

The method follows a top-down, multi-stage design that extends the "detect the person first, then estimate the pose inside the box" idea from body pose estimation to hands.

  1. Person detection. Each synchronized camera view is processed with YOLOv11 to detect the surgeon, since large-scale detectors remain reliable under unseen conditions.
  2. Whole-body pose estimation as initialization. A whole-body 2D pose estimator is applied inside each person bounding box to produce full-body keypoints including coarse hand locations. These are treated as priors, not final answers.
  3. Keyframe selection. For each hand, the frame with the highest average hand-joint confidence above 0.3 initializes tracking with Efficient Track Anything (TAM). Tracking is propagated forward and backward; if tracking fails or hits sequence boundaries, the successfully tracked frames for that hand are removed and the process restarts from the next highest-confidence frame until no valid detections remain. This preserves hand identity through occlusions, rapid movement, and intermittent detections.
  4. Fine 2D hand pose refinement. Each tracked hand crop is passed to a specialized 2D hand pose estimator operating at or near native image resolution, recovering joint detail lost at the whole-body stage. Individual keypoints with confidence below 0.1 are discarded to avoid propagating bad 2D predictions into 3D.
  5. Constrained 3D optimization. Starting from a canonical hand attached to the triangulated body pose, the pipeline minimizes a weighted sum of a confidence-weighted reprojection loss, a temporal smoothness loss penalizing joint changes between consecutive frames, and a shape consistency loss that uses an orthogonal Procrustes (SVD-based) alignment so that hand shape stays constant while orientation changes. An optional biomechanical loss enforcing joint-angle limits and anatomical feasibility is also studied. Frames are optimized sequentially using the previous frame's result, solved with L-BFGS-B.

Dataset. Data were recorded in a replica operating room across two sessions: SPP (spinal preparation) with 12 synchronized static GoPro Hero Black 12 cameras at 4K, 30 fps, fixed shutter, and SPI (spinal instrumentation) with 11 GoPro Hero Black 12 cameras. Cameras were synchronized via a custom PCB controller ensuring sub-frame alignment, and calibrated with OpenCV and Flückiger et al. (2025). The appendix specifies 4 far-field plus 8 near-field cameras for SPP and 4 far-field plus 7 near-field for SPI, all at 4K (9:16 aspect ratio), 30 fps, Linear lens mode, white balance fixed at 4500 K, far-field shutter 1/240 s and ISO 400, near-field shutter 1/960 s with ISO 400 for SPP and 100 for SPI. Six complexity levels L0–L5 are defined — from single-person static scenes (L0) to dynamic multi-person scenes with frequent occlusions (L5) — with the study focusing on L0–L2 and L3–L5 reserved for future benchmarking. The dataset contains 20 sequences of 300 frames each, over 68,000 synchronized multi-view frames. For each sequence, 10 keyframes per camera were manually labeled, yielding approximately 3,000 hand pose annotations; 2D labels were triangulated into 3D ground truth, and their reprojections serve as the ground-truth 2D poses. Annotation quality was checked via multi-annotator cross-review.

Metrics. Mean per joint position error (MPJPE), mean reprojection error (MRE), PCK_2D@τ and PCK_3D@δ, mean PCK (mPCK) averaged over thresholds 5, 10, 20, 30 for 2D and 5, 10, 25, 50 for 3D, average precision (AP) following the COCO protocol, and mean joint error (MJE). Baselines are Sapiens, RTMPose, and DWPose as whole-body models, plus a hand pose model from RTMPose applied to crops (noted HM), with the notation A→B meaning model B is applied inside bounding boxes from model A.

Why This Matters

Impact on research. The paper provides both a strong baseline and a shared benchmark for a problem where annotated surgical data is scarce, giving other groups a reproducible reference point and a dataset with triangulated 3D ground truth rather than 2D labels alone. The demonstration that a training-free, off-the-shelf stack can outperform domain-adapted whole-body models is a meaningful result for surgical computer vision, where data collection and annotation are expensive.

Real-world applications:

  • Skill assessment — objective analysis of surgeon hand motion during procedures.
  • Workflow recognition — tracking surgical activity to understand and analyze procedural phases.
  • Robot-assisted and touch-free interaction — natural interaction with surgical robots, imaging systems, and augmented or virtual reality platforms while preserving the sterile field by reducing physical contact with control panels or non-sterile surfaces.
  • Surgical training — detailed motion analysis to support training and feedback.

Industry relevance. Surgical robotics companies, operating-room imaging vendors, AR/VR surgical platform developers, and hospital systems investing in procedural analytics all stand to benefit from robust hand tracking that does not require large in-domain training sets. The released benchmark also gives vendors a common yardstick for comparing hand-tracking claims.

Stated limitations. The current implementation focuses on controlled single-person scenarios; multi-person settings, frequent self-occlusions, and low hand visibility are not yet handled effectively. Runtime is a barrier: the optimization and tracking stages require around 30–45 minutes for a 300-frame sequence on an RTX 2060 GPU.

Future Directions

  1. Faster, learning-based feedforward models trained using the 3D annotations produced by the current pipeline, targeting near real-time performance.
  2. Extending the system to multi-person and occlusion-heavy scenes, which the authors identify as a key research direction and which correspond to complexity levels L3–L5 that the paper defines but does not yet evaluate.
  3. Benchmarking the higher complexity levels (L3–L5) with the released dataset, since those levels are explicitly included for future benchmarking.
  4. Reconsidering biomechanical priors, given that the biomechanical loss reduced measured performance in the ablation while temporal smoothness and shape consistency improved it.

Target Audience

Researchers and engineers working on pose estimation, surgical computer vision, and operating-room analytics; teams building surgical skill assessment, workflow recognition, or touch-free/robotic control interfaces; and dataset builders interested in multi-view acquisition and 3D annotation methodology. Graduate students entering surgical computer vision will find the pipeline design accessible and the benchmark immediately useful, while the appendix (camera settings, scene descriptions, ethics approval BASEC Nr. 2021-01196, and ground-truth bone-length optimization) is valuable for groups planning their own multi-view surgical capture.

Authors’ abstract

Purpose: Accurate 3D hand pose estimation supports surgical applications such as skill assessment, robot-assisted interventions, and geometry-aware workflow analysis. However, surgical environments pose severe challenges, including intense and localized lighting, frequent occlusions by instruments or staff, and uniform hand appearance due to gloves, combined with a scarcity of annotated datasets for reliable model training. Method: We propose a robust multi-view pipeline for 3D hand pose estimation in surgical contexts that requires no domain-specific fine-tuning and relies solely on off-the-shelf pretrained models. The pipeline integrates reliable person detection, whole-body pose estimation, and state-of-the-art 2D hand keypoint prediction on tracked hand crops, followed by a constrained 3D optimization. In addition, we introduce a novel surgical benchmark dataset comprising over 68,000 frames and 3,000 manually annotated 2D hand poses with triangulated 3D ground truth, recorded in a replica operating room under varying levels of scene complexity. Results: Quantitative experiments demonstrate that our method consistently outperforms baselines, achieving a 31% reduction in 2D mean joint error and a 76% reduction in 3D mean per-joint position error. Conclusion: Our work establishes a strong baseline for 3D hand pose estimation in surgery, providing both a training-free pipeline and a comprehensive annotated dataset to facilitate future research in surgical computer vision.

Read the original paper