Research
End2Reg: Learning Task-Specific Segmentation for Markerless Registration in Spine Surgery
Overview Research area: Computer vision for medical imaging — markerless RGB-D point cloud registration for image-guided spine surgery, combining task-specific semantic segmentation with deep point cl

- arXiv
- 2512.13402
- Published
- 2025-12-15
- Authors
- Lorenzo Pettinari, Sidaty El Hadramy, Michael Wehrli, Philippe C. Cattin, Daniel Studer, Carol C. Hasler, Maria Licci
AI summary
Overview
Research area: Computer vision for medical imaging — markerless RGB-D point cloud registration for image-guided spine surgery, combining task-specific semantic segmentation with deep point cloud registration.
Technical level: Advanced. The paper assumes familiarity with 3D point clouds, rigid transformation estimation in SE(3), point cloud convolutional networks (KPConv), transformer-based registration (GeoTransformer), and differentiable discrete sampling (Gumbel-Softmax / straight-through estimators).
One-sentence scope: The paper introduces End2Reg, an end-to-end deep learning framework that jointly trains a segmentation module and a registration module so that the segmentation mask is learned purely to minimize the registration loss, removing the need for segmentation labels.
What This Paper Is About
Intraoperative navigation in spine surgery needs millimeter-level alignment between preoperative scans (CT/MRI) and the patient's actual anatomy, but current practice relies on radiation-heavy intraoperative 3D imaging and invasive bone-anchored fiducial markers. Markerless RGB-D approaches are an attractive radiation-free alternative, yet they depend on noisy, automatically generated weak bone labels to segment the exposed anatomy, and those label errors can propagate into the registration step; many also still require a manual step to pick the expected overlap region.
End2Reg's goal is to remove both the segmentation labels and the manual steps by learning a segmentation mask that is optimized for registration rather than for anatomical correctness. The mask is never supervised; it is only shaped by whether the downstream rigid alignment improves.
Key Contributions
- Joint learning of segmentation and registration. The segmentation and registration networks are trained end-to-end with no segmentation labels, so End2Reg learns task-specific segmentations that are optimized for registration rather than for reproducing bone anatomy.
- State-of-the-art registration performance. Validated on the publicly available SpineDepth (ex vivo) and SpineAlign (in vivo) datasets, with reduced median Target Registration Error and mean Root Mean Square Error relative to prior work.
- Robustness to occlusions. The method shows improved robustness under partial occlusions compared to existing methods, including an explicit evaluation of TRE as occlusion increases from 0 to 0.5.
- An ablation isolating the value of end-to-end training. A controlled comparison against the identical architecture trained in a two-step (sequential segmentation → registration) manner, supported by a Wilcoxon signed-rank test.
Main Findings
- Overall error reduction: End2Reg reduces median Target Registration Error by 32% and mean Root Mean Square Error by 61% on the ex- and in-vivo benchmarks.
- SpineDepth (ex vivo) accuracy: Median TRE of 1.8 [1.2, 2.7] mm, RTE of 1.9 [1.2, 2.7] mm, and RRE of 1.3 [0.8, 1.8] degrees, at 620 ms per frame. This compares with AutoReg's reported TRE of 2.7 [1.7, 3.6] mm, and with GeoTransformer at 1.4 [1.0, 2.0] degrees / 1.9 [1.2, 3.1] mm / 2.2 [1.5, 3.4] mm at 638 ms.
- Comparison against classical and learning-based baselines: On SpineDepth, ICP reaches TRE 5.4 [3.8, 7.2] mm, RANSAC+ICP 4.2 [2.7, 6.7] mm, FGR+ICP 17.2 [10.8, 37.2] mm, GMCNet 12.7 [9.5, 17.1] mm, OverlapPredator 5.1 [3.8, 7.0] mm, and oGMM 6.7 [4.3, 9.7] mm. Several classical baselines required manual selection of the expected overlap region.
- SpineAlign (in vivo) results: End2Reg reports mean fitness .71 (±.18) and RMSE 2.80 (±1.31) mm at 450 ms, versus CorrNet at .58 (±.11) and 7.14 (±0.47) mm, and GeoTransformer at .70 (±.12) and 3.04 (±1.76) mm at 455 ms. The authors note End2Reg operates without the coarse pre-to-intraoperative alignment that Daly et al. use as initialization, and that these results are a preliminary benchmark because no reliable ground truth is available.
- Occlusion robustness: In the TRE-versus-occlusion analysis (occlusion ratio from 0 to 0.5), End2Reg shows the smallest TRE across all occlusion levels, which the authors describe as the highest robustness among the compared methods.
- Ablation confirms end-to-end value: Wilcoxon signed-rank tests on SpineDepth show the improvement of End2Reg over the non-joint two-step training is significant (p ≪ 0.05) with a moderate effect size (r = 0.47) and a reduced outlier ratio (3.8% vs 7.0%).
- Segmentation is deliberately not anatomical: End2Reg's masks score lower than weakly supervised KPConv on Dice (.50 [.42, .54] vs .73 [.68, .77]) and IoU (.70 (±.06) vs .84 (±.05)), but achieve lower one-sided Chamfer distance and HD95, indicating better containment of the weak bone labels. The authors state overlap-based metrics are not the primary evaluation criterion because the task-specific mask is not meant to reproduce anatomical ground truth.
Methodology in Plain English
The pipeline treats each intraoperative RGB-D frame as a point cloud (the "target") and pairs it with a point cloud sampled from the preoperative anatomy (the "source"). A segmentation network — a KPConv encoder-decoder in a U-Net-like shape — looks at each intraoperative point and predicts a binary label: keep or discard. The kept points are supposed to be the ones useful for alignment, mainly bone.
The trick is how that mask is fed into the registration network, which is based on GeoTransformer. In the standard KPConv setup, source points get a constant feature of 1 and the network relies only on geometry. Here, each target point instead gets its predicted binary mask value as its feature. Because the convolution only sums contributions from neighbors, points labeled 0 effectively vanish from the convolution output while still occupying space geometrically — so the registration module focuses on the regions the segmentation module flagged.
The problem is that choosing a hard 0/1 label with arg max is not differentiable, so gradients from the registration loss cannot reach the segmentation network. The authors solve this with the Gumbel-Softmax estimator: Gumbel noise is added to the logits, and a temperature-controlled softmax produces a continuous, differentiable approximation of the discrete sample. They then use a Straight-Through Gumbel-Softmax construction — forward pass uses the hard discrete mask, backward pass uses the soft relaxation via a stop-gradient operator — so the whole system trains end-to-end with the registration loss alone. Training uses GeoTransformer's dual-phase coarse-to-fine loss.
Experiments used SpineDepth (ten cadaver specimens, pedicle screw placement, preoperative 3D models from CT, ground-truth poses from optical tracking, 8-fold cross-validation, frames up to 50% occlusion retained, no viewpoint filtering) and SpineAlign (24 lumbar open spine surgeries, preoperative models from CT or MRI, 20 patients for training and 4 for testing). Point clouds were normalized to a unit sphere with random rigid transformations applied; translations were capped at 10% of point cloud size (≈51 mm on SpineDepth, ≈13 mm on SpineAlign) and rotations at 45°. Training ran on an NVIDIA RTX 5090 GPU (32 GB) with batch size 1, initial learning rate 10⁻⁴, a 10,000-iteration warm-up, and cosine annealing decay. Baseline segmentation for comparison used weak bone labels from points within 3 mm (SpineDepth) or 10 mm (SpineAlign) of the ground-truth-aligned preoperative surface.
Why This Matters
Impact on research: The paper challenges a common assumption in surgical registration pipelines — that segmentation should be trained to match anatomy and then handed to a registration algorithm. By showing that a mask optimized directly for registration generalizes and outperforms anatomically supervised masks on downstream alignment, it suggests that task-specific, label-free intermediate representations may be more useful than accurate ones in other medical and robotics pipelines with heterogeneous sensor data.
Real-world applications:
- Pedicle screw insertion in spine surgery, where misplacement can cause neural or vascular injury and millimeter-level accuracy is required.
- Pediatric and complex spinal deformity correction, the clinical context of the co-authors at the University Children's Hospital Basel, where repeated intraoperative scans carry cumulative radiation burden.
- Reduction of intraoperative radiation dose for both patients and operating room staff by replacing repeated 3D radiographic imaging.
- Removal of workflow friction from bone-anchored fiducials, which can loosen during surgery and require repeated registrations, and from manual overlap-region selection.
- Generalization to other RGB-D-guided orthopedic procedures where an exposed bony surface is captured, such as trauma or joint surgery.
Industry relevance: RGB-D sensors are compact, low-cost, and easy to integrate into the operating room, unlike intraoperative CT or fluoroscopy systems that are expensive and occupy a large footprint. A method that is fully automatic and label-free lowers the annotation cost of building training data for surgical navigation products and reduces the number of manual steps a surgeon must perform. Code and interactive visualizations are released at https://lorenzopettinari.github.io/end-2-reg/.
Future Directions
- Model intervertebral deformation. The authors note SpineDepth lacks preoperative deformation, so the rigid-transformation assumption is not tested against realistic anatomical change.
- Acquire an in vivo dataset with CT-based ground-truth alignment. SpineAlign provides only coarse manual alignment; the current in vivo results are explicitly labeled a preliminary benchmark rather than a definitive comparison.
- Close the loop on the coarse-alignment prior. End2Reg operates without the initialization that CorrNet uses, so a controlled comparison using the same prior would clarify where the remaining in-vivo error comes from.
- Preserve robustness under realistic operating-room occlusion. The up-to-50% occlusion setting is an important step, but instrument and surgeon motion in live cases is dynamic; how the learned mask behaves on unseen occlusion patterns is not reported.
Target Audience
This paper is most valuable to medical computer vision and surgical navigation researchers working on markerless registration, RGB-D-based intraoperative guidance, and point cloud registration under low overlap. It is also relevant to machine learning researchers interested in differentiable discrete sampling and in replacing supervision signals with task-driven objectives. Clinician-engineers and regulatory-adjacent readers in orthopedic and spine surgery will find the clinical framing useful, but the method sections require a solid background in deep learning for 3D data.
Authors’ abstract
Intraoperative navigation in spine surgery demands millimeter-level accuracy. Currently, this is achieved through radiation-intensive intraoperative imaging and bone-anchored markers that are invasive and disrupt surgical workflow. Markerless RGB-D registration methods offer a promising alternative. However, existing approaches rely on weak segmentation labels to isolate relevant anatomical structures, potentially propagating errors through the registration process. We present End2Reg, an end-to-end deep learning framework that jointly optimizes segmentation and registration, eliminating the need for segmentation labels and manual steps. The network learns task-specific segmentation masks optimized for registration, guided solely by the registration objective without explicit segmentation supervision. End2Reg achieves state-of-the-art performance on ex- and in-vivo benchmarks, reducing median Target Registration Error by 32% and mean Root Mean Square Error by 61%, while maintaining robust performance under partial occlusions. Ablation results confirm that end-to-end optimization significantly improves registration accuracy. Overall, End2Reg advances towards fully automatic, markerless intraoperative navigation. Code and interactive visualizations are available at: https://lorenzopettinari.github.io/end-2-reg/.