Research
C3VDReg: A Benchmark for Local-to-Local Colonoscopic Registration toward Anatomical Localization
Overview Research area: Medical computer vision, specifically 3D point cloud registration for colonoscopy, with an emphasis on benchmark design and evaluation methodology. Technical level: Advanced. T

- arXiv
- 2511.00260
- Published
- 2025-10-31
- Authors
- Linzhe Jiang, Jiayuan Huang, Sophia Bano, Matthew J. Clarkson, Zhehua Mao, Mobarak I. Hoque
AI summary
Overview
Research area: Medical computer vision, specifically 3D point cloud registration for colonoscopy, with an emphasis on benchmark design and evaluation methodology.
Technical level: Advanced. The paper assumes familiarity with point cloud registration families (ICP, PointNetLK, DCP, RegTR, GeoTransformer), rigid pose metrics (RRE, RTE, registration recall), and point cloud construction from depth maps and meshes.
Scope: The paper introduces C3VDReg, a dataset and fixed evaluation protocol for rigid, viewpoint-matched, partial-to-partial point cloud registration between colonoscopic video depth reconstructions and CT-derived colon surfaces, and uses it to show that translation ambiguity, not low overlap, is the dominant failure mode.
What This Paper Is About
During colonoscopy the endoscopist sees only a small local patch of the mucosal surface, and linking that live view to a stable 3D anatomical reference (for example a preoperative CT-derived colon model) could support coverage assessment and revisited-region awareness. The paper argues that directly benchmarking this as a local-to-global search problem entangles two hard subproblems: finding the right region on a long, repetitive CT surface, and then aligning the observed patch to that region. C3VDReg isolates the second subproblem by building viewpoint-matched, local-to-local point cloud pairs from released C3VD depth maps, camera poses, and CT meshes, and then measures whether current registration algorithms can solve that controlled task.
Key Contributions
-
A new dataset and benchmark (C3VDReg). It contains 10,015 viewpoint-matched partial-to-partial point cloud pairs, generated by reprojecting each C3VD depth map into a source cloud and raycasting the CT mesh from the matched camera pose to produce the target cloud.
-
A fixed evaluation contract. The protocol specifies 8,192 points per source and target cloud, source-only perturbations of R25–90deg/T100–500mm, a source-to-target pose convention, strict registration recall at 5deg/5mm (RR@5) and 10deg/10mm (RR@10), plus rotation/translation error and geometry diagnostics, so that methods cannot redefine the task through custom preprocessing, splits, or metrics.
-
A systematic evaluation of classical and learning-based baselines. ICP, PointNetLK, PointNetLK Revisited, PointNetLK-Mamba, DCP, RegTR, and GeoTransformer are all run under the same contract, with runtime and memory profiled as a diagnostic.
-
A failure-mode analysis identifying translation ambiguity as the bottleneck. The paper shows that high ground-truth overlap does not guarantee successful registration, and decomposes translation error into axial (along the camera trajectory) and radial (across the wall) components.
Main Findings
-
The benchmark is far from saturated. GeoTransformer is the strongest baseline but reaches only 17.43% RR@5 and 35.63% RR@10. DCP records 0.00% at both thresholds. ICP reaches 5.46% RR@5 and 10.73% RR@10, RegTR 5.32% and 11.59%, PointNetLK-Mamba 3.16% and 11.11%, PointNetLK Revisited 2.39% and 10.58%, and PointNetLK 0.14% and 1.01%.
-
Rotation is recovered far more often than translation. GeoTransformer places 56.18% of pairs within 5deg rotation error but only 18.15% within 5mm translation error. The same rotation-over-translation gap appears for RegTR (18.87% vs 5.41%), ICP (21.12% vs 5.70%), PointNetLK-Mamba (24.28% vs 3.64%), PointNetLK Revisited (30.32% vs 2.78%), and the others.
-
Aggregate recall hides unstable behaviour. ICP's 114/2088 strict RR@5 successes come with a median RTE of 123.72mm and a 90th percentile RTE of 570.22mm, indicating convergence on a small subset rather than global robustness. GeoTransformer's median RTE is 15.86mm with a 90th percentile of 82.12mm, and median RRE 4.29deg with a 90th percentile of 17.16deg.
-
Performance is strongly anatomy dependent. GeoTransformer reaches 63.04% RR@5 on the held-out scene
cecum_t1_abut only 0.93% onsigmoid_t3_b. -
High ground-truth overlap does not eliminate failure. All 2,088 test pairs have high overlap after ground-truth alignment, with equal-size quantiles spanning 74.3–78.1%, 78.1–83.8%, 83.8–87.8%, and 87.9–93.1%. Strict RR@5 is nonmonotonic across these quantiles: GeoTransformer and RegTR improve in Q3 but fall in Q4, and DCP stays at zero.
-
Translation errors have long tails in both axial and radial directions. GeoTransformer's 90th percentile errors are 42.62mm axial and 64.77mm radial, with an axial fraction of 0.50. The axial fraction stays close to 0.5 across learned methods (0.45 for PointNetLK-Mamba, 0.50 for RegTR, 0.52 for PointNetLK Revisited, 0.50 for PointNetLK, 0.51 for DCP), meaning both sliding along the lumen and mislocalization across the wall occur.
-
Local residual fit can be misleading. In one GeoTransformer case the nearest-neighbor p90 residual is only 5.40mm while the RTE is 1004.0mm; in another, residuals and pose error are both poor (NN p90 18.57mm, RTE 911.2mm).
-
Recall alone is a poor model-selection signal. RegTR has higher RR@5 than PointNetLK-Mamba in the main table, yet PointNetLK-Mamba shows smaller translation tails and less drift on the same test pair.
-
There is a robustness/efficiency tradeoff. DCP is the fastest learned baseline at 64.3ms latency but has zero strict successes; RegTR is faster than GeoTransformer (83.0ms vs 186.8ms latency) but much less accurate; ICP is CPU-only at 787.7ms latency.
Methodology in Plain English
The authors did not collect new patient data. They built everything from the publicly released Colonoscopy 3D Video Dataset (C3VD), which provides colonoscopy videos along with depth maps, camera poses, surface models, and coverage annotations.
For each C3VD frame they construct two point clouds. The source cloud comes from the released depth map: valid pixels are unprojected using the camera intrinsics and transformed into the global C3VD coordinate frame. The target cloud comes from the CT mesh: a virtual camera is placed at the same released pose and rays are cast into the mesh, keeping the nearest intersection per image-plane sample using the Moller–Trumbore algorithm. This yields pairs that share anatomy but differ in sensing process, density, and noise, since one side is CT-derived geometry and the other is video-depth reconstruction.
The target cloud is deliberately an "oracle" view of the visible CT surface. This removes full-surface retrieval from the main task, so the benchmark measures only whether a method can align corresponding local observations. The authors report a full-surface diagnostic with GeoTransformer to show why this separation matters: against the full CT surface, the same source frames produce severe pose errors.
Before evaluation the protocol is frozen: 8,192 points per cloud, perturbations of R25–90deg and T100–500mm applied only to the source, an unperturbed target, and a source-to-target pose convention. Predictions are scored with relative rotation error (geodesic angle in degrees), relative translation error (Euclidean, in millimetres), joint recall at 5deg/5mm and 10deg/10mm, marginal rotation-only and translation-only hit rates, medians, 90th percentiles, and empirical CDFs. Geometry diagnostics (visible nearest-neighbor distance and 90% trimmed Chamfer distance) are computed after applying the predicted transform. Ground-truth overlap is computed independently of any prediction, using a 5mm radius and 2048 deterministic points per cloud, and is used only to stratify the fixed test set. Target registration error (TRE) is not reported because there are no independent landmark targets.
The 10,015 pairs split into 5,952 training pairs, 1,975 validation pairs, and 2,088 test pairs. The test set is defined at scene level and comprises six held-out C3VD sequences: cecum_t1_a (276), cecum_t3_a (730), sigmoid_t3_b (536), trans_t1_a (61), trans_t2_b (103), and trans_t4_a (382).
Why This Matters
Impact on research. The paper challenges a common assumption in registration work: that improving overlap or local correspondence quality is sufficient for accurate registration. It shows that on repetitive tubular anatomy, a prediction can look locally plausible yet be globally wrong, and that strict recall is nonmonotonic across overlap quantiles. It also supplies a reproducible protocol for a setting that medical registration literature has argued should make modality assumptions, coordinate conventions, and success criteria explicit rather than hidden in implementation details.
Real-world applications.
- Anatomy-aware navigation. Linking each live endoscopic view to a stable 3D reference could help endoscopists understand where they are within the colon.
- Coverage assessment. Knowing which mucosal regions have been observed could support systematic inspection and reduce missed areas.
- Revisited-region awareness. Detecting that the scope has returned to a previously seen location could reduce redundant inspection.
- Lesion site contextualization. Recording a polyp or lesion position in a common coordinate system would support documentation and follow-up.
Industry relevance. Colonoscopy is central to colorectal cancer detection and surveillance. Any system that augments guidance with preoperative CT anatomy will depend on the rigid partial-to-partial alignment step this benchmark isolates. The runtime and memory profiling in the paper also sets a practical constraint for near-online use: the fastest learned baseline (DCP, 64.3ms) has zero strict successes, while the most robust (GeoTransformer, 186.8ms) is slower, leaving an open engineering gap. Code, model checkpoints, and the dataset are released publicly, including at https://github.com/linzhe001/C3VDReg, with the C3VD-Raycasting-10k dataset at https://doi.org/10.5522/04/30640043 and checkpoints at https://huggingface.co/linzher/C3VDReg-checkpoints.
Future Directions
-
Constraints beyond local patch matching. The authors suggest sequence context, lumen topology, anatomical priors, retrieval-aware registration, and uncertainty-aware alignment with multiple hypotheses as directions that could address the translation ambiguity their analysis exposes.
-
Moving toward full local-to-global localization. The raycast target supplies an oracle visible CT surface, so visible-region selection and retrieval on the full CT surface remain unsolved. The full-surface diagnostic with GeoTransformer indicates this is substantially harder.
-
Extending to non-rigid and realistic conditions. The benchmark is rigid, viewpoint-matched, and limited to partial source and partial target clouds. It does not evaluate nonrigid deformation, unmatched viewpoints, debris, fluid, tissue motion, insufflation changes, in vivo workflow variability, or sequence-level trajectory estimation.
-
Broadening the baseline set. The authors note that additional robust classical methods and alternative retrieval-oriented pipelines remain to be studied, and that the axial/radial analysis currently relies on a trajectory-derived axis from C3VD camera centers rather than a manually annotated anatomical centerline.
Target Audience
Researchers and engineers working on point cloud registration, medical image registration, and surgical/endoscopic computer vision will benefit most. The paper is also relevant to groups building colonoscopy datasets and evaluation protocols, since its main argument concerns how a benchmark contract should be specified. Clinically oriented readers interested in anatomy-aware navigation and coverage assessment will find the framing useful, though the technical content (RRE/RTE metrics, registration recall, axial/radial decomposition, CUDA memory profiling) assumes a computer vision background.
Authors’ abstract
Anatomy-aware colonoscopic navigation requires localizing partial endoscopic observations on a stable 3D reference to support coverage assessment, revisited-region awareness, and CT-guided navigation. However, rigid point cloud registration in the colon differs fundamentally from standard benchmarks: surfaces are locally homogeneous, haustral folds are repetitive, views are highly partial, and reconstructed depth is noisy. We present C3VDReg, a dataset and benchmark derived from the Colonoscopy 3D Video Dataset (C3VD). For each frame, C3VDReg generates source point clouds via depth reprojection and target point clouds by raycasting CT meshes from matched camera poses. The benchmark comprises 10,015 viewpoint-matched partial-to-partial point cloud pairs (including 2,088 held-out test pairs) and evaluates baseline models under a standardized protocol: 8,192 points per cloud, source-only perturbations, fixed pose conventions, and unified metrics. Crucially, C3VDReg enables a systematic investigation of failure modes. We find that high geometric overlap alone is insufficient for reliable pose recovery: despite 74.3-93.1% ground-truth overlap, registration recall remains low across all evaluated methods. Through overlap, pose error, and translation decomposition analyses, we identify translation ambiguity along repetitive tubular anatomy as the primary bottleneck. This challenges the common assumption that increasing overlap or correspondence quality guarantees accurate registration, highlighting the need for stronger anatomical and contextual constraints. Code, model checkpoints, and data are available at https://github.com/linzhe001/C3VDReg .