Research
Pumpire: Unified Benchmark for Metric Distance Estimation
Pumpire: Unified Benchmark for Metric Distance Estimation Overview Research area: Computer vision / 3D scene reconstruction and evaluation benchmarks for metric 3D foundation models. Technical level:
- arXiv
- 2610.12423
- Published
- 2026-10-08
- Authors
- Siyu Chen, Zehan Wang, Jiayang Xu, Yihan Wu, Jialei Wang, Junming Chen, Ziang Zhang, Yutong Ying, Zhou Zhao
AI summary
Pumpire: Unified Benchmark for Metric Distance EstimationOverview
- Research area: Computer vision / 3D scene reconstruction and evaluation benchmarks for metric 3D foundation models.
- Technical level: Advanced (assumes familiarity with depth estimation, camera intrinsics, back-projection, and point clouds).
- Scope: A single benchmark, dataset, and metric suite that directly measures how accurately image- and video-level 3D foundation models recover the physical distance between two annotated points in a real scene.
What This Paper Is About
Existing evaluations test depth maps and camera intrinsics separately, or score reconstructed point clouds with geometric similarity metrics such as Chamfer Distance and F1-Score, neither of which directly reports whether a model can measure the real-world distance between two arbitrary points. The authors build Pumpire: a dataset of 100 real-world scenes with tape-measured point-pair distances, plus a protocol that back-projects model predictions and compares the resulting 3D distance against that physical ground truth. The goal is to evaluate the fundamental ability to perceive and estimate physical scale in reconstructed 3D space, which prior protocols have largely overlooked.
Key Contributions
- First unified benchmark for direct metric point-pair distance evaluation, enabling comparison across image-level and video-level 3D foundation models in both estimation (RGB-only) and completion (depth-prior-conditioned) settings.
- The pumpire-6k dataset: 100 diverse real-world scenes (40 indoor, 60 outdoor), each contributing 64 uniformly sampled frames from 200–300-frame captures, for 6,400 total frames at 1280×720, annotated with physically measured point-pair distances (tape measure) and per-frame sticker pixel coordinates.
- A metric suite and split strategy: Absolute Error (AE), Relative Error (RE), Logarithmic Error (LE), threshold accuracy (δ_τ), and Relative Standard Deviation (RSD) for cross-view consistency, evaluated separately on normal (76) and challenging (24) scenes split by a sensor-derived RSD threshold of 0.1.
- Comprehensive evaluation of 29 baseline configurations (8 image-level estimation, 10 image-level completion, 6 video-level estimation, 5 video-level completion; grouped as 26 method entries in Appendix B) with six empirical findings.
Main Findings
- Sensor baseline is strong on normal scenes, weak on challenging ones: Raw RealSense D435 depth achieves δ_1.05 = 79.32, RE = 0.038 on normal scenes (N=76) but δ_1.05 = 21.03, RE = 8.390 on challenging scenes (N=24). Overall: AE = 0.991, RE = 2.043, LE = 0.322, RSD = 0.260, δ_1.05 = 65.33, δ_1.10 = 76.97, δ_1.25 = 83.80.
- Normal-scene image-level estimation winner: Metric3D v2 has the best average rank (1.00) on normal scenes, followed by Unidepthv2 (2.43) and MetricAnything (3.43). On challenging scenes, Depth Pro (2.86) and MoGe-3 (3.00) rank first and second while Metric3D v2 drops to a rank of 4.00, showing that strong normal-condition performance does not generalize.
- Video-level estimation winner: Depth Anything 3 ranks first on both subsets (rank 1.00 normal, 1.29 challenging), with AE = 0.078, RE = 0.108 on normal scenes. On challenging scenes, pi3x improves from fourth to second (rank 2.14) whereas MUSt3R drops from third to fifth (4.86).
- Image-level completion winner: Lingbot-Depth ranks first on both subsets (rank 1.00 normal, 1.14 challenging), with AE = 0.022, RE = 0.029, δ_1.05 = 84.52 on normal scenes. The ordering of other methods shifts sharply: CDM-L515 and CDM-D435 move from eighth and seventh on normal scenes to second and third on challenging scenes, while Omni-dc and LDCM drop.
- Video-level completion winner: The three CAPA variants take the top three ranks on both subsets, with CAPA-VGGT first (rank 1.29 normal, 1.43 challenging). On challenging scenes CAPA-MoGe2 achieves the lowest AE, RE, and RSD, whereas CAPA-VGGT achieves the lowest LE and the highest threshold accuracies.
- Finding 1 — Separate depth and intrinsics benchmarks are insufficient: Rankings diverge across tasks: Unidepth ranks best on depth (2.08) but only fourth on intrinsics (3.46), while MoGe-2 shows the opposite trend (2.83 depth, 1.00 intrinsics). MoGe-2 has the best combined rank (1.92) yet is outperformed by Unidepthv2 on Pumpire. Replacing predicted intrinsics with ground-truth intrinsics during back-projection degrades performance in most cases.
- Finding 2 — Learned models beat sensors under hard conditions: On normal scenes only 5 of the 18 baseline configurations approach or surpass raw RealSense D435 δ_1.05, all of them completion models (Lingbot-Depth, PriorDA, Omni-dc, LDCM, InfiniDepth), while the best estimation model reaches barely half that accuracy. On challenging scenes, 16 of the 18 configurations achieve a lower RE than the sensor.
- Finding 3 — Different robustness patterns: Challenging-to-normal RE ratios do not overlap: estimation models stay within 1–20×, while completion models degrade by 20–210×. Even so, completion models still attain the highest δ_1.05 on both subsets; they are precise on most frames but suffer catastrophic outliers on optically and structurally degraded regions, whereas estimation models trade peak accuracy for uniformly bounded errors.
- Finding 4 — Denser depth priors are not always better: Using 100, 1,000, or 10,000 randomly sampled depth pixels (with the full depth map as reference), increasing the number of depth points does not consistently improve completion accuracy, and several models perform best at intermediate densities. Sparsity and sensor noise are two distinct challenges that jointly determine performance.
- Finding 5 — More views with priors helps video-level completion: Sweeping the fraction of views with depth priors over ρ ∈ {0.1, 0.2, 0.3, 0.5, 0.7, 1.0}, RE decreases consistently for both MapAnything and pi3x, with a substantially larger improvement for MapAnything and a more gradual reduction for pi3x.
- Finding 6 — Combined setting performs best: Comparing best models per setting on normal scenes, image-level estimation (Metric3D v2) and video-level estimation (Depth Anything 3) have comparable RE (0.106 vs. 0.109) but the video model has substantially lower RSD (0.067 vs. 0.029). Image-level completion (Lingbot-Depth) achieves substantially lower RE than the estimation models with an RSD comparable to video-level models (0.035 vs. 0.028 and 0.029), and video-level completion (CAPA-VGGT) achieves the lowest RE and RSD overall.
Methodology in Plain English
The authors avoid trusting sensor depth as ground truth: sensor-based distances fail on reflective or transparent surfaces, thin structures and object boundaries, and ordinary surfaces affected by noise, while synthetic data miss the real-world domain gap. Instead, they place two circular stickers (10 mm radius) in each scene, measure the distance between their centers with a tape measure, and record the sticker pixel coordinates in every frame. Coordinates are annotated with a coarse-to-fine, human-in-the-loop pipeline: manual clicks on an anchor frame, CoTracker3 propagation, re-annotation when tracking drift exceeds tolerance, and final frame-by-frame pixel-level refinement.
For evaluation, each model's depth map (or the z-coordinates of its predicted point map) is back-projected using camera intrinsics and the annotated pixel coordinates, and the Euclidean distance between the two resulting 3D points is compared against the measured ground truth. Estimation models use their own predicted intrinsics (canonical intrinsics for Metric3D and Metric3D v2); completion models conditioned on sensor depth use ground-truth intrinsics. Scenes are split into normal and challenging groups using the cross-view variation (RSD) of distances from raw sensor depth with an empirical threshold of 0.1. Video-level models are evaluated on 32 groups of 10 frames per scene, with identical sampling and a fixed random seed across all models.
Why This Matters
Impact on research. The paper shows that strong scores on depth and intrinsics benchmarks — even considered jointly — do not predict point-pair distance accuracy, and that point-cloud similarity metrics do not capture it either. This argues for direct, task-level evaluation of physical scale and gives the community a shared protocol for comparing model families that existing fragmented benchmarks could not compare side by side.
Real-world applications:
- Autonomous driving: translation errors between 3D detections and ground truth are associated with collision rates in the paper's cited work (Taamazyan et al., 2024; Caesar et al., 2020).
- Robotic manipulation: grasp success rates depend on 3D spatial measurements, not pixel-level depth accuracy alone (Hietanen et al., 2021; Hu et al., 2026).
- Vision-language models and spatial reasoning: limitations in recovering metric distances and object scales remain a bottleneck for spatial reasoning and physical task execution.
- Sensor-degraded environments: the normal/challenging split quantifies when active depth sensing breaks down and learned models become the more reliable option.
Industry relevance. The completion-versus-estimation trade-off is directly actionable: completion models give higher peak accuracy but can fail catastrophically by 20–210× on challenging surfaces, while estimation models keep bounded errors. Systems that need guaranteed worst-case behavior may prefer estimation; systems that mostly operate on benign surfaces may prefer completion, and combining multi-view context with depth priors (CAPA-VGGT) performs best overall.
Future Directions
- Mitigating catastrophic completion failures: completion models degrade by 20–210× on challenging scenes; how to bound these outlier errors without sacrificing peak accuracy is open.
- Jointly handling sparsity and sensor noise: since denser priors do not reliably help, models that explicitly separate the two failure sources are a natural next step.
- Better cross-view metric integration: MapAnything benefits far more from added depth priors than pi3x, so the ability to fuse metric cues across views is not yet uniform across models.
- Broadening the benchmark: the dataset covers 100 scenes (40 indoor, 60 outdoor) and 6,400 frames; whether these findings hold for other sensors, scales, or scene types remains untested within the paper.
Target Audience
Researchers and engineers working on metric 3D foundation models, monocular and multi-view depth estimation, depth completion, and spatial reasoning systems; benchmark designers who need a protocol for physical scale; and applied teams in robotics, autonomous driving, and augmented reality who must choose between RGB-only estimation and depth-prior-conditioned completion for metric distance tasks.
Authors’ abstract
We present Pumpire, a unified benchmark for evaluating metric point-pair distance estimation capability of both image- and video-level 3D foundation models, with or without depth priors. In contrast to previous approaches that normally evaluate depth and camera intrinsics separately or evaluate point-clouds with geometric similarity metrics, which cannot directly reflect models' point-to-point distance estimation capability, Pumpire directly assesses point-to-point distances from the reconstructed geometry. To this end, we collect a large-scale and diverse dataset (pumpire-6k) comprising 100 real-world scenes, each annotated with physically measured point-pair distances and containing 64 frames, for a total of 6,400 frames. Building on this dataset, we establish a holistic evaluation protocol that covers both image- and video-level 3D foundation models and enables direct assessment of point-pair distance errors and cross-setting comparison. We conduct extensive experiments across 29 baseline configurations of representative 3D foundation models and provide a comprehensive analysis of the results. By offering this benchmark, we target the more fundamental ability to perceive and estimate physical scale in the reconstructed 3D space, which prior evaluation protocols have largely overlooked. The project page can be found at https://pumpire.github.io/