Research
LuMon: A Comprehensive Benchmark and Development Suite with Novel Datasets for Lunar Monocular Depth Estimation
Overview Research area: Computer vision / monocular depth estimation (MDE), specifically for autonomous lunar rover navigation using electro-optical cameras. Technical level: Advanced. The paper assum

- arXiv
- 2604.09352
- Published
- 2026-04-10
- Authors
- Aytaç Sekmen, Fatih Emre Gunes, Furkan Horoz, Hüseyin Umut Işık, Mehmet Alp Ozaydin, Onur Altay Topaloglu, Şahin Umutcan Üstündaş, Yurdasen Alp Yeni, Halil Ersin Soken, Erol Sahin, Ramazan Gokberk Cinbis, Sinan Kalkan
AI summary
Overview
Research area: Computer vision / monocular depth estimation (MDE), specifically for autonomous lunar rover navigation using electro-optical cameras.
Technical level: Advanced. The paper assumes familiarity with depth-estimation metrics and architectures (foundation models, diffusion-based depth, LoRA fine-tuning, relative pose estimation), though its main arguments are stated in plain terms.
Scope: A benchmarking and development suite that curates six lunar and planetary-analog datasets, adds new ground-truth depth for Chang'e-3 imagery and the new CHERI dark-analog dataset, and systematically evaluates 14 monocular depth estimation models zero-shot plus one domain-adapted baseline.
What This Paper Is About
Terrestrial monocular depth estimation networks work well on Earth, but the Moon presents a severe visual domain gap: harsh high-contrast shadows, textureless regolith, no atmospheric scattering, and no large-scale metric ground truth to train or test against. Existing evaluations lean on Earth analogs that do not replicate these conditions. The paper introduces LuMon, a benchmarking framework with new datasets—including stereo-derived ground truth depth for real Chang'e-3 imagery and a new CHERI dark-analog dataset—to measure how 14 leading depth models actually perform on lunar-like imagery, and to test whether fine-tuning on synthetic data bridges the sim-to-real gap.
Key Contributions
- A benchmark framework and six curated datasets. LuMon evaluates MDE methods on LunarSim and LuSNAR (simulation), Etna-LRNT and Etna-S3LI (Earth analogs), and two new sources: the CHERI dark lunar-analog dataset and newly constructed stereo ground truth depth for real Chang'e-3 mission imagery. Ground truth for CHERI and Chang'e-3 was built via stereo reconstruction, refined with sparse ORB feature correspondences, sub-pixel rectification, and FoundationStereo disparity generation, then validated against known physical dimensions.
- A systematic zero-shot analysis of 14 models. The study spans metric, relative, generative, and video-consistent paradigms, testing robustness against craters, rocks, extreme shading, and varying depth ranges, plus accuracy and efficiency.
- A sim-to-real domain adaptation baseline. DepthAnything v2 (metric-outdoor ViT-L) was fine-tuned on the synthetic LuSNAR dataset using LoRA with rank r = 8, freezing the encoder and updating only the decoder, reducing trainable parameters from 335.3M to roughly 34.1M.
- Downstream task integration and an open-issues audit. Depth predictions were fed into a relative pose estimation pipeline (MADPose with MASt3R correspondences) on the LuSNAR test set to quantify how MDE quality affects robotic autonomy, and the paper catalogues limitations including sensitivity to stereo-rectification black borders and on-board compute constraints.
Main Findings
- Metric foundation models dominate zero-shot. Metric depth estimators consistently outrank relative and diffusion-based models across the benchmark. MapAnything, VDA, and DepthAnything 3 show the most robust cross-domain zero-shot performance; on real Chang'e-3 imagery, DepthAnything 3 reaches δ1 = 0.96, A.Rel = 0.05, RMSE = 0.44, and MapAnything reaches δ1 = 0.95, A.Rel = 0.07, RMSE = 0.52. Because the protocol applies per-image least-squares affine alignment uniformly to all models, this advantage reflects stronger learned geometric priors rather than global scale recovery.
- Earth analogs are deceptive. Most models achieve near-perfect accuracy on the well-lit Etna-LRNT dataset (δ1 = 1.00 for many architectures) because its volcanic terrain resembles terrestrial training data, yet performance plummets on Etna-S3LI with its stricter LiDAR ground truth, and on CHERI under harsh high-contrast polar lighting, where the best metric result is VDA at δ1 = 0.43, A.Rel = 0.44.
- Craters and rocks hurt more than shadows. On LuSNAR, the fine-tuned DAv2-FT dominates topological cases (regolith δ1 = 0.93, rocks 0.92, craters 0.87), but among zero-shot models craters cause severe degradation—MapAnything leads at δ1 = 0.42, A.Rel = 0.26 and UniDepth v2 at δ1 = 0.45, A.Rel = 0.25. Shading causes less overall degradation than complex surface geometry.
- A catastrophic failure case. Depth Pro breaks down completely on the Etna-LRNT dataset, producing A.Rel = 241.71 and RMSE = 450.05 in the overall comparison and A.Rel = 559.37, RMSE = 818.97 in the shaded split.
- Training scale helps only selectively. On authentic Chang'e-3 imagery, training scale correlates significantly with higher accuracy (δ1: ρ = 0.718, p = 0.004) and lower error (RMSE: ρ = −0.619, p = 0.018), but these correlations weaken or turn negative on CHERI and synthetic benchmarks. The ratio of real terrestrial training data does not correlate with improvement (ρ = −0.677, p = 0.008), and no statistically significant correlation was found between metric supervision percentage and zero-shot performance (p > 0.05).
- Fine-tuning wins in simulation, not in the real world. DAv2-FT achieves δ1 = 0.93, A.Rel = 0.07, RMSE = 1.07 on its LuSNAR training domain and transfers strongly to LunarSim (δ1 = 0.84, A.Rel = 0.13, RMSE = 13.96), but only marginally improves on real and analog datasets such as Etna-LRNT, Etna-S3LI, and Chang'e-3. On CHERI it reaches only δ1 = 0.36, A.Rel = 0.50, RMSE = 3.90.
- Distance behaviour separates model families. Relative and diffusion-based models (Lotus, DepthCrafter) show strong non-linear distortions and fail to transfer scale consistently to near and far boundary regions, while metric models (MapAnything, Metric-3D v2) and structurally constrained methods (VDA, MiDaS v3.1) preserve internal linearity. Near-range performance nonetheless degrades sharply on authentic Chang'e-3 and CHERI imagery.
- Efficiency varies enormously. Measured at 1280 × 720 on an Intel Xeon Gold 6326 CPU with an NVIDIA RTX A6000 GPU, DepthAnything-AC runs at 34.24 FPS using 0.509 GB VRAM and 0.130 TFLOPs, while DepthCrafter runs at 0.01 FPS using 20.166 GB and Marigold at 0.02 FPS using 272.109 TFLOPs. The fine-tuned DAv2-FT runs at 4.64 FPS with 1.926 GB VRAM and 1.309 TFLOPs.
- MDE is viable for downstream pose estimation. Using MADPose with MASt3R, DAv2-FT achieves the lowest median rotation error of 0.12 among MDE models, with AUC values of 57.70, 78.05, and 89.03 percent at 5, 10, and 20 degrees, compared with 0.11 and 61.78, 80.25, 90.12 for the MASt3R + ground-truth reference. Metric models consistently provide more geometrically reliable priors than relative ones.
Methodology in Plain English
The authors assembled six datasets that span three levels of realism. Two are photorealistic simulators: LunarSim, a Unity-based simulator providing stereo imagery and pixel-aligned depth from a custom captured trajectory including pose data, and LuSNAR, which contains nine Unreal Engine 4 sequences with stereo images, dense depth, and semantic segmentation masks for five geological classes. Four are real or analog sources: two Mount Etna analog datasets (Etna-LRNT with SGM-generated stereo depth, Etna-S3LI with LiDAR depth), the new CHERI dataset captured in a dark lunar-analog environment, and real Chang'e-3 mission stereo imagery.
Because CHERI and Chang'e-3 had no native depth, the team generated ground truth through stereo reconstruction: they refined geometry using sparse ORB feature correspondences for epipolar consistency, applied sub-pixel rectification, generated disparity maps with FoundationStereo, converted to metric depth, masked invalid pixels, and validated scale against known physical dimensions.
All 14 models were then run under one unified evaluation protocol: predictions were resized to ground-truth resolution, inverse-depth outputs were converted with a numerical stabilization constant of 1 × 10⁻⁶, invalid and out-of-range pixels were masked, and stereo datasets were clipped using KITTI-adapted thresholds. LunarSim remained unclipped because its relative depth lacks an absolute metric scale. A per-image least-squares affine alignment was applied uniformly to every model—including metric ones—so that the evaluation isolates structural fidelity rather than penalizing domain-induced global scale drift. Reported metrics are limited to δ1, AbsRel, and RMSE; MAE and SILog were also computed but not reported due to page limits.
Beyond the overall comparison, the authors ran targeted experiments: semantic-aware analysis using LuSNAR's segmentation labels for regolith, rocks, and craters; shadow analysis using dataset-specific brightness thresholds refined with a 5 × 5 morphological opening; near/far distance analysis; a Spearman rank correlation study relating performance to training scale, real data ratio, and metric supervision ratio; a LoRA fine-tuning baseline; and integration into a MADPose relative pose estimation pipeline with sky regions masked out. All LuSNAR evaluations were conducted strictly on the designated test set, while other datasets were evaluated in their entirety.
Why This Matters
Impact on research. The paper establishes that raw data scaling and metric supervision do not guarantee cross-domain robustness for extraterrestrial perception, and that Earth-based analogs can mislead benchmark designers by inflating apparent performance. It provides a standardized suite—novel ground truth, zero-shot rankings, a pose estimation pipeline, and a fine-tuning baseline—for comparing future methods. Code, datasets, and the benchmark are publicly available at https://metulumon.github.io/.
Real-world applications:
- Autonomous lunar rover navigation, where depth maps feed obstacle avoidance, path planning, and 3D mapping from low-mass electro-optical cameras.
- Relative pose estimation and visual odometry for planetary rovers operating without GPS or reliable global positioning.
- Small-satellite and mini-rover missions that need low-power, low-mass perception pipelines rather than heavy LiDAR or stereo rigs with large baselines.
- Terrestrial spin-offs in GPS-denied and extreme-lighting environments, such as mining, cave exploration, underwater inspection, and disaster response robotics.
Industry relevance. The efficiency table gives space-hardware engineers concrete FPS, VRAM, and TFLOPs figures for choosing between accuracy and on-board feasibility, and the discovery that models fail on black stereo-rectification borders is directly actionable for anyone processing raw mission data. The finding that targeted fine-tuning buys sim-to-sim improvements but not sim-to-real ones informs how mission budgets should be allocated between simulation and real-data collection.
Future Directions
- Architectural innovation over data scaling. Since neither terrestrial data volume nor raw metric supervision guarantees cross-domain robustness, the authors call for models that explicitly decouple illumination from geometric features or integrate physics-based rendering priors.
- Robustness to sensor calibration artifacts. The systematic vulnerability of architectures—notably Depth Pro—to the zero-value black borders produced by stereo rectification needs to be addressed, so networks natively ignore unstructured boundaries rather than requiring cropping.
- Accuracy–efficiency trade-offs on space hardware. All reported evaluations used the most capable weights of each architecture as an upper bound; the authors recommend using LuMon to identify or distil lightweight architectures suitable for power- and memory-constrained rovers.
- Closing the sim-to-real gap. The persistent stagnation of the LoRA-adapted DepthAnything v2 on authentic lunar imagery raises the open question of what combination of real mission data, diverse analogs, and adaptation techniques would actually transfer to the Moon.
Target Audience
Robotics and computer vision researchers working on extraterrestrial perception, domain adaptation, and depth estimation; planetary rover mission engineers evaluating on-board perception options; and benchmark designers interested in how terrestrial proxies can distort evaluation. The paper is most useful to readers with some background in depth estimation metrics and model families, since much of its evidence is presented in dense comparison tables, though the framing arguments about the sim-to-real bottleneck are accessible to a broader space-robotics audience.
Authors’ abstract
Monocular Depth Estimation (MDE) is crucial for autonomous lunar rover navigation using electro-optical cameras. However, deploying terrestrial MDE networks to the Moon brings a severe domain gap due to harsh shadows, textureless regolith, and zero atmospheric scattering. Existing evaluations rely on analogs that fail to replicate these conditions and lack actual metric ground truth. To address this, we present LuMon, a comprehensive benchmarking framework to evaluate MDE methods for lunar exploration. We introduce novel datasets featuring high-quality stereo ground truth depth from the real Chang'e-3 mission and the CHERI dark analog dataset. Utilizing this framework, we conduct a systematic zero-shot evaluation of state-of-the-art architectures across synthetic, analog, and real datasets. We rigorously assess performance against mission critical challenges like craters, rocks, extreme shading, and varying depth ranges. Furthermore, we establish a sim-to-real domain adaptation baseline by fine tuning a foundation model on synthetic data. While this adaptation yields drastic in-domain performance gains, it exhibits minimal generalization to authentic lunar imagery, highlighting a persistent cross-domain transfer gap. Our extensive analysis reveals the inherent limitations of current networks and sets a standard foundation to guide future advancements in extraterrestrial perception and domain adaptation.