Skip to content
AI.info

Research

The 3D Mirage: Probing and Taming 3D Hallucinations

Overview Research area: Computer vision, specifically monocular depth estimation (MDE) and 3D perception robustness, with connections to autonomous driving safety. Technical level: Advanced. The paper

The 3D Mirage: Probing and Taming 3D Hallucinations
arXiv
2512.15423
Published
2025-12-17
Authors
Hoang Nguyen, Xiaohao Xu, Xiaonan Huang

AI summary

Overview

  • Research area: Computer vision, specifically monocular depth estimation (MDE) and 3D perception robustness, with connections to autonomous driving safety.
  • Technical level: Advanced. The paper assumes familiarity with depth foundation models (Depth-Anything-V2, MiDaS, Marigold), transformer encoders, second-order differential operators, and parameter-efficient fine-tuning such as LoRA.
  • Scope: One sentence: the paper defines, benchmarks, quantifies, and then mitigates a failure mode it calls the "3D Mirage," in which monocular depth foundation models hallucinate nonexistent 3D structure from flat or low-curvature surfaces carrying illusory texture, especially when surrounding context is removed.

What This Paper Is About

Monocular depth foundation models appear to generalize well because they learn large-scale semantic priors, but those same priors cause them to invent 3D geometry that is not there. The authors show that when a model looks at 3D street art, chalk anamorphoses, forced-perspective murals, or large-format advertisements, it reads the painted illusion as real depth, and that this failure gets dramatically worse when the field of view is restricted so broad context is missing. The paper builds a benchmark, proposes two new reference-free scores for the failure, and presents a parameter-efficient training recipe that removes the hallucination without damaging the model's general depth knowledge.

Key Contributions

  1. A benchmark for probing the failure (3D-Mirage): The first benchmark that combines context variation with precise annotation for real-world illusions, including real object exclusions, support for multiple illusions per scene, and illusions that span surfaces. It contains 468 real-world RGB images of painting and street-art 3D illusions and 1,872 annotated full-crop instances.
  2. A new way to score the failure: A second-order magnitude-based evaluation framework with two metrics — the Deviation Composite Score (DCS), measuring spurious second-order 3D structure (hallucination intensity), and the Confusion Composite Score (CCS), measuring contextual instability (the mirage effect). Both are reference-free and require no ground-truth depth.
  3. A mitigation strategy (Grounded Self-Distillation, or GSD): A parameter-efficient adaptation that inserts low-rank adapters into the model's encoder to surgically suppress hallucinated curvature inside illusion regions of interest (ROIs), while a frozen teacher model enforces alignment on stable background and border regions to avoid catastrophic forgetting.
  4. A shift in evaluation philosophy: The authors argue that standard metrics such as Mean Absolute Error (MAE) and Root Mean Squared Error (RMSE) are perceptually blind to structural failures, and that MDE evaluation should move from pixel-wise accuracy toward structural and contextual robustness.

Main Findings

  • The failure is systemic, not anecdotal: All tested foundation models showed elevated DCS and CCS on 3D-Mirage — transformer-based (Depth-Anything-V2), diffusion-based (Marigold), generative (DepthFM), and commercially developed (Depth Pro), plus ZoeDepth and MiDaS. The baseline Depth-Anything-V2-Large (DAv2-L) scored a DCS of 994.6, confirming it perceives significant spurious 3D geometry.
  • The spread across models is large: On DCS, DepthFM reached 2.083e3 and Marigold 1.427e3, while DepthPro reached 649.1, MiDaS 670.2, ZoeDepth 589.3, DA-S 459.8, DA-B 482.7, DA-L 495.0, DAv2-S 1.007e3, DAv2-B 881.0, and the indoor/outdoor DAv2 specialists ranged from 706.9 (DAv2-IB) to 1.439e3 (DAv2-OB).
  • GSD reduces hallucination by roughly an order of magnitude: The method achieves DCS 58.64 and CCS 1.907e-4, versus 994.6 and 1.466e-3 for its DAv2-L teacher — a 94.10% reduction in DCS and 86.99% reduction in CCS (component reductions of 94.16% for d_cluster, 94.05% for d_avg, 86.59% for D_cluster, and 87.35% for D_avg).
  • The metrics are robust to artifacts: An image-edited, non-illusory counterpart at matched scene context scored DCS ≈ 54 and CCS ≈ 1.4 × 10⁻⁴, and adding Gaussian noise lifted DCS to at most ≈ 65, whereas the illusion reaches DCS ≈ 861 and CCS ≈ 7.4 × 10⁻⁴ — roughly a 16× DCS gap.
  • Corrections are surgically confined to the illusion: Error heatmaps show the model's changes are limited to the ROI, which the authors present as evidence that the non-hallucination knowledge preservation loss prevents the flattening objective from destroying valid background geometry.
  • The recipe transfers across architectures: On ZoeDepth, DCS dropped from 589.27 to 335.12 and CCS from 1.505e-3 to 1.136e-3; on 3DVI metrics ZoeDepth's EPE improved from 9.555 to 7.286, bad2 from 67.66 to 61.70, AbsRel from 0.1899 to 0.1273, RMSE from 0.2255 to 0.1400, and δ₁ from 76.16 to 84.76. On the Marigold diffusion backbone, a LoRA (r = 16) adaptation cut DCS and CCS by 44% and 40% while preserving DA-2K/DIW depth.
  • Zero-shot generalization to a separate illusion dataset: Evaluated on the 3DVI test split from Yao et al. (NeurIPS'25), the method reached EPE 1.75, bad2 26.67, bad3 15.52, bad5 6.59, AbsRel 0.03, RMSE 0.06, and δ₁ 99.50 — competitive with stereo and multi-view systems (Selective-RAFT: EPE 1.58, bad2 23.46, δ₁ 99.60) while using only monocular input.
  • Ablation results (partially reported): Removing hallucination knowledge re-editing gives DCS 988.60, removing knowledge preservation gives DCS 42.83, and the full method gives DCS 58.64. The corresponding CCS row is truncated in the available content, so the complete ablation comparison is not reported here.
  • Taming does not have to flatten curves: The plane-mixture and gating design means the model does not force a single fronto-parallel plane; in some cases it partially recovers a genuinely curved surface rather than treating everything as flat.
  • Small training budget suffices: Training for one epoch on a single NVIDIA A100 with batch size 8 produces the reported results; extending to 4 epochs yields only marginal gains on DCS/CCS, while other evaluation metrics show mixed behavior.

Methodology in Plain English

The authors attacked the problem in three stages.

Probe. They gathered 468 real-world photographs of painted and street-art 3D illusions — about 80% outdoor, 40% on pedestrian walkways, 8% billboard or advertisement-like, 16% spanning multiple support surfaces, and about 40% containing real objects or people near or on top of the illusion. Human annotators drew precise polygon masks around each illusion region, using nested polygons to carve out real objects that happen to sit inside the illusion. To simulate the limited fields of view and partial occlusions of driving scenes, they generated four random crops per image, each centered on the illusion with the ROI covering at least 40% of the crop area. This yields 1,872 full-crop instances. Illusion ROIs cover 49% of the image on average, while the crops cover 41% of the original image on average.

Score. For each instance, the same model is run twice: once on the full image and once on the crop. The full-image prediction is warped into the crop's coordinate frame so the two can be compared on identical pixels. Inside the ROI, depth values are quantile-normalized and passed through a second-order differential operator that responds to curvature rather than absolute depth. From this response map, two scalars are computed per branch: the sum of the largest 10% of responses, and a 10%-trimmed high-response mean. DCS combines the mean response vector across the dataset with the average per-sample response magnitude, capturing how much spurious structure the model invents. CCS projects the response difference between full and crop views onto the off-diagonal direction [1, −1], capturing how much the model's answer changes when context is removed. Both metrics are lower-is-better and require no ground truth.

Tame. Grounded Self-Distillation starts from a frozen Depth-Anything-V2-Large teacher. LoRA adapters (rank r = 16, α = 32, dropout 0.05, bias = none) are injected into the DINOv2 encoder's patch embedding layer and all MLP linear layers (fc1, fc2) across the 24 transformer blocks — only about 4M trainable parameters, roughly 1.2% of the DAv2-L backbone — while the decoder stays frozen. Training is dual-view with shared weights: one branch sees the full image, the other the crop. In each branch, the prediction is split into the illusion ROI and the non-illusion background plus an ROI-adjacent ring. The hallucination knowledge re-editing loss pushes second-order structure inside the ROI toward zero and re-targets each ROI toward locally fitted surface hypotheses derived from the teacher's ring neighborhood, using a small gating network over K hypotheses with a cross-entropy regularizer. The non-hallucination knowledge preservation loss distills the frozen teacher's stable geometry onto background and ring regions, including a smoothing term on the low-gradient seam and second-order matching on the high-gradient edge subset and a protective guard ring. Illusion batches are interleaved 4:1 with non-illusion images from Penn–Fudan (170 urban street images with 345 upright pedestrians) and CamVid (701 raw still frames of urban driving scenes) to avoid over-flattening. Augmentations are 50% horizontal flip and 5% photometric jitter. Optimization uses AdamW at learning rate 1 × 10⁻⁴, weight decay 0.01, and gradient clipping of 1.0, with fixed loss weights α₁ = 1.0, α₂ = 0.4, α₃ = 1.0, α₄ = 0.5, α₅ = 0.3, α₆ = 0.8, and α₇ = 0.3. The data is split at the source-image level (90/10) so crops from the same original image never span train and test: 421/47 scenes and 1684/188 paired full-crop instances. Knowledge preservation is checked with an ordinal pairwise accuracy protocol on NYU-v2.

Why This Matters

Impact on research. The work argues that the MDE field's evaluation culture — built around MAE, RMSE, REL, and scale-invariant objectives — cannot see structural hallucinations because pixel-wise averaging dilutes ROI-specific failures and treats each view independently. By introducing reference-free, second-order, dual-view metrics, the paper provides a template for measuring robustness properties that ground-truth depth benchmarks are not designed to capture. It also extends parameter-efficient fine-tuning into 3D hallucination mitigation, which the authors describe as a first.

Real-world applications:

  • Autonomous driving, where a hallucinated obstacle on a flat road could trigger unnecessary braking or evasive maneuvers, and where restricted fields of view are the norm rather than the exception.
  • Robotics and mobile navigation, where a robot relying on monocular depth could misjudge floor planarity or attempt to climb painted features.
  • Augmented reality and 3D reconstruction, where phantom geometry from advertisements, murals, or printed patterns would corrupt scene meshes and object placement.
  • Safety auditing and regression testing of depth models deployed in consumer devices, using the benchmark and metrics as a stress test rather than an average-case accuracy check.

Industry relevance. Every company shipping monocular depth — automotive perception teams, phone and camera makers, AR platforms, and robotics vendors — depends on models like Depth-Anything-V2, ZoeDepth, MiDaS, Marigold, DepthFM, and Depth Pro, all of which the paper shows failing on these inputs. Because GSD trains only about 4M parameters for one epoch on one A100, it is a practical, low-cost adaptation rather than a full retraining program, and the authors demonstrate it transfers to a second transformer backbone (ZoeDepth) and a diffusion backbone (Marigold).

Future Directions

  • Extending the benchmark beyond painting and street art. The current 468 images focus on painted illusions; other ambiguity sources such as reflections, transparent surfaces, and textureless regions are noted as adjacent but not covered.
  • Resolving the ablation tension. The partial ablation shows the full method at DCS 58.64 versus 42.83 without knowledge preservation,

Authors’ abstract

Monocular depth foundation models achieve remarkable generalization by learning large-scale semantic priors, but this creates a critical vulnerability: they hallucinate illusory 3D structures from planar/low-curvature but perceptually ambiguous inputs. We term this failure the 3D Mirage. This paper introduces a novel end-to-end framework to probe, score, and tame this under-quantified safety risk in monocular depth under context variation. To probe, we present 3D-Mirage, the first benchmark to combine context variation and precise annotation for real-world illusions with real object exclusions, multi-surface support; purpose-built to stress-test monocular depth on real-world illusions. To score, we propose a second-order magnitude-based evaluation with two metrics: the Deviation Composite Score (DCS) for high second-order 3D structure and the Confusion Composite Score (CCS) for contextual instability. To tame this failure, we introduce Grounded Self-Distillation, a parameter-efficient strategy on Depth-Anything-V2 baseline that surgically targets and resolves hallucination on illusion ROIs while preserving background knowledge, avoiding catastrophic forgetting. Our work provides an innovative pipeline for diagnosing and addressing this phenomenon, urging a necessary shift in the evaluation of MDE from pixel-wise accuracy to structural and contextual robustness.

Read the original paper