Skip to content
AI.info

Research

TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation

Overview Research area: Monocular geometry estimation — specifically single-image surface normal estimation — using generative diffusion/rectified-flow models, with a focus on transparent objects. Tec

TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation
arXiv
2609.06665
Published
2026-09-06
Authors
Mingwei Li, Yi Yang, Hehe Fan

AI summary

Overview

Research area: Monocular geometry estimation — specifically single-image surface normal estimation — using generative diffusion/rectified-flow models, with a focus on transparent objects.

Technical level: Advanced. The paper assumes familiarity with latent diffusion models (VAEs, DiTs, rectified flow), LoRA fine-tuning, spherical statistics (von Mises-Fisher distributions), and classical image-formation models.

Scope: The paper diagnoses a previously unquantified error source in VAE-based latent-diffusion geometry pipelines and proposes a FLUX.2-based framework, TransNormal-2, that corrects it through geometry-aware pixel-space training losses and a post-decode refinement module.

What This Paper Is About

Latent-diffusion models estimate surface normals by encoding images into a compressed latent space and decoding predictions back to pixels, but the VAE's 8× spatial compression blurs the sharp normal changes that occur at object boundaries. The paper measures this effect — even encoding and decoding ground-truth normal maps introduces 1.3°–8.5° of mean angular error, with edge error up to 2.8× the global average — and then builds a system that compensates for it on both the training and inference sides of the VAE decoder. The goal is precise normal estimation, especially on transparent objects where refraction corrupts conventional geometric cues.

Key Contributions

  1. Systematic analysis of VAE reconstruction degradation: The authors quantify, for the first time in this setting, the angular error introduced purely by the VAE encode-decode path when ground-truth normals are passed through it, showing the error is concentrated at object boundaries rather than uniformly distributed.

  2. Geometry-aware pixel-space training objectives: An inverse-rendering self-consistency loss based on Lambertian reflectance, combined with a von Mises-Fisher angular loss and wavelet edge-aware regularization, complements the latent MSE so that decoded predictions are judged on spherical geometry and image formation, not just latent-space distance.

  3. Geometric Refinement Module (GRM): A lightweight, RGB-guided post-decoder module that applies a constrained, gated residual correction to boundary-localized decoding errors, trained with frozen transformer weights.

  4. Strong results across seven benchmarks: TransNormal-2 matches or exceeds MoGe-2 on all eight reported general-scene metrics using only 1.4% as many task-specific normal annotations (122K versus 8.9M samples), and achieves the best reported transparent-object results.

Main Findings

  • VAE degradation is real and boundary-localized: Encoding and decoding ground-truth normals introduces 1.3°–8.5° MAE across four benchmarks. Edge/Global MAE ratios are 2.82× on ClearGrasp (1.29° global vs. 3.64° edge), 1.72× on iBims (8.50° vs. 14.61°), 1.34× on ScanNet (2.45° vs. 3.29°), and 1.26× on NYUv2 (1.81° vs. 2.27°).

  • The bottleneck is architectural, not VAE-specific: The paper states the same diagnostic applied to other latent-diffusion VAEs shows the same edge-concentrated degradation, so the problem is not unique to the FLUX.2 VAE.

  • General-scene parity at a fraction of the annotation cost: TransNormal-2 attains NYUv2 14.7° mean / 62.5% within 11.25°, ScanNet 12.7° / 69.2%, iBims 14.7° / 70.6%, and Sintel 29.2° / 27.5%, with an average rank of 1.4. MoGe-2 scores 14.7° / 62.3°, 12.8° / 68.4°, 14.7° / 70.4°, and 29.3° / 24.8° (average rank 2.3), while using 8.9M training samples versus 122K.

  • Largest gains on transparent objects: TransNormal-2 reaches 11.3° mean MAE on ClearGrasp (65.0% within 11.25°, 94.1% within 30°) and 19.1° on ClearPose (52.3% / 80.5°), an average rank of 1.0. This is a reduction of 4.2° on ClearGrasp and 3.1° on ClearPose over the strongest per-dataset prior baselines; against Lotus-2 specifically, the ClearPose gain is 4.3°.

  • TN-Syn is near-saturated: TransNormal-2 scores 3.6° mean MAE on the TN-Syn transparent-object benchmark, with 95.8% within 11.25° and 99.1% within 30°, leaving little headroom.

  • Single-step inference: The model predicts a normal latent in one deterministic forward pass at a fixed timestep (t=1) with empty-prompt conditioning, without iterative sampling.

  • Decoupled GRM training matters: The paper reports that Phase 2 freezes the core predictor and trains only the GRM, and states that joint training underperforms (reported in Section V-D), because a moving backbone would give the GRM a non-stationary error distribution to fit.

Methodology in Plain English

The authors start from FLUX.2[klein], a 9-billion-parameter rectified-flow Diffusion Transformer, and adapt it with LoRA adapters (rank r=256, α=256) while keeping all pre-trained weights frozen. An RGB image is compressed by the frozen VAE encoder, the DiT predicts a normal-map latent directly in one step, a Local Continuity Module smooths latent-level seams from patchification, and the frozen VAE decoder produces a coarse normal map.

Because that decoder inevitably blurs boundaries, two separate controls are added.

During training (Phase 1, 35K steps): alongside the standard latent MSE against the VAE-encoded ground-truth normal, three pixel-space losses are backpropagated through the frozen decoder. The von Mises-Fisher loss treats normals as unit directions on a sphere and penalizes the angle between prediction and ground truth rather than Euclidean distance. The wavelet loss decomposes both maps with a Haar wavelet and restricts high-frequency supervision to an edge mask derived from ground-truth normal gradients. The inverse-rendering loss estimates a light direction by least squares (with gradients detached), renders Lambertian shading from the predicted normals, and maximizes normalized cross-correlation with the grayscale input — a cue that is naturally robust to unknown albedo and ambient lighting because NCC is scale- and shift-invariant. Training progresses in stages: a general-scene warmup on Hypersim and Virtual KITTI, a continual stage adding transparent-object data, and a final stage enabling the rendering loss.

After decoding (Phase 2): the Geometric Refinement Module takes the coarse normal, a parameter-free anchor, the RGB image, and a binary domain flag. On opaque images the anchor is the coarse normal filtered by an RGB-guided filter, transferring full-resolution edge structure; on transparent images the anchor is the coarse normal itself, because RGB edges are unreliable on refractive surfaces. A small network predicts a per-pixel gated residual on top of the anchor, scaled by a global factor and L2-normalized, so the module can only correct residual error rather than rewrite the coarse prediction. GRM training has two substages: a base substage that learns the residual under a penalty for departing from the anchor, and a shorter calibration substage that tunes the confidence gate to activate on pixels whose anchor error exceeds a threshold while penalizing mean activation.

Training used 8 NVIDIA A100 GPUs (80 GB) with a total batch size of 32 and AdamW with random horizontal flipping. Training data is entirely synthetic, sampled in a 35:15:45:5 ratio from ClearGrasp (45,454 samples), TN-Syn (3,555 training, 395 testing), Hypersim (39,648 samples after filtering, resized to 576×768), and Virtual KITTI (four scenes, 33,580 samples, cropped to 352×1216).

Why This Matters

Research impact: The paper isolates and quantifies an error source that the authors argue is inherited by essentially every VAE-based latent-diffusion geometry method (Marigold, Lotus, E2E-FT and successors), yet has not been quantitatively studied for geometry estimation. Framing it as a boundary-localized blur proportional to the normal jump and the blur width — with an explicit one-dimensional analytical model — gives the community a concrete target for both architecture changes and post-hoc correction, rather than treating VAE artifacts as an unavoidable nuisance.

Real-world applications:

  • Robotic grasping and manipulation: The paper explicitly notes that accurate geometry for transparent objects underpins downstream robotic grasping, where agents must perceive and reason about physical properties; refraction and reflection break conventional depth and normal cues.
  • Scene understanding and augmented reality: Single-image normal estimation feeds shading, relighting, and physical plausibility in mixed-reality rendering, where boundary artifacts are visually conspicuous.
  • Industrial inspection and 3D asset capture: The qualitative examples include mechanical assemblies and gear teeth, where fine structural geometry is the object of interest.
  • Autonomous driving and indoor perception: Evaluation spans Sintel (street scenes) and ScanNet/NYUv2 (indoor scenes), the domains used by most geometry foundation models.

Industry relevance: TransNormal-2 is trained on 122K synthetic samples versus MoGe-2's 8.9M, which the authors frame as a route to competitive geometry estimation under far fewer task-specific normal annotations. Combined with single-step deterministic inference (no iterative sampling), that lowers both the annotation budget and the inference-cost profile relative to multi-step diffusion baselines.

The code is stated to be released at https://longxiang-ai.github.io/TransNormal-2.

Future Directions

  • Extending the degradation analysis to other VAE architectures: The paper notes that the same edge-concentrated degradation appears in other latent-diffusion VAEs, and that VA-VAE, VIVAT, and DC-AE suggest the standard SD-VAE is suboptimal for geometry-sensitive tasks. Quantifying whether geometry-aware VAEs (or pixel-space diffusion) actually close the gap remains open.

  • Video and multi-view transparent-object geometry: The authors observe that DKT repurposes video diffusion models to internalize refraction cues, but operates on video input-output and is therefore not directly comparable to their single-image setting. Whether the same pixel-space corrections transfer to cross-frame modeling is unaddressed.

  • Better treatment of RGB edges on refractive surfaces: The GRM deliberately falls back to the unfiltered coarse normal on transparent images because RGB edges are unreliable there — a limitation that hints at a need for refraction-aware edge cues.

  • Reducing reliance on the domain flag: The refinement module requires a per-image binary transparency flag, set from the source dataset during training and per benchmark at evaluation. Automating this would make the pipeline applicable to unlabeled in-the-wild images without manual domain assignment.

Target Audience

Researchers and graduate students working on monocular geometry estimation, diffusion-based dense prediction, and 3D vision foundations; practitioners building perception systems for robotics or AR/VR that must handle transparent and refractive objects; and engineers interested in how VAE compression artifacts propagate into downstream geometric tasks. The paper is written for readers already comfortable with latent diffusion architectures and spherical/directional statistics — the ablation details, loss coefficients, notation table, and GRM architecture are deferred to the Supplementary Material.

Note: The provided content is truncated mid-way through the discussion of Figure 6, so the ablations referenced as Sections V-D, V-E, and V-F (including the LoRA rank ablation) are cited but their numeric results are not present in the text supplied.

Authors’ abstract

Diffusion-based models enable monocular geometry estimation, yet their pixel-space precision is limited by a shared, under-studied error source: VAE reconstruction degradation. The 8x spatial compression in the VAE encoder-decoder degrades surface normals at object boundaries; even encoding and decoding ground-truth normals introduces 1.3--8.5° of mean angular error (MAE), with edge MAE reaching 2.8x the global MAE. We present TransNormal-2, a FLUX.2-based rectified-flow framework with single-step deterministic inference that addresses this degradation on both sides of the VAE decoder: in how latent predictions are supervised during training, and in how decoded normals are corrected at inference. First, geometry-aware pixel-space losses, including inverse rendering self-consistency, von~Mises-Fisher angular loss, and wavelet edge-aware regularization, complement latent MSE by enforcing spherical normal geometry and diffuse image-formation cues after VAE decoding. Second, a lightweight Geometric Refinement Module (GRM) applies an RGB-guided residual correction to reduce boundary-localized decoding errors without freely rewriting the coarse prediction. On general-scene benchmarks, TransNormal-2 matches or exceeds MoGe-2 on all eight reported metrics while using only 1.4% as many task-specific normal annotations. The gains are clearest for transparent objects, reducing MAE by 4.2° on ClearGrasp and 3.1° on ClearPose over the strongest prior baselines. Code will be released at https://longxiang-ai.github.io/TransNormal-2.

Read the original paper